Multimodal children's voice data processing method based on deep learning and federated learning
Through the multimodal children's voice data processing method of deep learning and federated learning, the problems of multimodal collaborative learning and data privacy protection in the existing technology are solved, and accurate diagnosis of children's voice diseases and safe and efficient data processing are realized.
Patent Information
- Application Number
- CN202510955059.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-11
- Publication Date
- 2025-09-05
- Estimated Expiration
- 2045-07-11
AI Technical Summary
The prior art has problems in the diagnosis of children's voice diseases with low efficiency in single mode information utilization, weak multimodal collaborative learning mechanism, insufficient federated learning application, weak modal alignment and feature interaction capabilities, and high difficulty in model deployment, resulting in insufficient diagnostic accuracy and efficiency.
Multimodal children's voice data processing method based on deep learning and federated learning is adopted, and pre-processing of laryngoscopy images and sound audio data, modal feature extraction and fusion, combined with VisionTransformer classifier model, and distributed training is used to ensure data privacy and model generalization capabilities.
It realizes the accurate classification of multimodal data, improves the accuracy and security of voice disease diagnosis, enhances the generalization ability and adaptability of the model in a multi-center environment, and ensures data privacy protection and compliance.
Smart Images

Figure CN120470245B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of children's voice data processing, and in particular to a multimodal children's voice data processing method based on deep learning and federated learning. Background Art
[0002] Children's voice disorders, such as vocal cord nodules, vocal cord polyps, and chronic laryngitis, are common laryngeal diseases in children, primarily manifesting as hoarseness, vocal fatigue, and abnormal voice tone. Early detection and timely intervention are crucial for the normal development of children's language skills. However, current clinical diagnosis of children's voice disorders still primarily relies on physicians' observation through laryngoscope images and auditory judgment of vocalization audio. This method is highly subjective and heavily reliant on the physician's experience.
[0003] Analyzing laryngoscope images or audio through artificial intelligence technology can assist doctors in diagnosis to a certain extent, but the following technical problems still exist, making it difficult to meet the comprehensive requirements of actual clinical applications for diagnostic accuracy, safety, generalization, and efficiency.
[0004] 1. Low efficiency of utilizing single modal information
[0005] Most current voice disease diagnosis systems are modeled based on single-modal image or audio data, making it difficult to comprehensively characterize pathological features from a multi-dimensional and complementary perspective. The lack of multimodal fusion mechanisms limits the recognition rate of disease differentiation, and the models are particularly unstable in the presence of sample noise or missing data. Existing technologies lack efficient multimodal information fusion mechanisms, making it impossible to fully exploit the complementary relationship between image and audio, thus affecting the accuracy of auxiliary diagnosis.
[0006] 2. Weak multimodal collaborative learning mechanism
[0007] While some studies have initially attempted to jointly model image and audio modalities, most employ static fusion methods, such as concatenating features or combining fixed weights. These methods are unable to dynamically adjust modal contributions based on the variability of input samples, resulting in insufficient responsiveness to diverse disease conditions. Existing fusion methods lack a dynamic modality weighting mechanism, making it impossible to intelligently adjust the proportion of modal information based on different disease manifestations, leading to insufficient model generalization.
[0008] 3. Lack or inadequacy of federated learning applications
[0009] Most current medical AI systems rely on centralized training, requiring the upload of raw data from multiple centers to a unified server, posing a risk of privacy breaches. While some research has introduced federated learning methods, these have primarily focused on single modalities such as images or text, and a multimodal federated learning system suitable for image and audio has yet to be established. Specifically, existing federated learning systems lack a mechanism for jointly training multimodal features of images and audio, making it difficult to achieve efficient fusion and modeling of multi-center data while simultaneously safeguarding data privacy.
[0010] 4. Weak modal alignment and feature interaction capabilities
[0011] The core of multimodal tasks lies in temporal alignment and semantic fusion between modalities. Most current methods ignore the heterogeneity between modalities and lack dedicated modality alignment modules and cross-attention mechanisms. This leads to redundant or distorted fusion information, which in turn affects the model's discriminative capabilities. In other words, the lack of effective modality alignment mechanisms and semantic interaction strategies prevents multimodal features from effectively complementing each other, limiting their fusion expressive power.
[0012] 5. Model deployment is difficult and has poor edge compatibility
[0013] Existing methods mostly use deep neural network models with large parameters, which are difficult to deploy directly on edge terminals or resource-constrained devices, limiting their application in scenarios such as primary healthcare and children's kindergartens.
[0014] That is, the model is large in scale, consumes high training resources, and lacks lightweight optimization strategies, which affects the deployment efficiency and operational stability of the system in the actual environment. Summary of the Invention
[0015] The purpose of the present invention is to overcome the shortcomings of the existing technology and provide a multimodal children's voice data processing method based on deep learning and federated learning, which realizes the accurate classification of children's voice data and can effectively assist in the diagnosis of children's voice diseases.
[0016] The present invention adopts the following technical solutions to achieve the above-mentioned objectives. The present invention provides a multimodal children's voice data processing method based on deep learning and federated learning, comprising:
[0017] S1. Collect laryngoscope image data related to children's voice diseases and corresponding children's vocal audio data;
[0018] S2. Preprocessing the collected laryngoscope image and vocalization audio data;
[0019] S3, extracting and fusing modal features of the preprocessed laryngoscope image and audio data;
[0020] S4. aligning the extracted laryngoscope image modal features with the audio data modal features;
[0021] S5. Based on federated learning, the classifier model is trained using the aligned laryngoscope image modality features and audio data modality features.
[0022] S6. Classify and identify children's voice data using the trained classifier model.
[0023] Furthermore, step S1 specifically includes:
[0024] Laryngoscope image data of children with voice diseases and corresponding children's vocal audio data are collected. Laryngoscope image data is collected by a high-resolution camera, and vocal audio data is obtained by a microphone or voice collection device.
[0025] Furthermore, in step S2, preprocessing the laryngoscope image data specifically includes:
[0026] Basic image processing:
[0027] First, noise reduction is performed on the original laryngoscope image. A combination of Gaussian filtering and median filtering is used to remove random noise and edge glitches in the image. Subsequently, uniform resizing and center cropping operations are performed to standardize the image to a set size. The pixel values are then normalized and mapped to the [0, 1] interval to match the neural network input requirements.
[0028] Image contrast and brightness enhancement:
[0029] Through Log transformation, the details of low-light areas are enhanced and the dynamic range of high-light areas is compressed;
[0030] The overall contrast of laryngoscope images is enhanced by histogram equalization;
[0031] Based on gamma transformation, the overall laryngoscope image contrast is dynamically adjusted according to the different brightness distributions of the laryngoscope image.
[0032] Furthermore, in step S2, preprocessing the audio data specifically includes:
[0033] The audio data is denoised using spectral subtraction, minimum mean square error estimation, Wiener filtering, and signal clarity assessment mechanisms.
[0034] After noise reduction, the original audio data is divided into frames according to a fixed time window, and the amplitude of each frame is normalized;
[0035] The audio signal is converted into spectral features using the Mel-frequency cepstral coefficient method. The specific process is as follows:
[0036] Perform short-time Fourier transform on each frame signal to obtain frequency domain representation;
[0037] Mapping to the Mel frequency scale to simulate the human auditory system's perception of different frequencies;
[0038] Calculate the power spectrum and take its logarithm. Finally, use discrete cosine transform to obtain a set of Mel-frequency cepstral coefficients with fixed dimensions. The Mel-frequency cepstral coefficients of the set order are used as the basic feature representation of the audio.
[0039] Finally, the audio data is enhanced, including pitch shifting, speech speed changes, background noise mixing, and echo and reverberation simulation.
[0040] Furthermore, in step S3, extracting modal features from the pre-processed laryngoscope image specifically includes:
[0041] The DLE module is used to extract local features in the laryngoscope image. The preprocessed laryngoscope image is input into the DLE module, and the local features of the laryngoscope image are output through the DLE module. ;
[0042] The DLE module includes a shallow dense convolution stack unit, a hole convolution path, an edge reinforcement branch and a residual connection mechanism;
[0043] The shallow dense convolution stack unit in the backbone path contains three consecutive 3×3 convolution layers, each with 64 output channels, and the input of each layer is the concatenation of the outputs of all previous layers;
[0044] First convolutional layer: convolution kernel size 3×3, number of input channels 3, number of output channels 64, stride 1, Padding=1, activation function is ReLU, Padding=1 means padding one circle of pixels around the input feature map;
[0045] Second convolutional layer: The convolution kernel size is 3×3, the input is the concatenation of the output of the first convolutional layer and the input image, the number of output channels is 64, and the activation function is ReLU;
[0046] The third convolutional layer has a convolution kernel size of 3×3, and its input is the concatenation of the output of the first convolutional layer, the output of the second convolutional layer, and the input image. The number of output channels is 64, and the activation function is ReLU.
[0047] The dilated convolution path is the first branch path, which includes inserting two 3×3 convolution layers with a dilation rate of 2 or 4 into the main path;
[0048] First dilated convolutional layer: convolution kernel size 3×3, dilation rate = 2, number of channels 64, activation function ReLU;
[0049] Second dilated convolutional layer: convolution kernel size 3×3, dilation rate = 4, number of channels 64, activation function ReLU;
[0050] The edge enhancement branch is a second branch path, comprising an edge feature convolution layer;
[0051] Edge detection operation: Perform Sobel operator processing on the input image to obtain the edge response map;
[0052] Edge feature convolution layer: convolution kernel size 1×1, number of output channels 64, activation function ReLU;
[0053] The residual connection mechanism includes adding the original input or the output of the previous module back to the main channel after passing it through a 1×1 convolution.
[0054] Furthermore, in step S3, extracting modal features from the pre-processed laryngoscope image specifically includes:
[0055] The GSA module is used to extract the global features of the laryngoscope image, and the local features of the laryngoscope image output by the DLE module are converted to Input the GSA module and output the global features of the laryngoscope image through the GSA module ;
[0056] The GSA module structure is as follows:
[0057] Patch division and coding layer:
[0058] The input local feature map is divided into several fixed-size patches. Each patch is encoded into a token vector through average pooling and 1×1 convolution, and finally a serialized token vector is obtained.
[0059] Spatial-aware position encoding layer:
[0060] A relative position encoding mechanism is introduced to define the relative position offset between tokens. The offset is mapped into a learnable bias vector and used in conjunction with token similarity to enhance spatial distribution modeling capabilities.
[0061] Global attention aggregation layer:
[0062] Use a single-head dual-channel attention structure, the first channel is structural relevance attention, the second channel is semantic response attention, the two-channel attention results are spliced and merged through MLP;
[0063] Token fusion and back-projection layer:
[0064] Reconstruct the aggregated token features into a two-dimensional feature map , represented as a global feature map.
[0065] Furthermore, the modal feature fusion of the preprocessed laryngoscope image specifically includes:
[0066] The MSFE module is used to fuse the local features and global features of the laryngoscope image. With global feature map Input the MSFE module to obtain the final fusion features of the laryngoscope image ;
[0067] The MSFE module consists of a multi-scale transformation layer, a feature fusion layer, a contextual attention weighting layer, and a scale alignment layer;
[0068] The multi-scale transformation layer contains three sub-layers of different scales:
[0069] The first scaling sub-layer uses the identity operation, that is, the transformation operation with exactly the same input and output, to maintain the size of the original image;
[0070] The second scale sub-layer uses average pooling and 1×1 convolution operations, with a pooling kernel of 2×2, a stride of 2, and an output channel of 256 to obtain the mesoscale structure;
[0071] The third scale sub-layer uses average pooling and 1×1 convolution operations, with a pooling kernel of 4×4, a stride of 4, and an output channel of 256, to model large-scale lesion morphology and background structure;
[0072] The three different scale sub-layers act on and , a total of 6 downsampled feature maps are generated, including 3 local feature maps and 3 global feature maps;
[0073] The feature fusion layer fuses the local feature map generated by the scale transformation layer with the global feature map at each scale. The fusion method is as follows:
[0074] ;
[0075] Where, Represents the local feature map at scale S, represents the global feature map at scale S, Represents a splicing operation, represents a 1×1 convolution operation, Represents the fused feature map at scale S;
[0076] Each scale is integrated once, and the final output is 、 、 , Represents the feature map of the first scale, 1 represents the first scale number, Represents the feature map of the second scale, 2 represents the second scale number, Feature map at the third scale, 4 represents the third scale number;
[0077] The fused feature map is dynamically weighted through the contextual attention weighting layer, and the spliced local and global feature maps are fed into a shared two-layer fully connected network. Softmax is used to obtain the attention weight. The attention weight scale is consistent with the fused feature map. The attention weight is applied to the fused feature map to achieve channel and spatial joint attention enhancement. The method is as follows:
[0078] ;
[0079] Where, represents the attention weight, Represents the fused feature maps after weighting at different scales;
[0080] Each scale is weighted once and the final output is , Represents the weighted fusion feature map at the first scale, Represents the weighted fusion feature map at the second scale, Represents the weighted fusion feature map at the third scale;
[0081] In the scale alignment layer, the weighted multi-scale fusion feature maps are spliced in the channel dimension to form the final fusion representation:
[0082] , Represents the final fused feature map after splicing;
[0083] Finally, it is compressed by 1×1 convolution , get the final fusion features of the laryngoscope image .
[0084] Furthermore, in step S3, the modal feature extraction and fusion of the pre-processed audio data specifically includes:
[0085] The AMFN module is used to extract local and global features of audio data from the preprocessed Mel-spectrogram and fuse the local and global features.
[0086] The AMFN module includes an MSFE-A module, a TCM module and an LGA module;
[0087] The local features of the audio data are extracted through the MSFE-A module, and the Mel spectrum of the audio data is input into the MSFE-A module. The audio features under different frequency receptive fields are extracted respectively through multiple parallel convolution branches of the MSFE-A module, including a low-frequency convolution branch with a convolution kernel size of 3×7, a medium-frequency convolution branch with a convolution kernel size of 5×5, and a high-frequency convolution branch with a convolution kernel size of 7×3. The output feature maps of each branch are spliced in the channel dimension and fused through 1×1 convolution to obtain the local feature map of the audio data. ;
[0088] The global features of the audio data are extracted through the TCM module, and the local feature maps of the audio data are converted into The TCM module is used to capture the long-distance time-dependent features in the audio signal. The spectrum is averaged and pooled in the frequency dimension. Then, one-dimensional time convolution is used in combination with a gating mechanism to generate the corresponding attention weight map. The attention weight map is multiplied with the local feature map to obtain the global feature map of the audio data. ;
[0089] The local feature map is obtained by LGA module With global feature map The outputs of the MSFE-A module and the TCM module are fused and spliced in the channel dimension. The channel attention mechanism and the frequency selection gating mechanism are introduced to perform channel compression and spectral importance weighting on the fused feature map to form the final fused audio modal feature. .
[0090] Furthermore, step S4 specifically includes:
[0091] Final fusion features of laryngoscope images by contrastive learning method and the final fused audio modal features Align and calculate and The similarity between them is calculated and the cross-modal feature alignment is optimized through the contrastive loss function.
[0092] Furthermore, step S5 specifically includes:
[0093] Final fusion features based on laryngoscope images and the final fused audio modal features Train the VisionTransformer classifier model. and The input is embedded into the mapping layer, uniformly mapped to the multi-dimensional feature space, and then used as the input of the Vision Transformer classifier model for multiple rounds of iterative training;
[0094] Federated learning technology is used during training, with multiple data centers completing local training separately and uploading the classifier model parameters to the server for aggregation, thus achieving distributed training without sharing the original data. During the federated learning process, a differential privacy perturbation mechanism is used to add noise to the local model parameters or gradients to protect the privacy of user data during the upload process.
[0095] The beneficial effects of the present invention are:
[0096] Compared to traditional single-modality diagnostic methods, this invention achieves multimodal data fusion by combining the complementary properties of images and audio. This fusion fully utilizes the different strengths of audio and images in the feature space, significantly improving the accuracy of voice disease data classification.
[0097] This invention uses federated learning technology to ensure the privacy protection of multi-center medical data and avoid the direct transmission of medical data. By training models in a distributed environment, data compliance and privacy security are guaranteed, while promoting cross-institutional data collaboration.
[0098] The present invention achieves stable training in the case of modality loss and uneven data distribution through a modality-aware federated aggregation strategy, significantly improving the generalization ability of the model in a multi-center environment. The federated personalization mechanism ensures the local diagnostic needs of different hospitals or kindergartens, and improves the practicality and adaptability of the model. The federated comparative learning mechanism improves the consistency of cross-modal feature alignment and ensures the semantic space uniformity of images and audio at different nodes. The federated differential privacy mechanism enhances the security of data transmission and model sharing processes, meeting medical data compliance requirements.
[0099] Based on the Vision Transformer classifier model, this invention can efficiently process complex laryngoscopic image data related to pediatric voice disorders, as well as the corresponding children's vocalization audio data. Through multimodal feature fusion, the classifier can extract deep features from the input data, providing fast and accurate classification results.
[0100] This invention introduces dynamic modal weighting based on an attention mechanism, which automatically adjusts modal weights based on the audio and image features in different diagnostic scenarios. This mechanism adaptively optimizes performance in processing different types of voice data, improving the accuracy and reliability of auxiliary diagnosis. BRIEF DESCRIPTION OF THE DRAWINGS
[0101] Figure 1 This is a flow chart of a multimodal children's voice data processing method based on deep learning and federated learning provided by the present invention. DETAILED DESCRIPTION
[0102] To make the objectives, technical solutions and advantages of the embodiments of the present invention more clear, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention.
[0103] like Figure 1 As shown, the present invention provides a multimodal children's voice data processing method based on deep learning and federated learning, which specifically includes the following steps:
[0104] Step 1: Acquisition of laryngoscope images and vocalization audio data;
[0105] During the implementation of the present invention, it is necessary to jointly collect laryngoscope images and vocal audio related to children's voice diseases. The collection process includes the following aspects:
[0106] 1. Collection Objects
[0107] The data were collected from children aged 3 to 12, covering the critical period of voice development from preschool to elementary school. All data came from children clinically diagnosed with voice disorders, including but not limited to the following three main conditions:
[0108] Vocal cord paralysis: manifested as weak voice and heavy breath;
[0109] Vocal cord polyps: Common in children who overuse their voices and have abnormal voice tone;
[0110] Vocal cord nodules: Early symptoms include hoarseness and significant vocal fatigue.
[0111] 2. Laryngoscope Image Acquisition
[0112] Collection equipment: Use an electronic laryngoscope system with high-definition electronic imaging capabilities (such as Olympus ENF-VQ series, Storz HDScope, etc.);
[0113] Imaging parameters: Image resolution should be no less than 1920 × 1080, frame rate should be controlled above 30 fps, and cold light source should be used to avoid thermal stimulation.
[0114] Filming method: Children are instructed to pronounce standard vowels (e.g., "ah"), and the laryngeal structure is ensured to be clear and intact during filming, with no obstruction of the glottal structure in the image;
[0115] Operation specification: The operation is performed by a qualified otolaryngologist, and the collection process is completed within 1 minute.
[0116] 3. Vocal audio collection
[0117] Capture equipment: Use a professional voice recording device (such as a Zoom H4n or Shure SM7B recording device);
[0118] Audio parameters: sampling rate 44.1kHz, 16-bit quantization accuracy, mono recording;
[0119] Recording environment: Sampling was carried out in a standard soundproof room with background noise below 30dB;
[0120] Voice content design:
[0121] Basic vowel combinations ( / a / , / e / , / i / , / o / , / u / );
[0122] Short numbers (1-10) and everyday phrases (such as "My name is Xiao Ming", "Please listen to my voice");
[0123] The length of each speech sample is controlled at 15 to 30 seconds, containing multiple morpheme changes.
[0124] 4. Multimodal synchronization mechanism
[0125] To ensure consistency between laryngoscope images and audio, the sampling system is equipped with a unified timestamp marking mechanism (using the RTC module to synchronize the image acquisition and audio recording systems), and all modality data use a unified coding naming method (such as ChildID_0001_T1) to achieve accurate pairing of cross-modality data.
[0126] 5. Data labeling and management
[0127] All samples were jointly diagnosed and confirmed by no less than three senior otolaryngologists;
[0128] After data collection, a dedicated person will establish a sample management database, marking the sample source, collection time, modality correspondence and disease label.
[0129] 6. Ethics and Privacy Compliance
[0130] Before data collection, written consent must be obtained from the guardian of the child being tested, and the guardian must sign an informed consent form;
[0131] The entire process has passed medical ethics review and filing to ensure the legal and compliant use of data;
[0132] All samples were anonymized after collection, and only the anonymous ID labels were retained for subsequent modeling and analysis.
[0133] Step 2: Data preprocessing;
[0134] Laryngoscope image data preprocessing:
[0135] To improve the expression quality and recognition stability of laryngoscope images in subsequent deep neural networks, this paper constructs a phased, modular image preprocessing and enhancement strategy, covering traditional image processing methods and deep learning-based image generation and enhancement methods. Specifically, it includes the following contents:
[0136] 1. Basic image processing
[0137] First, noise suppression is performed on the original laryngoscope image. A combination of Gaussian filtering and median filtering is used to remove random noise and edge burrs in the image. Then, uniform resizing and center cropping operations are performed to standardize the image to a size of 224×224. Finally, pixel values are normalized and mapped to the interval [0,1] to match the neural network input requirements.
[0138] 2. Image contrast and brightness enhancement
[0139] To adapt to non-ideal imaging conditions such as dark occlusion and highlight reflection in laryngoscope images, the present invention introduces three nonlinear image enhancement transformations:
[0140] Log transform enhances details in low-light areas and compresses the dynamic range in high-light areas. The transformation formula is as follows:
[0141] ;
[0142] Where, Represents the transformed pixel value, that is, the enhanced image grayscale value, Represents a scaling factor (constant) used to adjust the contrast range of the transformed image, usually used to normalize or map the result to a desired pixel value range. Represents the pixel adjustment factor, a parameter used to enhance the effect of nonlinear transformation and control the sensitivity or responsiveness of the transformation.
[0143] The image enhancement method based on Log transform can effectively enhance the details of low-brightness areas, suppress overexposure in bright areas, and improve the overall contrast of the image. It is suitable for the enhancement needs of faint lesion areas in medical images.
[0144] Histogram equalization enhances the overall contrast of the image by redistributing the grayscale frequency. It is suitable for improving the structural clarity of the lesion area. The formula is as follows:
[0145] ;
[0146] Where, Represents the grayscale value after processing, Represents the input grayscale value of the original image, Represents the number of pixels with gray level j, n represents the total number of pixels in the image, and L represents the number of gray levels in the image.
[0147] The histogram equalization method increases image contrast by redistributing the grayscale histogram, thereby enhancing the overall contrast of the image. It is particularly suitable for improving image clarity and structural detail recognition in children's laryngeal lesion areas.
[0148] Gamma transform dynamically adjusts the overall image contrast according to different brightness distributions and is suitable for nonlinear image brightness correction. The formula is as follows:
[0149] ;
[0150] Where, Indicates the gamma index, which is used to control the brightness enhancement level. When the dark details are enhanced, When , highlights are enhanced.
[0151] 3. Traditional data enhancement operations
[0152] This paper uses two image enhancement libraries, Augmentor and Imgaug, to perform various transformation operations on images, including:
[0153] Geometric enhancement: such as ±10° random rotation, horizontal mirror flip, and affine stretching;
[0154] Color perturbation: random brightness and contrast adjustment to simulate different lighting conditions;
[0155] Occlusion simulation: random noise and occlusion blocks are added to enhance the model's robustness to interference.
[0156] 4. Synthetic Feature Space Enhancement (SMOTE)
[0157] To alleviate the imbalanced distribution of voice disease categories in the training set, this paper uses the SMOTE (Synthetic Minority Over-sampling Technique) algorithm to generate synthetic samples in the image feature space. The synthesis strategy is as follows:
[0158] ;
[0159] Where, Represents the generated new sample feature vector, represents the feature vector of the original minority class sample, represents a neighboring sample vector of x in the feature space, represents the random interpolation factor.
[0160] In this way, the SMOTE algorithm generates new synthetic samples in the feature space, enabling the model to better learn the distribution characteristics of minority classes during training, thereby improving the overall classification performance.
[0161] 5. Generative Adversarial Enhancement Mechanism
[0162] In order to further expand image diversity and simulate the distribution of multi-morphological lesions, this paper introduces the Generative Adversarial Network (GAN) and its improved ConSinGAN (Contextual Single Image GAN) to synthesize lesion area images. By using a local area of a single image as the generator input, ConSinGAN can generate style, texture and original Figure 1 New sample images with rich structural details are used to enhance the model's robust modeling capabilities for lesions with fuzzy boundaries and unclear morphology.
[0163] 6. Dynamic Enhancement (Mixup) during Training
[0164] During the training process, the present invention adopts the Mixup strategy to generate new sample and label combinations, effectively improving the model's ability to learn transitions in feature interpolation regions. The calculation method is as follows:
[0165] ;
[0166] ;
[0167] Where, and Represent the input feature vectors of two different samples in the original training set, and They represent the label vectors of the corresponding samples (usually in one-hot encoding form), λ∈[0,1] is the mixing coefficient, which is usually sampled from the Beta distribution and controls the mixing ratio of the two samples. 、 Represents the synthesized new samples and their labels.
[0168] The above-mentioned multi-level image enhancement strategy is automatically triggered and combined by the image preprocessing module. The specific strategy is dynamically adjusted according to the image brightness distribution, texture complexity and category weight, achieving the goals of diversity modeling of multi-lesion samples, weak feature enhancement and inter-class balanced training.
[0169] Audio data preprocessing:
[0170] After denoising the audio data, the audio signal is converted into spectral features using MFCC (Mel-Frequency Cepstral Coefficients). The audio data is also enhanced by varying the pitch, tempo, and noise to increase the model's adaptability to diverse audio. This preprocessing includes the following key steps:
[0171] Denoising:
[0172] Considering the interference factors such as strong breathing sounds, background conversation sounds, and equipment noise in children's vocalization, the present invention adopts the following noise suppression algorithm to improve the signal-to-noise ratio of subsequent feature extraction:
[0173] Spectral Subtraction: Based on noise estimation of non-speech segments, the noise power spectrum of each frame is calculated in real time and subtracted from the total power spectrum to extract the pure speech component.
[0174] Minimum Mean Square Error Estimation (MMSE): uses the statistical distribution of speech and noise to estimate the maximum a posteriori probability of a clean speech signal, suitable for dynamic noise environments;
[0175] Wiener Filter: Combines the signal-to-noise power spectrum ratio to optimize spectrum recovery, and is particularly suitable for enhancing low-frequency resonance information.
[0176] Signal clarity evaluation mechanism: Uses short-term clarity indicators (such as Log-Spectral Distance, LSD) to dynamically select the above algorithm combination.
[0177] The above noise reduction processing ensures that environmental interference is effectively reduced and speech fidelity is improved without damaging the speech structure.
[0178] Audio segmentation and normalization:
[0179] Audio segmentation and amplitude normalization:
[0180] Frame division mechanism: Each audio segment is divided into frames with a 25 millisecond window and a 10 millisecond frame shift to extract local time segments;
[0181] Window function: Hamming window is used to improve the smoothness of time domain boundaries;
[0182] Amplitude normalization: Perform Min-Max normalization or Z-Score normalization on each frame of audio to ensure energy consistency of input samples at different recording intensities and suppress loudness deviation.
[0183] This module ensures that the speech signal has a uniform dynamic range before entering the spectrum conversion module.
[0184] Mel-frequency cepstral coefficient extraction:
[0185] To convert the audio signal into a two-dimensional spectrogram that can be processed by the model, the present invention uses MFCC to extract spectral features. The MFCC extraction process includes:
[0186] Perform short-time Fourier transform on each frame signal to obtain frequency domain representation;
[0187] Mapping to the Mel frequency scale to simulate the human auditory system's perception of different frequencies;
[0188] The power spectrum is calculated and its logarithm is taken. Finally, a set of Mel-frequency cepstral coefficients with fixed dimensions is obtained through discrete cosine transform. Usually, the first 13 or 20 MFCC coefficients are taken as the basic feature representation of the audio.
[0189] During calculation, for each frame of audio, first calculate its power spectrum, then map it to the Mel frequency axis, and finally obtain the nth MFCC feature through discrete cosine transform:
[0190] Audio Enhancement:
[0191] To improve the model's generalization across different ages, vocalization styles, and recording scenarios, this paper introduces several speech data enhancement techniques:
[0192] Pitch Shifting: Simulates pitch changes caused by differences in vocal cord tension and aging by raising or lowering the pitch by ±1 to 2 semitones.
[0193] Time Stretching: Adjusts the speaking speed to 0.9 to 1.1 times the original speed, keeping the pitch unchanged, to simulate the pronunciation difference between fast and slow speech;
[0194] Background noise mixing: Select noise clips from a preset noise library (including medical scenes, human voices, and shuffling papers), and control the mixed signal-to-noise ratio (SNR) between 15 and 30 dB to improve interference resistance;
[0195] Echo and Reverberation Simulation: Simulates the acoustic environments of different physical spaces (such as clinics and classrooms) by convolving the actual room response function (RIR).
[0196] To further enhance the audio modality’s ability to perceive disease-related time-frequency structures, the present invention also introduces the following advanced enhancement strategies:
[0197] Spectrum Augmentation: Randomly block several continuous bands of the MFCC image on the time axis or frequency axis to simulate band loss, instantaneous silence, etc., to enhance the robustness of the model;
[0198] Speed Perturbation + SNR Control: Jointly adjusts speech rate and signal-to-noise ratio to construct a diverse sample matrix, improving the model's adaptability to extreme samples.
[0199] Audio alignment and unified formatting:
[0200] After MFCC processing, all audio clips are uniformly cropped or padded to a fixed length (e.g., 3 seconds, corresponding to approximately 300 frames);
[0201] Use Zero Padding and Frame Repetition methods to ensure consistent data batch input structure;
[0202] Sample meta-information (such as original duration and enhanced scheduling labels) is added for subsequent dynamic attention mechanism modeling.
[0203] Step 3: Feature extraction and fusion;
[0204] To achieve multi-scale and multi-structure joint modeling of children's laryngoscope images and vocal audio, the present invention independently designed a cross-modal feature extraction and fusion framework consisting of local structure extraction, global semantic modeling, scale fusion, and context alignment modules, and constructed an end-to-end processing network with structural innovation and task-specificity. The details are as follows:
[0205] DLE (Dense Local Extractor) module, used for local feature extraction of images;
[0206] GSA (Global Semantic Aggregator) module, used for global feature extraction of images;
[0207] The MSFE (Multi-Scale Fusion Encoder) unit is used to fuse local and global image features;
[0208] AMFN (Audio Multi-scale Fusion Network);
[0209] The audio multi-scale fusion network includes the MSFE-A module, the TCM module, and the LGA module. The MSFE-A module is used for local audio feature extraction, the TCM module is used for global audio feature extraction, and the LGA module is used for audio fusion.
[0210] The present invention realizes two-level modeling from local fine-grained information to overall semantic relationships by dividing the image feature extraction module into the DLE module and the GSA module. The DLE module uses dense convolution, multi-channel fusion and edge enhancement mechanisms to focus on extracting the boundaries and texture features of small lesion areas such as vocal cord nodules and polyps. The GSA module captures the contextual relationships between different anatomical regions through serialized representation and spatial semantic attention mechanism, and enhances the ability to understand the global structure of the image. The two modules work together to ensure that the model has the ability to discriminate generalized structures while maintaining high-resolution details, significantly improving the accuracy and stability of voice-related lesion identification.
[0211] The DLE module includes shallow dense convolution stacking units, void convolution paths, edge enhancement branches and residual connection mechanisms, which extract fine-grained texture and boundary features while maintaining the spatial resolution of the input image.
[0212] 1. Shallow dense convolution stacking unit
[0213] It consists of three consecutive 3×3 convolutional layers, each with 64 output channels. The input of each layer is the concatenation of the outputs of all previous layers, that is:
[0214] ;
[0215] Where, represents the convolution operation of the 3×3 convolution kernel, Indicates the l The output of the convolutional layer, the input is the concatenation of the outputs of all previous layers, including the initial input, Indicates the input of the DLE module, represents the output of the first convolutional layer, Indicates the The output of a convolutional layer.
[0216] This structure is similar to the dense connections in DenseNet, but is limited to shallow stacking to avoid over-complexity. Each convolutional layer is followed by batch normalization and the ReLU activation function to achieve step-by-step detail extraction while improving gradient transfer efficiency and feature reuse.
[0217] 2. Dilated Convolution Path
[0218] Insert two 3×3 convolutional layers with dilation rate of 2 or 4 in the main path;
[0219] Purpose: Expand the receptive field to cover slightly larger lesion boundaries while maintaining the same resolution.
[0220] 3. Edge Strengthening Branch
[0221] Perform Sobel gradient filtering on the input image (or use Learnable Edge Attention);
[0222] The gradient map is used as an auxiliary channel to participate in feature fusion and highlight edge areas such as the glottis and fissure;
[0223] Fusion method:
[0224] ;
[0225] Where, represents the local structural feature map extracted by the backbone path, represents the edge feature map extracted by the edge enhancement branch, express and The fused feature map.
[0226] 4. Residual Connection Mechanism
[0227] The original input or the output of the previous module is added back to the main channel after 1×1 convolution;
[0228] Avoid deep feature offset and enhance low-level feature retention capabilities.
[0229] The final output is a feature map: ;
[0230] H and W remain consistent with the original image. H represents the height of the feature map, and W represents the width of the feature map to avoid loss of spatial information. C is the number of output channels (such as 128 or 256). The specific setting depends on the backbone structure. Each position vector encodes the texture, boundary, and lesion activation response of the area.
[0231] The advantages of the DLE module are as follows:
[0232] Multi-layer 3×3 convolution stacking, deep and fine-grained feature extraction, and enhanced edge detection capabilities;
[0233] Dense connections promote feature sharing and gradient flow, alleviating information bottlenecks;
[0234] The dilated convolution path introduces a large receptive field perception capability to identify large-area blurred lesions;
[0235] Edge-guided branch to improve detection of microstructural areas (e.g., glottic fissures, marginal polyps);
[0236] Residual channel bridging preserves the original spatial information of the input and improves the efficiency of feature fusion;
[0237] The structure of the DLE module is as follows:
[0238] Input size: 224 × 224 × 3 (normalized RGB image);
[0239] Backbone path: shallow dense convolution stacking unit;
[0240] Convolutional layer 1 (D-Conv1): convolution kernel size 3×3, number of input channels 3, number of output channels 64, stride 1, padding = 1, activation function is ReLU;
[0241] Convolutional layer 2 (D-Conv2): The convolution kernel size is 3×3, the input is the concatenation of D-Conv1 and the input image (a total of 67 channels), the output channel number is 64, and the activation function is ReLU;
[0242] Convolutional layer 3 (D-Conv3): convolution kernel size 3×3, input is the concatenation of D-Conv1, D-Conv2 and the input image (a total of 131 channels), output channels number 64, activation function ReLU;
[0243] Output splicing features: The output splicing of convolution layer 1, convolution layer 2, and convolution layer 3 is used as a dense feature map , the size is 224×224×192 (or unified to 256 channels after compression).
[0244] Branch path 1: dilated convolution path;
[0245] Dilated-Conv1: kernel size 3×3, dilation rate = 2, number of channels 64, activation function ReLU;
[0246] Dilated-Conv2 (optional): kernel size 3×3, dilation ratio = 4, number of channels 64, activation function ReLU;
[0247] The output feature map size remains 224×224×64.
[0248] Branch path 2: edge reinforcement branch;
[0249] Edge detection operation: Perform Sobel operator processing on input image II to obtain edge response map , size is 224×224×1;
[0250] Edge-Conv: convolution kernel size 1×1, 64 output channels, ReLU activation function;
[0251] The output edge feature map size is 224×224×64.
[0252] Feature fusion and output stage:
[0253] Splicing and fusion: splicing the three path outputs in the channel dimension;
[0254] The fusion size is: 224×224×320;
[0255] Fuse-Conv: convolution kernel size 1×1, 320 input channels, 256 output channels, ReLU activation function;
[0256] Residual connection: The input image is convolved with 1×1 and then element-wise added to the fusion feature:
[0257] ;
[0258] Final output size: 224 × 224 × 256.
[0259] The GSA module is a custom neural network module designed to model structural relationships and global semantic information between long-range regions in images. Unlike traditional self-attention architectures, the GSA module emphasizes lightweight global modeling and structural location awareness to accommodate complex objects with irregular spatial distribution and diverse scales in medical images.
[0260] The GSA module, through spatial partitioning, position encoding, dual-channel attention modeling, and back-projection, enables modeling of structural semantic relationships between distant regions in the image. This module fully captures the contextual information between anatomical regions in laryngeal images, making it significantly valuable for identifying overall abnormal patterns.
[0261] The GSA module focuses on long-distance dependency modeling and global structure understanding of images, and is used to capture the cross-regional semantic relationships between the internal structures of the laryngeal cavity.
[0262] The input of the GSA module is the local feature map output by the DLE module , usual size: 224×224×256.
[0263] The GSA module structure is as follows:
[0264] 1. Patch Division and Coding Layer
[0265] Function: Convert the two-dimensional feature map into a patch sequence so that subsequent modules can model the relationship between regions;
[0266] Divide the input feature map into several small patches of fixed size, such as 16×16;
[0267] Each patch is encoded into a token (vector) through average pooling and 1×1 convolution;
[0268] Get the serialized token vector:
[0269] ;
[0270] Among them, X is the token sequence of all patches of the image, represents the encoding representation vector of the i-th patch, , N represents the number of patches obtained by image division, and D is the projection dimension (such as 256).
[0271] Division method: Divide the input feature map into patches of size 16×16, a total of 14×14=196 patches;
[0272] Operation: Each patch is linearly projected through a 1×1 convolution and encoded into a vector of length 256;
[0273] Output size: 196×256 (i.e. 196 tokens, each with 256 dimensions);
[0274] 2. Spatial Aware Position Encoding Layer
[0275] Function: Introducing structure-aware position bias to enable the model to have relative spatial position awareness and improve modeling accuracy;
[0276] Compared with standard position encoding, this layer introduces a relative position encoding mechanism to define the relative position offset between tokens;
[0277] The offset is mapped to a learnable bias vector and used in conjunction with token similarity to enhance the spatial distribution modeling capability. This mechanism is more suitable for anatomical structures in medical images that are rotated, deformed, and irregularly arranged.
[0278] Implementation: Construct a learnable relative position offset matrix to represent the position relationship between tokens;
[0279] Operation content: The offset matrix is added to the structure-related attention as a gain term of the attention score;
[0280] Output size: 196×256 (encoded token sequence with position information);
[0281] 3. Global Attention Aggregation Layer
[0282] Function: Modeling spatial structure correlation and semantic activation intensity separately, so that the model can focus on local structure and semantic global distribution at the same time;
[0283] The independently designed lightweight attention module does not use multi-head attention, but uses a single-head dual-channel attention structure:
[0284] Channel 1: structural relevance attention;
[0285] Input: token sequence (196×256)
[0286] operate:
[0287] Linear_Q: 256→256, Linear_Q represents a linear transformation layer used to generate a query vector (Q is Query);
[0288] Linear_K: 256→256, Linear_K represents a linear transformation layer used to generate a key vector (K is Key);
[0289] Linear_V: 256→256, Linear_V represents a linear transformation layer used to generate a value vector (V is Value);
[0290] Calculate the structural attention score: ;
[0291] Where, is a learnable relative position bias matrix used to enhance the attention mechanism’s ability to model the relative relationship between token spaces. Represents the structural relevance attention matrix, and the score after Softmax indicates the degree of structural attention each token pays to other tokens.
[0292] Get the structure attention output: , represents the structural attention output;
[0293] Channel 2: Semantic Affinity Attention;
[0294] Same structure, independent weights, focusing on semantically active areas;
[0295] Output: , Represents semantic response attention output;
[0296] The two attention results are concatenated and merged through MLP:
[0297] , Represents the merged output;
[0298] Output size: 196 × 256;
[0299] 4. Token Fusion and Back-Projection
[0300] Function: Reconstruct the aggregated token features into a two-dimensional feature map and restore the spatial topology to facilitate subsequent fusion and classification processing.
[0301] After all tokens are fused, they are mapped back to the original spatial dimension through MLP;
[0302] Operation method:
[0303] will sequence Reduction to a two-dimensional structure;
[0304] Each token corresponds to a 16×16 region, which is reconstructed by upsampling or deconvolution;
[0305] Back projection module:
[0306] Use 1×1 convolution to expand the channel to 512 (optional);
[0307] Use PixelShuffle or upsampling to reconstruct into a spatial feature map;
[0308] Output:
[0309] A global image feature map with cross-regional semantic understanding capabilities, emphasizing structural hierarchical relationships and overall abnormal area modeling.
[0310] Reorganize the serialized token into a two-dimensional feature map , as a global feature representation.
[0311] The MSFE module proposed in this paper is used to perform multi-scale hierarchical fusion of local feature maps (output by the DLE module) and global feature maps (output by the GSA module) to achieve the joint expression of detail information and semantic structure, thereby improving the detection and recognition capabilities of lesion areas of different scales (such as tiny nodules and large areas of congestion) in laryngoscope images.
[0312] The MSFE module constructs a multi-scale pyramid representation of local and global feature maps to capture the spatial structure, boundary information and semantic context of the image at different resolutions, and redistributes the weights of the fusion results through the channel-level attention mechanism. This significantly enhances the model's perception of multi-scale lesions (such as tiny polyps and extensive congestion) in laryngoscope images while maintaining feature richness.
[0313] Local and global collaborative perception: Fusing the detail edge information extracted by the DLE module with the contextual structure information extracted by the GSA module;
[0314] Multi-scale spatial adaptability: The three-scale sampling structure can capture the responses of small, medium, and large lesion areas at different resolutions;
[0315] Dynamic attention enhancement mechanism: The contextual attention module can significantly improve the model's perception of key areas (lesions, asymmetric structures);
[0316] Scale alignment and unified expression: Multi-scale features are uniformly upsampled and concatenated to ensure that the final features have cross-scale structural consistency.
[0317] MSFE module input:
[0318] Local feature map: ;
[0319] Global feature map: ;
[0320] Default dimensions: H=W=224, C=256.
[0321] MSFE module structure consists of:
[0322] 1. Multi-scale transformation layer
[0323] To capture the lesion response characteristics at different spatial resolutions, the MSFE module first constructs multi-scale representations for local and global feature maps. Specifically, it includes three scales:
[0324] Scale number, downsampling, proportional output, size description:
[0325] S = 1, original size, 224 × 224 × 256, retaining high-resolution detail features;
[0326] S = 2, downsampling,2×112 × 112 × 256, to capture mesoscale structures, such as blurred boundary areas;
[0327] S = 4, downsampling, 4 × 56 × 56 × 256, extracting large-scale background or symmetrical structures;
[0328] Implementation: Use 3×3 convolution + stride + padding, or average pooling + 1×1 convolution for dimensionality reduction at each scale;
[0329] Output: , represents the height of the feature map, Indicates the width of the feature map to avoid loss of spatial information, is the number of output channels, S represents the scale;
[0330] The specific implementation is shown in Table 1 below:
[0331] Table 1: Multi-scale transformation layer
[0332]
[0333] Each scale acts on and , a total of 6 downsampled feature maps (3 local and 3 global) are generated.
[0334] 2. Feature Fusion Layer
[0335] At each scale S, local and global features are fused separately to integrate micro-boundary features and macro-structural relationships.
[0336] Main fusion method: channel splicing + 1×1 convolution compression
[0337] ;
[0338] Where, Represents the local feature map at scale S, represents the global feature map at scale S, Represents a splicing operation, represents a 1×1 convolution operation, Represents the fused feature map at scale S;
[0339] The unified channel number after fusion is 256.
[0340] The specific implementation is shown in Table 2 below:
[0341] Table 2: Feature fusion layer
[0342]
[0343] Each scale is fused once and three feature maps are output: 、 、 .
[0344] 3. Contextual Attention Weighted Layer
[0345] To enhance the response to key lesion areas (such as edge micro-protrusions and glottal asymmetry), the MSFE module introduces a contextual attention mechanism to dynamically weight the fused feature map:
[0346] The local and global features are concatenated and fed into a shared two-layer fully connected network (FC).
[0347] Use Softmax to obtain attention weight map , dimension and consistent;
[0348] Apply the weight map to the fused feature map to achieve channel-space joint attention enhancement:
[0349] ;
[0350] The specific implementation is as shown in Table 3:
[0351] Table 3: Contextual Attention Weighted Layer
[0352]
[0353] The attention mechanism is a lightweight channel attention that enhances key semantic areas, such as lesion boundaries and morphological mutation areas.
[0354] 4. Scale alignment and fusion integration layer
[0355] Fusion feature maps of all scales Upsampled (if S>1) to 224×224 resolution;
[0356] Use bilinear interpolation or PixelShuffle method to restore spatial dimensions;
[0357] The specific implementation is shown in Table 4:
[0358] Table 4: Scale alignment layer
[0359]
[0360] All three scales are restored to 224×224×256, and the spatial dimension is unified for integration.
[0361] Multi-scale feature maps are concatenated in the channel dimension to form the final fused representation:
[0362] ;
[0363] If the number of output channels is too large (such as 768), 1×1 convolution is used to compress it to 256 or 512 channels as the final image feature representation. ;
[0364] Finally, through the fusion integration layer, the specific implementation is as shown in Table 5:
[0365] Table 5: Fusion Integration Layer
[0366]
[0367] Final output:
[0368] , is the number of channels after fusion (such as 256 or 512);
[0369] It represents the joint representation of the global semantics and local boundary information of the image modality under multi-scale alignment, and will subsequently be cross-modally aligned with the audio modality features.
[0370] Audio feature extraction:
[0371] The feature extraction of audio modalities in the present invention adopts a customized audio multi-scale fusion network to extract rich local and global temporal features from the preprocessed Mel-spectrogram, thereby achieving accurate modeling of spectral anomalies (such as frequency drift, energy loss, pitch instability, etc.) that may exist in children's voices.
[0372] AMFN includes MSFE-A module, TCM module and LGA module.
[0373] The Multi-Scale Frequency Encoder for Audio (MSFE-A) module extracts features at different frequency receptive fields using multiple parallel convolution branches. These branches include a low-frequency convolution branch (with a 3×7 kernel), a mid-frequency convolution branch (with a 5×5 kernel), and a high-frequency convolution branch (with a 7×3 kernel). The feature maps output by each branch are concatenated in the channel dimension and fused using a 1×1 convolution to form a preliminary frequency feature map.
[0374] The TCM (Temporal Context Modeling) module is used to capture long-range temporal dependencies in audio signals. It performs average pooling on the frequency dimension of the spectrogram, then uses one-dimensional temporal convolution combined with a gating mechanism to generate an attention weight map, and performs contextual enhancement on the local frequency feature map.
[0375] The Local-Global Attention Fusion (LGA) module concatenates the outputs of the MSFE-A and TCM modules in the channel dimension. It also introduces a channel attention mechanism and a frequency-selective gating mechanism to perform channel compression and spectral importance weighting on the fused feature maps to form the final audio modality representation.
[0376] The final output size of the AMFN module is 128×64×128, which is suitable for subsequent modality alignment, joint classification or contrastive learning tasks.
[0377] Compared with traditional audio feature extraction methods based on residual networks or Transformer structures, this invention has the following technical advantages by designing a three-stage audio multi-scale fusion network:
[0378] The multi-scale frequency convolution structure extracts high-frequency, mid-frequency, and low-frequency information, significantly enhancing the model's ability to perceive abnormal spectral areas. This is particularly useful for problems such as the loss of high-frequency details in children's throats or vocal cord resonance drift.
[0379] Use lightweight convolutional attention instead of Transformer to model long-term dependencies, avoid introducing too many parameters, and improve model stability;
[0380] The introduction of channel attention and frequency selection mechanisms achieves structural adaptive enhancement in feature expression, which helps the model focus on feature areas that are more discriminative for lesion identification.
[0381] Audio local feature extraction:
[0382] The MSFE-A module uses three parallel branches to extract low-frequency, mid-frequency, and high-frequency energy features, simulating changes in the vocal cord spectral structure. As the local feature modeling module of the audio modal feature extraction network, the MSFE-A module focuses on extracting structural details at different frequency scales in the Mel-spectrogram, including key pathological features such as energy concentration areas, formant bands, and frequency mutation regions.
[0383] The MSFE-A module utilizes a multi-scale frequency receptive field structure to simulate the human ear's sensitivity to different frequency bands, enabling hierarchical perception of spectral structure. The low-frequency branch perceives the overall tonal contour, the mid-frequency branch focuses on the direction of formant peaks, and the high-frequency branch enhances edge mutation responses in the affected area. After channel concatenation, compressed convolution is used to integrate the convolutional layers into a unified semantic space, providing high-resolution input for subsequent global modeling and modal alignment.
[0384] Input Mel-spectrogram: , represents the input mel spectrogram;
[0385] T=128: T represents the number of time frames, 4F=64: F represents the number of Mel frequency channels, single-channel grayscale image (1 represents the number of channels);
[0386] Structural branches (layer by layer)
[0387] MSFE-A consists of three parallel convolution branches, which extract low-frequency, mid-frequency, and high-frequency structural patterns respectively.
[0388] The structure of branch 1 is shown in Table 6:
[0389] Table 6: Low-Frequency Branch
[0390]
[0391] The structure of branch 2 is shown in Table 7:
[0392] Table 7: Mid-Frequency Branch
[0393]
[0394] The structure of branch three is shown in Table 8:
[0395] Table 8: High-Frequency Branch
[0396]
[0397] Fusion stage (Concat + Conv):
[0398] Channel stitching:
[0399] , Indicates the low-frequency channel output, Indicates the intermediate frequency channel output, Represents the high-frequency channel output, Represents the output after channel splicing;
[0400] Channel-compressed convolution:
[0401] Operation: 1×1 convolution;
[0402] Number of input channels: 192, number of output channels: 128;
[0403] Activation function: ReLU;
[0404] Output:
[0405] ;
[0406] Audio global feature extraction
[0407] The TCM module simulates the structural changes of audio over long periods of time, extracting semantic features such as pitch fluctuations, speech rate, and inter-frame resonance trends. Built on the temporal dimension of the spectrogram, this module uses lightweight convolution and gated attention mechanisms to model inter-frame contextual relationships, replacing the traditional Transformer architecture with advantages such as lightweight, efficient, and patentable.
[0408] TCM module input: ;
[0409] The implementation steps are as follows:
[0410] Step 1: Frequency Compression
[0411] The spectrogram is average-pooled along the frequency dimension, allowing the model to focus more on structural modeling in the time dimension and simulate the temporal evolution of the speech frame sequence. The specific implementation is shown in Table 9 below.
[0412] Table 9: Frequency-direction compression layers
[0413] Layer Name Operation Type parameter Output size Function AvgPool-F Average Pooling kernel size=(1×64) 128 ×1×128 Compress along the frequency axis to preserve the time series distribution
[0414] Get the compressed time series feature representation: ;
[0415] Step 2: Temporal Convolutional Context Modeling
[0416] One-dimensional temporal convolution is used to extract inter-frame contextual dependency information, replacing the QK (Query-Key) attention mechanism in the standard Transformer. The specific implementation is shown in Table 10.
[0417] Table 10: Temporal Convolutional Context Modeling Layer
[0418]
[0419] The output is recorded as: ;
[0420] Step 3: Gated Attention Weight Generation
[0421] The original sequence features are gated and fused with the convolutional context to form frame-level attention weights.
[0422] Concatenate two features:
[0423] ;
[0424] Fully connected → Activation (Sigmoid):
[0425] ;
[0426] Where b represents the bias vector, is the weight matrix of the fully connected layer or linear mapping layer, and is also a trainable parameter. The two together determine the generated mapping relationship of the attention weight. After Sigmoid activation, the output Represents the temporal attention weight of each frame, which is used to perform time-dependent enhancement on the local feature map.
[0427] Step 4: Attention Weighting
[0428] The generated attention weights are used to perform weighted amplification of the original input feature map on a time-frame basis to enhance the key structure frame response.
[0429] Attention map Broadcast to the frequency dimension, with Multiplication:
[0430] ;
[0431] The output size is the same as the input: ;
[0432] Representation: An audio feature map with enhanced temporal dependencies for fusion with local feature maps.
[0433] The innovations of the TCM module are as follows:
[0434] Frequency compression → time modeling, using spectrum mean to express abstract frame states and reduce redundant modeling;
[0435] Convolution replaces attention and simulates the receptive field of self-attention, but the calculation is more stable and controllable;
[0436] The gating mechanism generates frame weights to amplify key frame features without the need for QK calculations, thus avoiding dependence on the Transformer structure.
[0437] Local and global feature fusion mechanism:
[0438] In the audio modality, in order to improve the joint modeling capability of the local time-frequency details and the global temporal structure of the audio, the present invention also designs a fusion mechanism of local and global features.
[0439] The LGA module is used to fuse the output feature maps from the MSFE-A module and the TCM module, enhance them through channel attention and frequency gating mechanisms, and output a unified high-quality audio feature representation.
[0440] The innovations of the LGA module are as follows:
[0441] Channel attention (non-SE) finely models the importance of each spectral channel, making it lightweight and efficient;
[0442] Frequency Selective Gating (FSG) is particularly suitable for the problem of children's audio lesions being concentrated in frequency bands.
[0443] Autonomous splicing and fusion + multi-weighted structure. Not based on Transformer or ResNet, with independent structural protection value.
[0444] enter:
[0445] Local feature map: ;
[0446] Global feature map: ;
[0447] The specific fusion process is as follows:
[0448] Step 1: Feature splicing;
[0449] The feature concatenation layer is shown in Table 11:
[0450] Table 11: Feature concatenation layer
[0451] Layer Name Operation Type parameter Output size Function Concat-LG Splicing Channel dimension stitching 128 × 64 × 256 Joint modeling of local and global features
[0452] Step 2: Channel Attention Enhancement
[0453] A lightweight channel attention mechanism is introduced to dynamically improve the representation strength of key semantic channels (such as frequency resonance and energy jitter).
[0454] 1. Global Average Pooling
[0455] Operation: Yes Perform global average pooling to obtain the channel description vector;
[0456] 2. Fully connected mapping (FC → ReLU → FC → Sigmoid);
[0457] The compression ratio r=4, that is, the middle hidden layer is 64-dimensional;
[0458] Output channel attention weight vector ;
[0459] 3. Attention Weighting
[0460] The channel attention weight vector Acts on The channel dimension of :
[0461] ;
[0462] Step 3: Frequency selective gating mechanism;
[0463] It is used to enhance the response of specific frequency bands (such as 3kHz–6kHz) in the spectrogram, and has stronger structural recognition of common areas of vocal cord lesions in children.
[0464] 1. Average pooling (along the time axis):
[0465] Output frequency description vector ;
[0466] 2. Learnable frequency weight vector:
[0467] Defining trainable vectors , initialized to all 1s;
[0468] Get the frequency weighted graph ;
[0469] 3. Broadcast and multiplication:
[0470] Will The broadcast is 128×64×1128×64×1, acting on the output of the previous stage:
[0471] ;
[0472] Step 4: Channel compression and standardization;
[0473] Use 1×1 convolution for channel compression, compressing the dimension from 256 to 128:
[0474] ;
[0475] BatchNorm + ReLU can be added (optional)
[0476] Output feature map: ;
[0477] Represents: Final feature map of audio modality after structural enhancement;
[0478] It can be subsequently used for image-audio modality alignment (such as cross-modal contrastive learning) and classifier input.
[0479] Step 4: Feature alignment;
[0480] Contrastive learning and feature alignment: Final fusion features of laryngoscope images via contrastive learning method and the final fused audio modal features Align and calculate and The similarity between them is optimized through contrast loss function to ensure that the features of different modalities remain consistent in the shared feature space.
[0481] The present invention adopts a customized cross-modal contrastive learning framework based on the principles of InfoNCE and dynamic weighting mechanism. The dynamic weighting and customized improvement of cross-modal contrast loss introduced by it are particularly suitable for cross-modal feature learning in voice disease data processing.
[0482] Design of cross-modal contrast loss:
[0483] C.1 A positive sample pair refers to an image and audio sample of the same condition. For example, a child's throat image and the corresponding voice sample are both from a case of muscle tension voice disorder, and their features should be brought closer.
[0484] Image features: features extracted by the image modality feature extraction network , Right now ;
[0485] Audio features: Features extracted by the audio modality feature extraction network , Right now ;
[0486] The goal of the positive pair is to ensure and As close as possible in the embedding space.
[0487] c.2 Negative pairs are images and audio samples from different conditions. For example, images of muscle tension voice disorder and audio samples of vocal cord nodules. They should be as far apart as possible in feature space.
[0488] c.3 Cross-modal contrast loss function design;
[0489] For the contrastive loss between image and audio modalities, the following contrastive loss function can be used:
[0490] ;
[0491] Where, represents the contrast loss function, N is the batch size, and are the image modality and audio modality feature vectors of the i-th sample, represents the Euclidean distance, is the temperature coefficient (hyperparameter) that controls the smoothness of the loss. The loss function helps the model align the features of the corresponding image and audio while increasing the distance between different samples.
[0492] There is a logarithmic part in the loss function, which is similar to the information contrast loss and can help the model better separate features of different categories.
[0493] c.4 Introducing dynamic weighting
[0494] Because features from different modalities may differ in expressive power, a dynamic weighting mechanism is introduced to optimize cross-modal contrast loss. During training, an attention mechanism is trained to dynamically adjust the weights of each modality's features in the loss function based on their importance.
[0495]
[0496] Where i represents the current anchor image sample index, j represents the traversal of all audio sample indexes (including positive and negative samples), represents the feature representation of the i-th image modality sample, represents the feature representation of the j-th audio modality sample, represents the feature representation of the i-th audio modality sample, represents the weight of the image modality, Represents the audio modality weight.
[0497] In voice disease data processing, cross-modal contrastive learning can effectively improve feature alignment between image and audio modalities, thereby enhancing classification accuracy. By designing an appropriate contrastive loss function and incorporating cross-entropy loss during training, the model can better learn diagnostic features extracted from both audio and images, providing a more precise auxiliary tool for clinical practice.
[0498] Step 5: Model training and optimization;
[0499] To protect the privacy of medical data, this invention utilizes federated learning technology during model training. Local data from multiple hospitals is not uploaded directly to a central server. Instead, only model parameters are uploaded after local training. This approach prevents the leakage of sensitive data while ensuring data privacy across different medical institutions.
[0500] Federated learning technology is used during training. To adapt to the inconsistent distribution of multimodal inputs across different medical centers, modality loss, privacy sensitivity, and other issues, the following strategies are combined to improve training stability and security:
[0501] 1. Modality-aware federated update mechanism
[0502] For each client node, check whether it has both image and audio modality data and train it according to the following strategy:
[0503] If it is a dual-modal node, complete image + audio + fusion network training is performed;
[0504] If it is a single-modal node, only the corresponding modal branch is trained, and the sub-network parameters where no modality appears are frozen to avoid gradient drift;
[0505] During aggregation, the modal presence matrix is used to adjust the parameter update weights of each branch to prevent a modal branch from being strongly dominated by a unilateral node.
[0506] 2. Federal Personalization Strategy
[0507] To adapt to feature shifts caused by differences in data collection conditions, laryngoscope equipment, and language habits among centers, a personalized federation mechanism is adopted:
[0508] Use FedBN: During the aggregation phase, only the convolutional / attention backbone network parameters are synchronized, and each center maintains independent BN parameters, which improves the accuracy in non-IID scenarios;
[0509] At the same time, the classification head is set to local personalization, and only the shared backbone network is aggregated, which effectively prevents the problem of inconsistent distribution of prediction targets of different institutions.
[0510] 3. Privacy protection mechanism
[0511] A differential privacy perturbation mechanism is introduced to add Gaussian noise to gradients or parameters before uploading them, achieving mathematical-level protection of patient privacy information during federated training.
[0512] The federated server uses encrypted aggregation operations based on homomorphic encryption to ensure that the server cannot directly obtain local model weights, thereby enhancing privacy protection capabilities.
[0513] 4. Federated Contrastive Learning Extensions:
[0514] A cross-modal contrastive loss is introduced as one of the local optimization objectives in federated training to ensure spatial consistency of alignment between different centers by maintaining the distance constraints of image or audio modalities in the shared embedding space.
[0515] Each node uses NT-Xent loss locally to compare modal features, and the federated server synchronously aligns the encoder sub-network parameters.
[0516] Model training in this step includes:
[0517] Based on voice audio data and laryngoscope image data collected from multiple hospitals, the DLE module + GSA module (image modality feature extraction network) and the AMFN module (audio modality feature extraction network) were trained.
[0518] Final fusion features based on image modality Final fusion features with audio modality , train the feature alignment module;
[0519] Final fusion features based on laryngoscope images and the final fused audio modal features Train the VisionTransformer classifier model. and The two modalities are jointly input into the embedding mapping layer and uniformly mapped to a 768-dimensional feature space before being used as the input of the Vision Transformer classification model. This process ensures that the two modalities are aligned at the input layer, providing a foundation for subsequent joint modeling and classification training.
[0520] Training set source: Based on voice and laryngoscopic image data collected from multiple hospitals;
[0521] Optimization algorithm: AdamW optimizer is used, with an initial learning rate of 1e-4 and a learning rate cosine annealing schedule.
[0522] Evaluation indicators: Use multiple indicators such as accuracy, AUC, F1-score, etc.
[0523] Training rounds: The standard setting is 50 to 100 rounds, which is dynamically adjusted based on the performance of the validation set.
[0524] The classifier model will undergo multiple rounds of iterative optimization on a large-scale joint audio-image dataset, and its performance will be evaluated on an independent test set. By adjusting hyperparameters, the model's generalization capabilities will be further improved to ensure its efficiency and accuracy in real-world applications.
[0525] Step 6: Data classification and processing;
[0526] The fused features are input into the trained classifier model for classification.
[0527] Vision Transformer Classifier:
[0528] This paper introduces the Vision Transformer (ViT) as a multimodal fusion classifier within the audio-image joint diagnosis framework. It utilizes image and audio embeddings optimized through contrastive learning to classify voice disease data. The ViT consists of a cross-attention mechanism module, a ViT backbone network, and a classification output layer, demonstrating excellent cross-modal alignment and global dependency modeling capabilities.
[0529] Feature fusion and cross-attention mechanism
[0530] To further enhance the interaction between modal features, this paper also designs a cross-attention fusion module. The fusion method is as follows:
[0531] ;
[0532] in, and are the weight matrices of audio and image features respectively, It is the fused feature representation.
[0533] The Attention mechanism is used to capture the coupling characteristics of the two modalities in the space-frequency dimension and achieve cross-modal structural alignment.
[0534] This module contains 2 layers of cross-attention units (each layer contains 1 attention head), and the output is used for subsequent Transformer encoder processing.
[0535] Vision Transformer structural parameters
[0536] The ViT backbone network structure adopted by the present invention is as follows:
[0537] Input block method: fusion features Divide into a Patch sequence of length 16, generating a total of about 49 tokens;
[0538] Positional encoding: Introducing learnable positional encoding to preserve temporal and spatial order;
[0539] Transformer encoder layer: a total of 6 layers of encoder, each layer contains:
[0540] Multi-head self-attention module (8 heads)
[0541] Feedforward network module (MLP layer width is 3072 dimensions)
[0542] Residual connections and LayerNorm;
[0543] Embedding dimension: Each token dimension is 768;
[0544] Dropout rate: 0.1, used to prevent overfitting.
[0545] During the training process, the ViT network can effectively model the global contextual relationship of the input sequence and improve the overall recognition ability of complex voice abnormality patterns.
[0546] The fused multimodal features are input into the fully connected layer for classification prediction to determine the patient's voice disease category. During the training process, the model uses the cross entropy loss function As the main optimization target, cross-modal contrast loss is introduced at the same time And the joint loss function To ensure that audio and image features remain consistent during training, the specific optimization goals are:
[0547] ;
[0548] in, is the cross entropy loss function, is the cross-modal contrast loss, For joint losses, and is the balance parameter.
[0549] The foregoing description is merely a preferred embodiment of the present invention. It should be understood that the present invention is not limited to the form disclosed herein and should not be construed as excluding other embodiments. Rather, the present invention can be used in various other combinations, modifications, and environments and can be modified within the scope of the concept described herein through the above teachings or techniques or knowledge in the relevant field. Modifications and variations made by those skilled in the art that do not depart from the spirit and scope of the present invention are intended to be protected by the appended claims.
Claims
1. A multimodal children's voice data processing method based on deep learning and federated learning, characterized by: include: S1. Collect laryngoscope image data related to children's voice diseases and corresponding children's vocal audio data; S2. Preprocessing the collected laryngoscope image and vocalization audio data; S3, extracting and fusing modal features of the preprocessed laryngoscope image and audio data; The modal feature extraction of the preprocessed laryngoscope image specifically includes: The DLE module is used to extract local features in the laryngoscope image. The preprocessed laryngoscope image is input into the DLE module, and the local features of the laryngoscope image are output through the DLE module. ; The DLE module includes a shallow dense convolution stack unit, a hole convolution path, an edge reinforcement branch and a residual connection mechanism; The shallow dense convolution stack unit is the backbone path, which contains three consecutive 3×3 convolution layers, each with 64 output channels, and the input of each layer is the concatenation of the outputs of all previous layers; First convolutional layer: convolution kernel size 3×3, number of input channels 3, number of output channels 64, stride 1, Padding=1, activation function is ReLU, Padding=1 means padding one circle of pixels around the input feature map; Second convolutional layer: The convolution kernel size is 3×3, the input is the concatenation of the output of the first convolutional layer and the input image, the number of output channels is 64, and the activation function is ReLU; The third convolutional layer has a convolution kernel size of 3×3, and its input is the concatenation of the output of the first convolutional layer, the output of the second convolutional layer, and the input image. The number of output channels is 64, and the activation function is ReLU. The dilated convolution path is the first branch path, which includes inserting two 3×3 convolution layers with a dilation rate of 2 or 4 into the main path; First dilated convolutional layer: convolution kernel size 3×3, dilation rate = 2, number of channels 64, activation function ReLU; Second dilated convolutional layer: convolution kernel size 3×3, dilation rate = 4, number of channels 64, activation function ReLU; The edge enhancement branch is a second branch path, comprising an edge feature convolution layer; Edge detection operation: Perform Sobel operator processing on the input image to obtain the edge response map; Edge feature convolution layer: convolution kernel size 1×1, number of output channels 64, activation function ReLU; The residual connection mechanism is to add the original input or the output of the previous module back to the main channel after passing it through 1×1 convolution; Modal feature extraction and fusion of preprocessed audio data specifically include: The AMFN module is used to extract local and global features of audio data from the preprocessed Mel-spectrogram and fuse the local and global features. The AMFN module includes an MSFE-A module, a TCM module and an LGA module; The local features of the audio data are extracted through the MSFE-A module, and the Mel spectrum of the audio data is input into the MSFE-A module. The audio features under different frequency receptive fields are extracted respectively through multiple parallel convolution branches of the MSFE-A module, including a low-frequency convolution branch with a convolution kernel size of 3×7, a medium-frequency convolution branch with a convolution kernel size of 5×5, and a high-frequency convolution branch with a convolution kernel size of 7×3. The output feature maps of each branch are spliced in the channel dimension and fused through 1×1 convolution to obtain the local feature map of the audio data. ; The global feature map of the audio data is extracted by the TCM module, and the local feature map of the audio data is converted into The TCM module is used to capture the long-distance time-dependent features in the audio signal. The spectrum is averaged and pooled in the frequency dimension. Then, one-dimensional time convolution is used in combination with a gating mechanism to generate the corresponding attention weight map. The attention weight map is multiplied with the local feature map to obtain the global feature map of the audio data. ; The local feature map is obtained by LGA module With global feature map The outputs of the MSFE-A module and the TCM module are fused and spliced in the channel dimension. The channel attention mechanism and the frequency selection gating mechanism are introduced to perform channel compression and spectral importance weighting on the fused feature map to form the final fused audio modal feature. ; S4. aligning the extracted laryngoscope image modal features with the audio data modal features; S5. Based on federated learning, the classifier model is trained using the aligned laryngoscope image modality features and audio data modality features. S6. Classify and identify children's voice data using the trained classifier model.
2. The multimodal children's voice data processing method based on deep learning and federated learning according to claim 1 is characterized in that: Step S1 specifically includes: Laryngoscope image data of children with voice diseases and corresponding children's vocal audio data are collected. Laryngoscope image data is collected by a high-resolution camera, and vocal audio data is obtained by a microphone or voice collection device.
3. The multimodal children's voice data processing method based on deep learning and federated learning according to claim 1 is characterized in that: In step S2, preprocessing the laryngoscope image data specifically includes: Basic image processing: First, noise reduction is performed on the original laryngoscope image. A combination of Gaussian filtering and median filtering is used to remove random noise and edge glitches in the image. Subsequently, uniform resizing and center cropping operations are performed to standardize the image to a set size. The pixel values are then normalized and mapped to the [0, 1] interval to match the neural network input requirements. Image contrast and brightness enhancement: Through Log transformation, the details of low-light areas of the image are enhanced and the dynamic range of high-light areas of the image is compressed; The overall contrast of laryngoscope images is enhanced by histogram equalization; Based on gamma transformation, the overall laryngoscope image contrast is dynamically adjusted according to the different brightness distributions of the laryngoscope image.
4. The multimodal children's voice data processing method based on deep learning and federated learning according to claim 1 is characterized in that: In step S2, preprocessing the audio data specifically includes: The audio data is denoised using spectral subtraction, minimum mean square error estimation, Wiener filtering, and signal clarity assessment mechanisms. After noise reduction, the original audio data is divided into frames according to a fixed time window, and the amplitude of each frame is normalized; The audio signal is converted into spectral features using the Mel-frequency cepstral coefficient method. The specific process is as follows: Perform short-time Fourier transform on each frame signal to obtain frequency domain representation; Mapping to the Mel frequency scale to simulate the human auditory system's perception of different frequencies; Calculate the power spectrum and take its logarithm. Finally, use discrete cosine transform to obtain a set of Mel-frequency cepstral coefficients with fixed dimensions. The Mel-frequency cepstral coefficients of the set order are used as the basic feature representation of the audio. Finally, the audio data is enhanced, including pitch shifting, speech speed changes, background noise mixing, and echo and reverberation simulation.
5. The multimodal children's voice data processing method based on deep learning and federated learning according to claim 1 is characterized in that: In step S3, extracting modal features from the pre-processed laryngoscope image specifically includes: The GSA module is used to extract the global features of the laryngoscope image, and the local features of the laryngoscope image output by the DLE module are converted to Input the GSA module and output the global features of the laryngoscope image through the GSA module ; The GSA module structure is as follows: Patch division and coding layer: The input local feature map is divided into several fixed-size patches. Each patch is encoded into a token vector through average pooling and 1×1 convolution, and finally a serialized token vector is obtained. Spatial-aware position encoding layer: A relative position encoding mechanism is introduced to define the relative position offset between tokens. The offset is mapped into a learnable bias vector and used in conjunction with token similarity to enhance spatial distribution modeling capabilities. Global attention aggregation layer: Use a single-head dual-channel attention structure, the first channel is structural relevance attention, the second channel is semantic response attention, the two-channel attention results are spliced and merged through MLP; Token fusion and back-projection layer: Reconstruct the aggregated token features into a two-dimensional feature map , represented as a global feature map.
6. The multimodal children's voice data processing method based on deep learning and federated learning according to claim 5 is characterized in that: In step S3, the modal feature fusion of the pre-processed laryngoscope image specifically includes: The MSFE module is used to fuse the local features and global features of the laryngoscope image. With global feature map Input the MSFE module to obtain the final fusion features of the laryngoscope image ; The MSFE module consists of a multi-scale transformation layer, a feature fusion layer, a contextual attention weighting layer, and a scale alignment layer; The multi-scale transformation layer contains three sub-layers of different scales: The first scaling sub-layer uses the identity operation, that is, the transformation operation with exactly the same input and output, to maintain the size of the original image; The second scale sub-layer uses average pooling and 1×1 convolution operations, with a pooling kernel of 2×2, a stride of 2, and an output channel of 256 to obtain the mesoscale structure; The third scale sub-layer uses average pooling and 1×1 convolution operations, with a pooling kernel of 4×4, a stride of 4, and an output channel of 256, to model large-scale lesion morphology and background structure; The three different scale sub-layers act on and , a total of 6 downsampled feature maps are generated, including 3 local feature maps and 3 global feature maps; The feature fusion layer fuses the local feature map generated by the scale transformation layer with the global feature map at each scale. The fusion method is as follows: ; Where, Represents the local feature map at scale S, represents the global feature map at scale S, Represents a splicing operation, represents a 1×1 convolution operation, Represents the fused feature map at scale S; Each scale is integrated once, and the final output is 、 、 , Represents the feature map of the first scale, 1 represents the first scale number, Represents the feature map of the second scale, 2 represents the second scale number, Feature map at the third scale, 4 represents the third scale number; The fused feature map is dynamically weighted through the contextual attention weighting layer, and the spliced local and global feature maps are fed into a shared two-layer fully connected network. Softmax is used to obtain the attention weight. The attention weight scale is consistent with the fused feature map. The attention weight is applied to the fused feature map to achieve channel and spatial joint attention enhancement. The method is as follows: ; Where, represents the attention weight, Represents the fused feature maps after weighting at different scales; Each scale is weighted once and the final output is , Represents the weighted fusion feature map at the first scale, Represents the weighted fusion feature map at the second scale, Represents the weighted fusion feature map at the third scale; In the scale alignment layer, the weighted multi-scale fusion feature maps are spliced in the channel dimension to form the final fusion representation: , Represents the final fused feature map after splicing; Finally, it is compressed by 1×1 convolution , get the final fusion features of the laryngoscope image .
7. The multimodal children's voice data processing method based on deep learning and federated learning according to claim 6 is characterized in that: Step S4 specifically includes: Final fusion features of laryngoscope images by contrastive learning method and the final fused audio modal features Align and calculate and The similarity between them is calculated and the cross-modal feature alignment is optimized through the contrastive loss function.
8. The multimodal children's voice data processing method based on deep learning and federated learning according to claim 7 is characterized in that: Step S5 specifically includes: Final fusion features based on laryngoscope images and the final fused audio modal features Train the VisionTransformer classifier model. and The input is embedded into the mapping layer, uniformly mapped to the multi-dimensional feature space, and then used as the input of the Vision Transformer classifier model for multiple rounds of iterative training; Federated learning technology is used during training, with multiple data centers completing local training separately and uploading the classifier model parameters to the server for aggregation, thus achieving distributed training without sharing the original data. During the federated learning process, a differential privacy perturbation mechanism is used to add noise to the local model parameters or gradients to protect the privacy of user data during the upload process.
Citation Information
Patent Citations
Laryngoscope image multi-attribute classification method based on multi-modal information fusion
CN116664929A
Vocal cord lesion detection method and system based on multi-modal information fusion
CN119128786A