Microphone defect detection method and system based on image model

Through the combination of multimodal data flow fusion, physical optimization and deep learning, the shortcomings of traditional microphone defect detection methods in data processing, model accuracy and generalization capabilities are solved, and efficient microphone defect detection is achieved.

CN120107212AInactive Publication Date: 2025-06-06KEITHY INNOVATION TECH CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510188257.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-20
Publication Date
2025-06-06
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Existing microphone defect detection methods rely on single modal data, making it difficult to handle complex relationships between multimodal data, resulting in low recognition accuracy and poor adaptability.

Method used

By acquiring multimodal data such as images, vibration waveforms and sound pressure signals, performing spatiotemporal alignment processing, multi-material coupling model is used to use the preset MEMS microphone structure, and applying physical-GAN enhancement optimization to generate a microphone defect coupling model. Then, domain-invariant feature extraction and feature network construction are carried out, and a dual-branch feature decoupling network and meta-learning framework are combined to construct an image microphone detection defect model.

Benefits of technology

It effectively solves the problem of inaccurate time alignment between multimodal data, reduces noise interference, improves the accuracy and reliability of the model, and improves the accuracy, efficiency and flexibility of defect detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120107212A_ABST
    Figure CN120107212A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of image recognition, in particular to a microphone defect detection method and system based on an image model. The method comprises the following steps: acquiring a multi-modal data set; performing data preprocessing on the multi-modal data set to generate a space-time alignment multi-modal data stream; performing multi-material coupling modeling on the time-space alignment multi-modal data flow by using a preset MEMS microphone structure, and performing physical-GAN enhancement optimization to generate a microphone defect coupling model; therefore, through combination of multi-modal data stream fusion, physical optimization and deep learning, the defects of a traditional microphone defect detection method in data processing, model accuracy and generalization ability are overcome, and the defect detection precision, efficiency and reliability are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image recognition, and in particular to a method and system for detecting microphone defects based on an image model. Background Art

[0002] The detection method based on acoustic signals and vibration signals relies on manually set feature extraction rules and simple classification algorithms. Its recognition ability is limited by noise interference and data quality, and the recognition accuracy of complex defects is low. Secondly, the detection model of traditional technology lacks the ability to adapt to specific defects, and the efficiency and accuracy of defect detection in different microphone types or working environments are poor. Thirdly, most of the existing microphone defect detection systems rely on test data in laboratory environments. In practical applications, especially in noisy environments, the detection accuracy and reliability are often affected. In addition, in the defect detection process, most of the existing technologies rely on single-mode data input, which leads to a lack of deep learning capabilities for complex correlations between multiple different features. For example, the defects of microphones are manifested as the joint effects of multiple signals such as images, vibrations, and sound pressure, but the existing systems are often unable to effectively process the complex relationships between these multimodal data, resulting in the inability of the model to fully identify potential defects. Moreover, the traditional model is based on predefined features for classification, ignoring the adaptive learning ability of the model in diverse actual scenarios, and it is difficult to adapt to changes in different environments. Summary of the invention

[0003] Based on this, it is necessary to provide a microphone defect detection method and system based on an image model to solve at least one of the above technical problems.

[0004] To achieve the above object, a microphone defect detection method based on an image model is provided, the method comprising the following steps:

[0005] Step S1: obtaining a multimodal data set; performing data preprocessing on the multimodal data set to generate a spatiotemporally aligned multimodal data stream;

[0006] Step S2: Use the preset MEMS microphone structure to perform multi-material coupling modeling on the spatiotemporal aligned multimodal data stream, and perform physical-GAN enhanced optimization to generate a microphone defect coupling model;

[0007] Step S3: extract domain-invariant features of the microphone defect coupling model, and construct a feature network to generate a microphone double-branch feature decoupling network; build a meta-learning framework for the microphone double-branch feature decoupling network to the microphone defect coupling model, and perform image space interpolation optimization to construct an image microphone detection defect model;

[0008] Step S4: Perform microphone defect detection evaluation based on the image microphone defect detection model and generate a microphone image model evaluation report.

[0009] The beneficial effect of the present invention is that by acquiring multimodal data such as images, vibration waveforms and sound pressure signals, and preprocessing these data, a time-space aligned data stream is generated. This process ensures the synchronization of data from different sources in time, and provides unified input data for subsequent model training. This measure effectively solves the problem of inaccurate time alignment between multimodal data and reduces noise interference between data. By presetting the MEMS microphone structure, multi-material coupling modeling is performed on the multimodal data stream, and the physical-generative adversarial network (GAN) is applied for optimization, and the generated microphone defect coupling model has high accuracy and reliability. Physical-GAN enhanced optimization not only improves the generalization ability of the model, but also can effectively capture the interaction between different materials, avoiding the errors and deficiencies of the single data source model. By extracting domain-invariant features and constructing feature networks for the defect coupling model, combining the dual-branch feature decoupling network with the meta-learning framework, the distinguishability and stability of the defect features are further enhanced. Image space interpolation optimization improves the quality of image data, ensures that the fusion of images and other modal data is more accurate, and finally constructs an efficient image microphone defect detection model. The detection model is used for evaluation and a microphone image model evaluation report is generated, which provides a quantitative detection effect for practical applications. This process avoids the shortcomings of defect detection models in traditional methods that are difficult to process multimodal data by integrating multiple optimization strategies, and significantly improves the accuracy of detection and the flexibility of application. Therefore, the present invention solves the deficiencies of traditional microphone defect detection methods in data processing, model accuracy and generalization ability through the combination of multimodal data stream fusion, physical optimization and deep learning, and improves the accuracy, efficiency and reliability of defect detection.

[0010] Preferably, step S1 comprises the following steps:

[0011] Step S11: acquiring a multimodal data set, wherein the multimodal data set includes a microphone image sequence, a microphone vibration waveform, and a microphone sound pressure signal;

[0012] Step S12: calculating the zero-crossing rate of the multimodal data set to obtain a zero-crossing rate signal set; performing nonlinear quantization processing on the zero-crossing rate signal set to obtain microphone multimodal sampling quantization data;

[0013] Step S13: preprocess the microphone multimodal sampling quantized data to generate a time-space aligned multimodal data stream.

[0014] The present invention forms a multi-dimensional data set by acquiring multi-modal data such as microphone image sequence, vibration waveform and sound pressure signal. The diversity of this data set ensures that the different physical characteristics and potential defect modes in the microphone working process can be fully reflected. The zero-crossing rate of the multi-modal data set is calculated, and the zero-crossing rate signal set is obtained based on this. The zero-crossing rate is an important signal feature that can effectively capture the drastic changes and mutation points in the signal waveform, reflecting the dynamic characteristics of the signal. The nonlinear quantization processing of the zero-crossing rate signal set can further enhance the resolution and expression ability of the signal features, especially when there are uneven changes or mutations in the signal. This nonlinear quantization processing helps to improve the sensitivity of the model to tiny defects. Through this process, redundant information in the multi-modal data can be removed and important features can be compressed, thereby improving data processing efficiency. The obtained microphone multi-modal sampling quantization data is preprocessed to generate a time-space aligned multi-modal data stream. At this stage, alignment and standardization processing are performed for different types of data to ensure that data such as images, vibration waveforms and sound pressure signals are consistent in time and space. This step effectively avoids interference between different data sources due to time deviation or inconsistent sampling frequency, and lays the foundation for subsequent feature extraction and modeling. The spatiotemporal aligned data stream can provide more accurate data input for defect detection, ensure data consistency in subsequent model training and reasoning, and further improve detection accuracy and reliability.

[0015] Preferably, performing data preprocessing on microphone multimodal sampling quantized data comprises the following steps:

[0016] Perform image denoising on the microphone image sequence through a filter to generate microphone image denoising data; perform histogram equalization illumination correction on the microphone image denoising data to generate microphone illumination correction data; perform resolution scale normalization on the microphone illumination correction data to generate microphone image normalization data;

[0017] Perform vibration noise reduction processing on the microphone vibration waveform by using high-pass filtering to generate microphone vibration waveform noise reduction data; perform Fu value normalization processing on the microphone vibration waveform noise reduction data to generate microphone vibration waveform standardized data;

[0018] De-noising the microphone sound pressure signal by using wavelet transform to generate microphone de-noised pressure signal data; extracting the frequency component of the microphone de-noised pressure signal data by fast Fourier transform, and performing normalization processing to generate microphone sound pressure standardized data;

[0019] The microphone image standardized data, the microphone vibration waveform standardized data and the microphone sound pressure standardized data are windowed and timestamp aligned to generate a spatiotemporally aligned multimodal data stream.

[0020] The present invention performs denoising on the microphone image sequence through a filter, which can effectively reduce high-frequency noise and background interference in the image, making the image data more representative and helpful for subsequent image feature extraction. This step not only enhances the image quality, but also provides a clearer data basis for subsequent illumination correction and standardization processing. Illumination correction is performed through histogram equalization, which solves the problem of uneven illumination in the image acquisition process, so that the microphone image still maintains stable visual quality under different environmental conditions. This process optimizes the contrast and brightness of the image, making the defect details more prominent, and helps to improve the sensitivity of defect detection. Next, resolution scale standardization is performed to unify the image size and scale to ensure that all image data have the same resolution and processing standards, thereby providing consistent data input for subsequent feature fusion and model training. For the microphone vibration waveform, noise reduction is performed through high-pass filtering, which effectively removes the low-frequency noise component, improves the clarity and accuracy of the vibration waveform data, and reduces the interference caused by environmental noise or vibration of the device itself. Subsequently, the vibration waveform is standardized by using Fu value normalization to ensure the uniform scale of the data, which is conducive to the fusion and model training of multimodal data. For the sound pressure signal data, denoising is performed through wavelet transform, which can remove useless noise while retaining the key features of the signal and improve the signal-to-noise ratio. The frequency components are extracted through fast Fourier transform and normalized to make the frequency characteristics of the sound pressure signal more prominent and standardized, which helps to better identify the frequency patterns related to defects in subsequent steps. Finally, through window timestamp alignment, the microphone image standardized data, vibration waveform standardized data, and sound pressure standardized data are aligned in time and space, providing a unified time scale and spatial position for multimodal data fusion and defect identification.

[0021] Preferably, step S2 comprises the following steps:

[0022] Step S21: performing multimodal data coupling on the spatiotemporally aligned multimodal data stream, and outputting microphone multimodal coupling data;

[0023] Step S22: using a preset MEMS microphone structure to perform physical modeling on the microphone multi-modal coupling data to generate a MEMS microphone physical characteristic model;

[0024] Step S23: Perform physical-GAN enhancement optimization based on the MEMS microphone physical characteristic model to generate a microphone defect coupling model.

[0025] The present invention generates microphone multimodal coupling data by coupling multimodal data streams aligned in time and space. This process integrates different modal data such as images, vibration waveforms and sound pressure signals, can fully explore the correlation between multidimensional signals, and improve the comprehensive expression ability of data. Through coupling processing, data of different modes are converted into a multimodal data set with collaborative information, which helps to capture potential complex defect modes, especially when the interaction between different modal signals cannot be analyzed independently by a single mode. By using a preset MEMS microphone structure to perform physical modeling on multimodal coupling data, a MEMS microphone physical property model is generated. This modeling process can more accurately simulate the actual working state of the microphone by introducing physical structural characteristics, taking into account the design characteristics and working environment of the MEMS microphone, thereby improving the applicability and credibility of the model in practical applications. The physical property model can reflect the influence of different materials and different structural parameters on the performance of the microphone, and provide a more scientific basis for subsequent defect analysis. Physical-GAN enhanced optimization is performed based on the physical property model of the MEMS microphone, thereby generating a microphone defect coupling model. In this process, Physics-GAN is optimized through generative adversarial network technology, and the model's expressiveness and generalization ability are enhanced by combining physical property data. Physics-GAN can not only better learn complex physical relationships, but also effectively handle noise and outliers in the data, improving the model's ability to detect tiny defects. Through these three steps, the microphone defect coupling model generated in the end has higher accuracy and robustness, can still effectively identify microphone defects in complex noise environments, and provide reliable defect prediction results.

[0026] Preferably, step S21 includes the following steps:

[0027] Step S211: capturing surface physical defects of the microphone image standardization data to generate microphone surface physical defect data; performing crack pattern analysis on the microphone surface physical defect data, and performing smoothness corrosion scanning to generate microphone physical characteristic data;

[0028] Step S212: performing mechanical response waveform analysis on the standardized data of microphone vibration waveform to generate microphone mechanical response waveform data; performing spectrum feature analysis on the microphone mechanical response waveform data to generate microphone time domain spectrum feature data; performing physical-time domain feature analysis on the microphone time domain spectrum feature data and the microphone physical characteristic data to integrate them to generate microphone physical-time domain feature data;

[0029] Step S213: extracting the acoustic signal amplitude from the microphone sound pressure normalization data, and performing two-dimensional drawing to generate a two-dimensional drawing of the microphone acoustic signal; performing peak analysis on the two-dimensional drawing of the microphone acoustic signal to generate microphone acoustic characteristic data;

[0030] Step S214: performing multimodal acoustic coupling on the microphone acoustic characteristic data and the microphone physical-time domain characteristic data to generate microphone multimodal coupling data.

[0031] The present invention successfully extracts the surface defect features by capturing the surface physical defects of the microphone image standardized data, and further processes the image data through crack pattern analysis and smoothness corrosion scanning to generate microphone surface physical defect data. This processing method can carefully capture the tiny physical defects on the microphone surface, especially factors such as cracks and surface roughness, which are important factors affecting the performance of the microphone. The mechanical response waveform analysis is performed on the microphone vibration waveform standardized data, and the time domain spectrum feature data is obtained through spectrum feature analysis. This process extracts the time domain spectrum features related to the mechanical response from the vibration waveform, which is crucial for revealing the physical response characteristics of the microphone during operation. By integrating the microphone time domain spectrum feature data with the physical characteristic data through physical-time domain feature analysis, the correlation between the vibration waveform and the physical characteristics is further enhanced, providing comprehensive space-time information for subsequent defect analysis. The acoustic signal amplitude is extracted from the microphone sound pressure standardized data, and two-dimensional drawing is performed, and then the peak analysis is performed on the drawing to generate the microphone acoustic characteristic data. This processing method extracts acoustic features closely related to microphone performance, such as amplitude changes and frequency response, by analyzing the waveform changes of acoustic signals, thus providing another important feature dimension for defect detection. Finally, the microphone acoustic characteristic data is multimodally acoustically coupled with the physical-time domain feature data to generate microphone multimodal coupling data. This step integrates multidimensional information such as images, vibration waveforms, and sound pressure signals by fusing the features of different data sources, so that the data is richly expressed in multiple feature spaces, which helps to capture the multidimensional features of potential defects.

[0032] Preferably, step S23 includes the following steps:

[0033] Step S231: performing microphone vibration-acoustic response simulation on the microphone acoustic characteristic data and the microphone physical characteristic data to generate microphone vibration-acoustic response constraint items;

[0034] Step S232: using the microphone vibration-acoustic response constraint item to perform physical constraint processing on the MEMS microphone physical characteristic model to generate a microphone defect model; constructing a GAN discriminator for the microphone defect model to generate a microphone GAN discriminator;

[0035] Step S233: Perform physical-GAN enhancement optimization on the microphone defect model based on the microphone GAN discriminator to generate a microphone defect coupling model.

[0036] The present invention generates a microphone vibration-acoustic response constraint item by simulating the vibration-acoustic response of microphone acoustic characteristic data and physical characteristic data. This constraint item can accurately capture the performance changes of the microphone in actual use by simulating the behavior of the microphone under vibration and acoustic response, especially the defect performance caused by the interaction of vibration and sound waves. Through this process, the relationship between the physical characteristics and acoustic characteristics of the microphone can be closely combined, thereby providing a more realistic and comprehensive physical description, and providing an important physical constraint basis for subsequent model optimization. The generated microphone vibration-acoustic response constraint item is applied to the MEMS microphone physical characteristic model for physical constraint processing, and then a microphone defect model is generated. Through the processing of physical constraints, the microphone defect model can more accurately reflect the actual defect performance of the microphone under the interaction of vibration and acoustics, avoiding performance distortion caused by overfitting of the model. This step improves the generalization ability of the model in practical applications by introducing physical knowledge and data constraints. At the same time, the step of constructing a microphone GAN discriminator uses the discriminator of the generative adversarial network to further optimize the model, and the data generated by the model is judged by the discriminator, which can further distinguish between defective and normal data. This process helps improve the accuracy of the generated model and reduce the occurrence of misjudgments or false data when generating defect models. Finally, the microphone defect model is physically enhanced and optimized based on the microphone GAN discriminator to generate a microphone defect coupling model. Through GAN optimization and combined with physical constraints, the generated defect coupling model has been significantly improved in accuracy and robustness, and can better identify microphone defects in complex situations, especially in detection tasks under noisy and interference environments.

[0037] Preferably, step S3 comprises the following steps:

[0038] Step S31: extracting the sound-wave coupling shared encoder of the microphone defect coupling model to generate a microphone dual-domain shared encoder; performing domain-specific screening based on the microphone dual-domain shared encoder to generate a microphone private domain encoder;

[0039] Step S32: performing maximum mean distribution alignment based on the microphone dual-domain shared encoder and the microphone private-domain encoder, and constructing a feature network to generate a microphone dual-branch feature decoupling network;

[0040] Step S33: Use the microphone dual-branch feature decoupling network to build a meta-learning framework for the microphone defect coupling model, and perform image space interpolation optimization to construct an image microphone detection defect model.

[0041] The present invention extracts the features of the microphone defect coupling model through an acoustic-wave coupling shared encoder, and generates a microphone dual-domain shared encoder. The shared encoder can effectively extract the common features of the microphone in acoustic and wave response, and provides a unified representation of multimodal features for subsequent analysis. In addition, domain-specific screening is performed based on the dual-domain shared encoder, and a microphone private domain encoder is further generated. The private domain encoder can independently process features related to a specific domain, thereby refining and enhancing the unique feature expression of each domain, so that when facing complex defects or data noise, the model can more accurately capture the key features in each domain and avoid cross-domain feature interference. Step S32 performs maximum mean distribution alignment on the microphone dual-domain shared encoder and the microphone private domain encoder, and constructs a feature network to generate a microphone dual-branch feature decoupling network. Through the maximum mean distribution alignment (MMD) method, the distribution differences between different data sources or domains are further reduced, so that the model can effectively align information between multiple feature domains and eliminate the existing inconsistencies between domains. The construction of the dual-branch feature decoupling network separates domain features from common features, thereby enhancing the network's ability to independently process features from different domains and improving the specificity and robustness of the model. Step S33 uses the microphone dual-branch feature decoupling network to build a meta-learning framework for the microphone defect coupling model, and finally builds an image microphone detection defect model through image space interpolation optimization. The meta-learning framework allows the network to quickly adapt to different tasks and improves the model's migration ability when facing new data, while the image space interpolation optimization further refines the spatial features of the image data, making the image recognition of microphone defects more accurate.

[0042] Preferably, step S33 includes the following steps:

[0043] Step S331: constructing an edge-cloud collaborative architecture using a microphone dual-branch feature decoupling network to generate a microphone edge-cloud collaborative architecture;

[0044] Step S332: generating decoupled samples for the microphone defect coupling model based on the microphone edge-cloud collaborative architecture to obtain microphone feature space interpolation data;

[0045] Step S333: performing virtual defect sample space interpolation optimization on the microphone defect coupling model according to the microphone feature space interpolation data to obtain an image microphone detection defect model.

[0046] The present invention builds an edge-cloud collaborative architecture through a microphone dual-branch feature decoupling network, and generates a microphone edge-cloud collaborative architecture. The architecture realizes hierarchical management of data processing, in which the edge end is responsible for fast data preprocessing and preliminary defect analysis, while the cloud end undertakes more complex model training and optimization tasks. Through the synergy of the edge and the cloud, efficient processing of real-time data and model updating can be achieved, while avoiding the delay problem in the data transmission process, and improving the real-time performance and computing efficiency of the system. This architecture not only improves the response speed of the system, but also effectively reduces the burden on the cloud through distributed computing, so that microphone defect detection can be carried out simultaneously in multiple devices and multiple scenarios. Decoupled sample generation is performed based on the microphone edge-cloud collaborative architecture to obtain microphone feature space interpolation data. Through decoupled sample generation, the system can extract feature differences from different sources or different fields, form more representative feature space data, and then perform spatial interpolation processing on the data. This processing method can generate more diverse and comprehensive sample data, making the training set richer, thereby improving the learning ability of the model. According to the microphone feature space interpolation data, the microphone defect coupling model is optimized by virtual defect sample space interpolation, and finally an image microphone detection defect model is obtained. Through interpolation optimization of virtual defect samples, the system can generate some defect samples that are difficult to obtain in the real environment, thereby introducing more defect types during the training process and increasing the robustness and generalization ability of the model. This not only improves the detection accuracy of the model for complex defects, but also enhances the adaptability of the model in the face of different environments and noise conditions.

[0047] Preferably, step S4 comprises the following steps:

[0048] Step S41: acquiring a microphone defect image; performing microphone defect detection evaluation on an image microphone defect detection model to generate image microphone defect evaluation data;

[0049] Step S42: comparing the microphone defect image with the image microphone defect evaluation data to generate image microphone defect comparison data;

[0050] Step S43: constructing a defect detection report based on the image microphone defect comparison data to generate an image microphone defect detection report.

[0051] The present invention obtains a microphone defect image, evaluates the image in combination with an image microphone detection defect model, and generates image microphone defect evaluation data. This step can perform in-depth analysis on the image data, capture the defect features therein, and ensure the efficiency and accuracy of defect identification by comparing the existing model evaluation data. By performing a multi-dimensional evaluation on the image, it can ensure that the coverage of defect detection is wider, thereby improving the accuracy and reliability of the detection. The microphone defect image is compared with the image microphone defect evaluation data to generate image microphone defect comparison data. This process helps to verify the accuracy of defect detection and improve the system's perception of different types of defects. By comparing the image with the evaluation data, the detection model is further optimized so that it can extract more feature information from the image data to form defect comparison data with strong contrast. Through comparative analysis, it is possible to effectively identify defects that are ignored or misjudged, reduce the probability of missed judgment and misjudgment, and ensure the reliability of the system. Finally, the image microphone defect comparison data is used to construct a defect detection report to generate an image microphone defect detection report. Through this step, the detection data is converted into a standardized and structured detection report, which not only provides intuitive detection results for technicians, but also provides data support for subsequent optimization and improvement. The report can summarize key information such as the frequency and severity of different types of defects, so that decision-makers can make reasonable technical adjustments and improvements based on the report to further improve the reliability and performance of the microphone.

[0052] In this specification, a microphone defect detection system based on an image model is provided, which is used to perform the above-mentioned microphone defect detection method based on an image model. The microphone defect detection system based on an image model includes:

[0053] The data acquisition and preprocessing module is used to obtain multimodal data sets; perform data preprocessing on the multimodal data sets to generate a spatiotemporally aligned multimodal data stream;

[0054] Multi-material coupling modeling and optimization module, which is used to perform multi-material coupling modeling on the spatiotemporally aligned multimodal data stream using the preset MEMS microphone structure, and perform physical-GAN enhanced optimization to generate a microphone defect coupling model;

[0055] The feature extraction and decoupling network module is used to extract domain-invariant features of the microphone defect coupling model, construct a feature network, and generate a microphone double-branch feature decoupling network; the microphone double-branch feature decoupling network is used to build a meta-learning framework for the microphone defect coupling model, and perform image space interpolation optimization to build an image microphone detection defect model;

[0056] The defect detection and evaluation module is used to perform microphone defect detection evaluation based on the image microphone defect detection model and generate a microphone image model evaluation report.

[0057] The beneficial effect of the present invention is that by acquiring multimodal data such as images, vibration waveforms and sound pressure signals, and preprocessing these data, a time-space aligned data stream is generated. This process ensures the synchronization of data from different sources in time, and provides unified input data for subsequent model training. This measure effectively solves the problem of inaccurate time alignment between multimodal data and reduces noise interference between data. By presetting the MEMS microphone structure, multi-material coupling modeling is performed on the multimodal data stream, and the physical-generative adversarial network (GAN) is applied for optimization, and the generated microphone defect coupling model has high accuracy and reliability. Physical-GAN enhanced optimization not only improves the generalization ability of the model, but also can effectively capture the interaction between different materials, avoiding the errors and deficiencies of the single data source model. By extracting domain-invariant features and constructing feature networks for the defect coupling model, combining the dual-branch feature decoupling network with the meta-learning framework, the distinguishability and stability of the defect features are further enhanced. Image space interpolation optimization improves the quality of image data, ensures that the fusion of images and other modal data is more accurate, and finally constructs an efficient image microphone defect detection model. The detection model is used for evaluation and a microphone image model evaluation report is generated, which provides a quantitative detection effect for practical applications. This process avoids the shortcomings of defect detection models in traditional methods that are difficult to process multimodal data by integrating multiple optimization strategies, and significantly improves the accuracy of detection and the flexibility of application. Therefore, the present invention solves the deficiencies of traditional microphone defect detection methods in data processing, model accuracy and generalization ability through the combination of multimodal data stream fusion, physical optimization and deep learning, and improves the accuracy, efficiency and reliability of defect detection. BRIEF DESCRIPTION OF THE DRAWINGS

[0058] Figure 1 A schematic diagram of the steps of a microphone defect detection method based on an image model;

[0059] Figure 2 for Figure 1 Detailed implementation steps of step S2 in the flowchart;

[0060] Figure 3 for Figure 1 Detailed implementation steps of step S3 in FIG.

[0061] Figure 4 for Figure 1 Detailed implementation steps of step S4 in FIG.

[0062] The realization of the purpose, functional features and advantages of the present invention will be further explained in conjunction with embodiments and with reference to the accompanying drawings. DETAILED DESCRIPTION

[0063] The following is a clear and complete description of the technical method of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by technicians in this field without creative work are within the scope of protection of the present invention.

[0064] In addition, the accompanying drawings are only schematic illustrations of the present invention and are not necessarily drawn to scale. The same reference numerals in the figures represent the same or similar parts, and their repeated description will be omitted. Some of the block diagrams shown in the accompanying drawings are functional entities and do not necessarily correspond to physically or logically independent entities. The functional entities can be implemented in software form, or implemented in one or more hardware modules or integrated circuits, or implemented in different networks and / or processor methods and / or microcontroller methods.

[0065] It should be understood that, although the terms "first", "second", etc. may be used herein to describe various units, these units should not be limited by these terms. These terms are used only to distinguish one unit from another unit. For example, without departing from the scope of the exemplary embodiments, the first unit may be referred to as the second unit, and similarly the second unit may be referred to as the first unit. The term "and / or" used herein includes any and all combinations of one or more of the listed associated items.

[0066] To achieve this, please refer to Figures 1 to 4 , a microphone defect detection method based on an image model, the method comprising the following steps:

[0067] Step S1: obtaining a multimodal data set; performing data preprocessing on the multimodal data set to generate a spatiotemporally aligned multimodal data stream;

[0068] Step S2: Use the preset MEMS microphone structure to perform multi-material coupling modeling on the spatiotemporal aligned multimodal data stream, and perform physical-GAN enhanced optimization to generate a microphone defect coupling model;

[0069] Step S3: extract domain-invariant features of the microphone defect coupling model, and construct a feature network to generate a microphone double-branch feature decoupling network; build a meta-learning framework for the microphone double-branch feature decoupling network to the microphone defect coupling model, and perform image space interpolation optimization to construct an image microphone detection defect model;

[0070] Step S4: Perform microphone defect detection evaluation based on the image microphone defect detection model and generate a microphone image model evaluation report.

[0071] The beneficial effect of the present invention is that by acquiring multimodal data such as images, vibration waveforms and sound pressure signals, and preprocessing these data, a time-space aligned data stream is generated. This process ensures the synchronization of data from different sources in time, and provides unified input data for subsequent model training. This measure effectively solves the problem of inaccurate time alignment between multimodal data and reduces noise interference between data. By presetting the MEMS microphone structure, multi-material coupling modeling is performed on the multimodal data stream, and the physical-generative adversarial network (GAN) is applied for optimization, and the generated microphone defect coupling model has high accuracy and reliability. Physical-GAN enhanced optimization not only improves the generalization ability of the model, but also can effectively capture the interaction between different materials, avoiding the errors and deficiencies of the single data source model. By extracting domain-invariant features and constructing feature networks for the defect coupling model, combining the dual-branch feature decoupling network with the meta-learning framework, the distinguishability and stability of the defect features are further enhanced. Image space interpolation optimization improves the quality of image data, ensures that the fusion of images and other modal data is more accurate, and finally constructs an efficient image microphone defect detection model. The detection model is used for evaluation and a microphone image model evaluation report is generated, which provides a quantitative detection effect for practical applications. This process avoids the shortcomings of defect detection models in traditional methods that are difficult to process multimodal data by integrating multiple optimization strategies, and significantly improves the accuracy of detection and the flexibility of application. Therefore, the present invention solves the deficiencies of traditional microphone defect detection methods in data processing, model accuracy and generalization ability through the combination of multimodal data stream fusion, physical optimization and deep learning, and improves the accuracy, efficiency and reliability of defect detection.

[0072] In the embodiment of the present invention, reference Figure 1 FIG. 1 is a schematic diagram of a step flow of a microphone defect detection method based on an image model of the present invention. In this example, the microphone defect detection method based on an image model includes the following steps:

[0073] Step S1: obtaining a multimodal data set; performing data preprocessing on the multimodal data set to generate a spatiotemporally aligned multimodal data stream;

[0074] In the embodiment of the present invention, obtaining a multimodal data set is a key link, which involves obtaining various types of data including microphone image sequences, vibration waveforms and sound pressure signals from multiple sensors or devices. These data come from different sensor sources, and there are differences in their data structure, sampling frequency, time domain, etc., so data preprocessing is required to ensure that multimodal data can be compared and analyzed at the same time. In the data preprocessing stage, the microphone image sequence is first standardized, such as using image enhancement algorithms for denoising, illumination correction and resolution scale standardization to ensure the quality and consistency of image data. Then, the vibration waveform data is analyzed in the time domain, and high-frequency noise is removed by filtering, smoothing and other methods, and normalization is performed as needed to ensure that the data range of the vibration signal is unified, which is convenient for subsequent analysis and modeling. In addition, the preprocessing of the sound pressure signal cannot be ignored, including denoising, frequency analysis and normalization processes, and frequency domain features are extracted by fast Fourier transform (FFT) to further enhance the identifiability of the signal. The core goal of these data processing steps is to eliminate various types of noise interference, achieve consistency in data format and improve the signal-to-noise ratio of the data. Finally, after all the data has been preprocessed as above, it will be aligned in time and space through methods such as timestamp alignment to ensure that data of different modalities can be effectively registered in time and space dimensions, so that these data can be fused and analyzed in a unified format. At this point, the generated time-space aligned multimodal data stream has higher accuracy and consistency, and can provide high-quality input for subsequent modeling, analysis, and detection.

[0075] Step S2: Use the preset MEMS microphone structure to perform multi-material coupling modeling on the spatiotemporal aligned multimodal data stream, and perform physical-GAN enhanced optimization to generate a microphone defect coupling model;

[0076] In an embodiment of the present invention, a coupling model based on physical principles is established to reflect the interaction between different materials inside the microphone. Multi-material coupling modeling technology relies on physical modeling methods to combine multimodal information such as vibration waveforms, sound pressure signals and image data to ensure that each mode is mapped to each other in time and space, and to show the coupling relationship between them in the model. In the modeling process, by accurately modeling the structural characteristics of the MEMS microphone, such as the vibration membrane of the microphone, the sensor response characteristics, etc., these different physical phenomena are combined using physical simulation methods such as finite element analysis to form a comprehensive model that can reflect the dynamic performance and defect characteristics of the microphone. Next, the existing physical modeling results are further optimized by physical-GAN (generative adversarial network) enhanced optimization technology. Specifically, physical-GAN enhanced optimization introduces a generative adversarial network, so that the model can not only generate defect features according to the physical model, but also further improve the generation effect of the model through the game process between the discriminator and the generator. This process takes advantage of GAN and can automatically adjust the generated microphone defect coupling model to make it more consistent with the characteristics of the actual defect. At the data level, this technical approach improves the model's adaptability and optimization effects by making comprehensive use of physical modeling and machine learning techniques while ensuring the physical foundation.

[0077] Step S3: extract domain-invariant features of the microphone defect coupling model, and construct a feature network to generate a microphone double-branch feature decoupling network; build a meta-learning framework for the microphone double-branch feature decoupling network to the microphone defect coupling model, and perform image space interpolation optimization to construct an image microphone detection defect model;

[0078] In an embodiment of the present invention, by introducing a domain-invariant feature extraction method, the problem of data source heterogeneity can be effectively solved, ensuring that the extracted features can maintain stability and representativeness across different modes (such as images, vibration waveforms, and sound pressure signals). Specific technical means can adopt a convolutional neural network (CNN) or autoencoder structure in deep learning, and transform complex microphone defect features into low-dimensional, domain-independent representations by high-dimensional mapping of the data, so that different data modalities can obtain a unified feature space representation. Then, by constructing a feature network, the extracted invariant features are input into a neural network with a branch structure to form a so-called "microphone dual-branch feature decoupling network". The design of this network enables the model to independently process data features from different modes, and extract key information of different features such as images, vibrations, and sound pressure signals through a dual-branch network structure, thereby avoiding interference between different features and improving the recognition accuracy of the model. Furthermore, for the microphone defect coupling model, a meta-learning framework is used to build it, with the aim of enabling the model to have the ability to quickly adapt to new environments and data. Meta-learning uses a small amount of sample learning to enable the network to quickly adjust and optimize when faced with new samples or different data distributions, thereby improving the generalization ability of the model. At the same time, image space interpolation optimization technology is introduced to make detailed spatial adjustments to image data, so that changes in image resolution and scale are effectively compensated, ensuring that image data can provide more accurate feature information in different data sets or experimental environments, thereby improving the accuracy and robustness of the defect detection model.

[0079] Step S4: Perform microphone defect detection evaluation based on the image microphone defect detection model and generate a microphone image model evaluation report.

[0080] In an embodiment of the present invention, the input microphone image data is processed and passed to a previously trained detection model, which performs feature extraction and pattern recognition through a deep convolutional neural network (CNN) or other deep learning models suitable for image recognition. Specific technical means include extracting significant features related to microphone defects from the image through a pre-trained feature extraction network, and using a convolution layer to process the image at multiple scales and multiple levels, so as to capture tiny defect traces from a complex image. Then, the model further processes the extracted features through a classification layer or a regression layer to determine whether there are defects in the image and give the corresponding defect type or probability. Then, an evaluation report is generated based on the test results. The process of generating a report involves multiple layers of data processing, including but not limited to post-processing of the results, statistical analysis, and anomaly detection. Specifically, during the report generation process, the defect information output by the model will first be refined, for example, using a probability threshold to determine whether a defect exists. If the model detects a defect, the defect category will be further confirmed through relevant algorithms such as a decision tree or a support vector machine (SVM). Secondly, in order to ensure that the evaluation results of the model are scientific and reliable, the evaluation process involves the calculation of multiple evaluation indicators, such as precision, recall, F1 value, etc., which are used to measure the performance of the model, especially the accuracy and robustness in defect detection tasks.

[0081] Preferably, step S1 comprises the following steps:

[0082] Step S11: acquiring a multimodal data set, wherein the multimodal data set includes a microphone image sequence, a microphone vibration waveform, and a microphone sound pressure signal;

[0083] Step S12: calculating the zero-crossing rate of the multimodal data set to obtain a zero-crossing rate signal set; performing nonlinear quantization processing on the zero-crossing rate signal set to obtain microphone multimodal sampling quantization data;

[0084] Step S13: preprocess the microphone multimodal sampling quantized data to generate a time-space aligned multimodal data stream.

[0085] In an embodiment of the present invention, obtaining a multimodal data set refers to collecting data from different sensors, including microphone image sequences, microphone vibration waveforms, and microphone sound pressure signals. Each type of data is collected by a corresponding sensor. The microphone image sequence can record images of microphone surface or structural changes through a high-resolution camera, while the microphone vibration waveform records vibration information through an accelerometer or a vibration sensor, and the sound pressure signal records the intensity change of the sound wave through an acoustic sensor. Calculating the zero-crossing rate is a method for statistical analysis of the signal. The zero-crossing rate is the number of times the signal passes through the zero axis, which is of great significance for analyzing the frequency characteristics and noise level of the signal. This step first requires processing the original microphone vibration waveform and sound pressure signal, calculating the number of times each data point crosses the zero axis, and obtaining a zero-crossing rate signal set. Then, this signal set is subjected to nonlinear quantization processing. The processing method reduces data redundancy and quantization error by mapping continuous signals to finite discrete values, thereby obtaining microphone multimodal sampling quantization data. Nonlinear quantization can compress the signal by means such as Log quantization, compression algorithm, etc., so that the data is more representative and convenient for subsequent processing. Finally, the sampled quantized data is preprocessed, mainly including denoising, normalization, filtering, and time alignment. These processing steps ensure the consistency of different modal data in time and space, thereby generating a multimodal data stream that is aligned in time and space. The process of time and space alignment requires the use of interpolation methods to synchronize the data to ensure that each data modality presents corresponding values ​​at the same timestamp, which is crucial for the comprehensive analysis of multimodal data.

[0086] Preferably, performing data preprocessing on microphone multimodal sampling quantized data comprises the following steps:

[0087] Perform image denoising on the microphone image sequence through a filter to generate microphone image denoising data; perform histogram equalization illumination correction on the microphone image denoising data to generate microphone illumination correction data; perform resolution scale normalization on the microphone illumination correction data to generate microphone image normalization data;

[0088] Perform vibration noise reduction processing on the microphone vibration waveform by using high-pass filtering to generate microphone vibration waveform noise reduction data; perform Fu value normalization processing on the microphone vibration waveform noise reduction data to generate microphone vibration waveform standardized data;

[0089] De-noising the microphone sound pressure signal by using wavelet transform to generate microphone de-noised pressure signal data; extracting the frequency component of the microphone de-noised pressure signal data by fast Fourier transform, and performing normalization processing to generate microphone sound pressure standardized data;

[0090] The microphone image standardized data, the microphone vibration waveform standardized data and the microphone sound pressure standardized data are windowed and timestamp aligned to generate a spatiotemporally aligned multimodal data stream.

[0091] In an embodiment of the present invention, a filter is used to perform image denoising on a microphone image sequence, the purpose of which is to remove noise signals in the image and make the image clearer. Commonly used image denoising methods include Gaussian filtering, mean filtering, etc. These methods smooth the image by blurring and reduce unnecessary interference. The denoised image data is then subjected to illumination correction by histogram equalization technology to adjust the brightness and contrast of the image to eliminate differences under different illumination conditions. This processing can make the details of the image more obvious under different illumination conditions by stretching or compressing the grayscale value distribution of the image. Next, the corrected image is subjected to resolution scale standardization, aiming to unify the size and resolution of the image for subsequent analysis and processing. This step involves methods such as image scaling or cropping, so that all images have the same spatial resolution. For the vibration waveform data, a high-pass filter is first used to perform vibration denoising processing on it, the purpose of which is to remove low-frequency noise and retain high-frequency vibration signals, which is crucial to the effectiveness and accuracy of the signal. Afterwards, the vibration waveform data is normalized by using Fu value normalization processing so that it is within the same numerical range, which is convenient for subsequent comparison and processing. For the sound pressure signal of the microphone, wavelet transform is used for denoising. Wavelet transform can effectively decompose the signal into different frequency bands and suppress the noise, thereby retaining the key information of the signal. The denoised sound pressure signal data is then extracted through fast Fourier transform (FFT) to extract the frequency components, convert the time domain signal into a frequency domain signal, and further analyze the frequency characteristics of the signal. Finally, all data (including images, vibration waveforms, and sound pressure signals) are normalized to ensure that different modal data are in the same numerical range. In order to ensure the temporal consistency of different modal data, all standardized data need to be aligned with window timestamps, which can be processed by interpolation methods so that each data modality has a corresponding value at the same time point, thereby generating a multimodal data stream aligned in time and space.

[0092] As an example of the present invention, refer to Figure 2 As shown, in this example, step S2 includes:

[0093] Step S21: performing multimodal data coupling on the spatiotemporally aligned multimodal data stream, and outputting microphone multimodal coupling data;

[0094] Step S22: using a preset MEMS microphone structure to perform physical modeling on the microphone multi-modal coupling data to generate a MEMS microphone physical characteristic model;

[0095] Step S23: Perform physical-GAN enhancement optimization based on the MEMS microphone physical characteristic model to generate a microphone defect coupling model.

[0096] In an embodiment of the present invention, the time-space aligned multimodal data stream is subjected to multimodal data coupling processing. The core of this step is to convert data of different modes (such as images, vibration waveforms and sound pressure signals, etc.) into a unified data set by fusing them. This process includes data alignment, feature matching and modal cross-reference to achieve information complementarity and collaborative expression between different data sources. Through the coupling operation, data of different modes can be associated in a common feature space, so as to better reveal the correlation and potential laws hidden between different modes, and provide rich feature information for subsequent analysis. The multimodal coupling data is physically modeled using a preset MEMS (micro-electromechanical system) microphone structure. This step combines the physical properties of the sensor with the measurement data to establish a physical model to describe the working principle and response characteristics of the MEMS microphone. The purpose of physical modeling is to construct a mathematical model that can accurately reflect its working state and behavior by considering the mechanical, electrical, acoustic and other characteristics of the microphone. The model combines factors such as the size, material properties, structural characteristics and response characteristics of the microphone, and establishes a model related to sensor performance through physical laws (such as laws in the fields of mechanics, electromagnetism and acoustics). Based on the physical property model of the MEMS microphone, the physical-GAN enhancement optimization technology is used to generate a microphone defect coupling model. Here, physical-GAN optimization combines physical modeling with generative adversarial networks, and enhances and optimizes the model through the generator and discriminator of the GAN framework. The generator is responsible for generating the microphone defect model, while the discriminator compares the difference between the generated defect model and the real data and performs error feedback to further optimize the accuracy and robustness of the model.

[0097] Preferably, step S21 includes the following steps:

[0098] Step S211: capturing surface physical defects of the microphone image standardization data to generate microphone surface physical defect data; performing crack pattern analysis on the microphone surface physical defect data, and performing smoothness corrosion scanning to generate microphone physical characteristic data;

[0099] Step S212: performing mechanical response waveform analysis on the standardized data of microphone vibration waveform to generate microphone mechanical response waveform data; performing spectrum feature analysis on the microphone mechanical response waveform data to generate microphone time domain spectrum feature data; performing physical-time domain feature analysis on the microphone time domain spectrum feature data and the microphone physical characteristic data to integrate them to generate microphone physical-time domain feature data;

[0100] Step S213: extracting the acoustic signal amplitude from the microphone sound pressure normalization data, and performing two-dimensional drawing to generate a two-dimensional drawing of the microphone acoustic signal; performing peak analysis on the two-dimensional drawing of the microphone acoustic signal to generate microphone acoustic characteristic data;

[0101] Step S214: performing multimodal acoustic coupling on the microphone acoustic characteristic data and the microphone physical-time domain characteristic data to generate microphone multimodal coupling data.

[0102] In an embodiment of the present invention, the microphone image standardized data is captured by surface physical defects and used to generate microphone surface physical defect data. This step mainly analyzes the microphone surface through image processing algorithms (such as edge detection, feature point extraction, etc.) to capture existing physical defects such as cracks and flaws. Through crack pattern analysis, the morphological characteristics of the defects can be deeply revealed, and then the image can be further processed through smoothness corrosion scanning to highlight the defect area and remove noise, thereby generating more accurate microphone physical characteristic data. The microphone vibration waveform standardized data is subjected to mechanical response waveform analysis, mainly relying on time domain signal processing methods such as autocorrelation analysis or wavelet transform to extract the vibration characteristics of the microphone and generate mechanical response waveform data. Then, through spectral feature analysis, frequency domain features can be extracted from the mechanical response waveform, for example, the spectral data of the microphone is obtained through fast Fourier transform (FFT). Finally, the time domain spectral feature data of the microphone is integrated with the physical characteristic data through physical-time domain feature analysis, and different types of data features can be integrated through data fusion methods (such as multidimensional interpolation, weighted average, etc.) to form richer microphone physical-time domain feature data. The standardized data of microphone sound pressure is extracted through acoustic signal amplitude. First, the sound pressure signal captured by the microphone is analyzed to extract its amplitude characteristics, and then it is visualized as an image through two-dimensional drawing. Then, the key peaks in the sound pressure signal are identified through peak analysis technology to extract feature data related to acoustic characteristics. This process helps to quantify the spatiotemporal changes of acoustic signals and reveal the acoustic performance of the microphone. The acoustic characteristic data of the microphone is multimodally acoustically coupled with the physical-time domain feature data. Through this process, combining acoustic characteristics, physical characteristics and time domain characteristics, multiple data sources can be integrated into a unified feature space, and data fusion technology (such as multimodal learning) can be used to generate multimodal coupling data of the microphone.

[0103] Preferably, step S23 includes the following steps:

[0104] Step S231: performing microphone vibration-acoustic response simulation on the microphone acoustic characteristic data and the microphone physical characteristic data to generate microphone vibration-acoustic response constraint items;

[0105] Step S232: using the microphone vibration-acoustic response constraint item to perform physical constraint processing on the MEMS microphone physical characteristic model to generate a microphone defect model; constructing a GAN discriminator for the microphone defect model to generate a microphone GAN discriminator;

[0106] Step S233: Perform physical-GAN enhancement optimization on the microphone defect model based on the microphone GAN discriminator to generate a microphone defect coupling model.

[0107] In an embodiment of the present invention, the acoustic characteristic data and physical characteristic data of the microphone are subjected to vibration-acoustic response simulation. This process establishes a multi-dimensional simulation model by combining the vibration response and acoustic characteristics of the microphone. The model uses physical models and acoustic data to predict the behavior of the microphone in an actual working environment, and then generates microphone vibration-acoustic response constraints. These constraints are a mathematical expression of the performance limit or performance of the microphone, which is used in subsequent physical modeling and optimization processes. Based on the microphone vibration-acoustic response constraints, the physical characteristic model of the MEMS microphone is physically constrained. This step uses physical constraint theory to constrain the correlation between the physical characteristics of the microphone and the vibration-acoustic response through mathematical optimization methods (such as the Lagrange multiplier method) to ensure that the physical characteristic model of the microphone is more in line with the working state in actual operation. In addition, a GAN discriminator is constructed for the microphone defect model. In this process, the defect model is trained by the discriminator of the generative adversarial network (GAN), and the discriminator improves the accuracy of the defect model by comparing the difference between the generated defect data and the real defect data. Based on the GAN discriminator, the microphone defect model is enhanced and optimized by physical-GAN to further improve the performance of the defect model. This process combines the physical model with the GAN method, so that the actual physical laws can be incorporated into the data-driven model optimization, thereby enhancing the model's expressiveness while ensuring that the model output results are consistent with the actual physical phenomena. Through this physical-GAN optimization, the generated microphone defect coupling model not only has powerful feature extraction capabilities, but also can further improve the accuracy and robustness of defect detection through the optimization of the generative adversarial network.

[0108] As an example of the present invention, refer to Figure 3 As shown, in this example, step S3 includes:

[0109] Step S31: extracting the sound-wave coupling shared encoder of the microphone defect coupling model to generate a microphone dual-domain shared encoder; performing domain-specific screening based on the microphone dual-domain shared encoder to generate a microphone private domain encoder;

[0110] Step S32: performing maximum mean distribution alignment based on the microphone dual-domain shared encoder and the microphone private-domain encoder, and constructing a feature network to generate a microphone dual-branch feature decoupling network;

[0111] Step S33: Use the microphone dual-branch feature decoupling network to build a meta-learning framework for the microphone defect coupling model, and perform image space interpolation optimization to construct an image microphone detection defect model.

[0112] In an embodiment of the present invention, the microphone defect coupling model is feature extracted by the sound-wave coupling shared encoder to generate a microphone dual-domain shared encoder. This step uses a shared encoder to extract and share features between multiple data domains, thereby realizing the fusion of multimodal data, especially the coupling between the acoustic and vibration characteristics of the microphone. The core of the shared encoder is that it can simultaneously process and extract information from microphone acoustic signals and vibration signals, thereby generating a unified feature space, and further extracting features with stronger recognition capabilities under specific tasks through domain-specific screening to generate a microphone private domain encoder. Based on the microphone dual-domain shared encoder and the private domain encoder, alignment is performed by the maximum mean distribution alignment (MMD) technology to reduce the distribution differences between different data domains. This technology minimizes the mean difference between the source domain and the target domain by comparing the statistical characteristics of the source domain and the target domain, ensuring that the model can effectively transfer learning and generalization. The purpose of MMD alignment is to eliminate the deviation between the source domain and the target domain in data distribution and improve the robustness of the model in an unknown environment. Through the aligned feature data, a feature network is further constructed to generate a microphone dual-branch feature decoupling network. The network has two branches, which process the features of acoustic and vibration data respectively, helping the model to learn the feature representations of different data sources more accurately. The meta-learning framework is built using the microphone dual-branch feature decoupling network. Meta-learning is a technology that enables the model to quickly adapt to new tasks by learning how to learn. In this step, the adaptability and accuracy of the microphone defect detection model are further improved through the meta-learning method combined with the image space interpolation optimization technology. Image space interpolation optimization can enhance the local features in the image by spatially interpolating the image data, so that the model can more accurately identify and locate defects when detecting defects with image microphones.

[0113] Preferably, step S33 includes the following steps:

[0114] Step S331: constructing an edge-cloud collaborative architecture using a microphone dual-branch feature decoupling network to generate a microphone edge-cloud collaborative architecture;

[0115] Step S332: generating decoupled samples for the microphone defect coupling model based on the microphone edge-cloud collaborative architecture to obtain microphone feature space interpolation data;

[0116] Step S333: performing virtual defect sample space interpolation optimization on the microphone defect coupling model according to the microphone feature space interpolation data to obtain an image microphone detection defect model.

[0117] In an embodiment of the present invention, a microphone dual-branch feature decoupling network is used to build an edge-cloud collaborative architecture. The construction of this architecture involves distributing the microphone data processing process to edge computing devices and cloud servers to achieve distributed processing of computing tasks. The edge computing device is responsible for real-time data processing and local feature extraction, while the cloud concentrates on processing large-scale data, training and optimizing deep learning models. Through this collaborative architecture, the response speed and computing efficiency of the system can be effectively improved, especially in large-scale microphone data processing and defect detection tasks, where edge computing provides fast local processing capabilities, while the cloud handles more complex global tasks. Decoupled sample generation is performed based on the microphone edge-cloud collaborative architecture. This step generates microphone feature space interpolation data by implementing a decoupling operation in the collaborative architecture. The goal of decoupled sample generation is to extract and integrate microphone data features, eliminate redundant information between different modal data, and thereby obtain more representative samples. Feature space interpolation further optimizes these samples so that the generated data can not only accurately capture the characteristics of the microphone, but also effectively fill in the gaps in the feature space, enhancing the diversity and robustness of the data. The generated microphone feature space interpolation data is used to perform virtual defect sample space interpolation optimization on the microphone defect coupling model. This technical means aims to simulate different types of microphone defects and enhance the model's ability to identify defect types by generating and optimizing virtual defect samples in the data space. Through interpolation optimization, more representative and diverse defect samples can be created. These samples not only increase the diversity of training data, but also improve the model's generalization ability when facing unknown defects. The core technology of this process is spatial interpolation based on the generative model, which can generate virtual defect samples through computers, further improving the accuracy and reliability of the model in image microphone detection.

[0118] As an example of the present invention, refer to Figure 4 As shown, in this example, step S4 includes:

[0119] Step S41: acquiring a microphone defect image; performing microphone defect detection evaluation on an image microphone defect detection model to generate image microphone defect evaluation data;

[0120] Step S42: comparing the microphone defect image with the image microphone defect evaluation data to generate image microphone defect comparison data;

[0121] Step S43: constructing a defect detection report based on the image microphone defect comparison data to generate an image microphone defect detection report.

[0122] In an embodiment of the present invention, microphone defect images are obtained as input data for subsequent analysis. These images include different types of defects or anomalies, and these defect images are obtained by sensors or imaging systems. On this basis, the existing image microphone detection defect model is used to perform defect detection evaluation on the acquired image data. By applying deep learning models, especially image classification and detection algorithms such as convolutional neural networks (CNNs), microphone defects in images are identified and image microphone defect evaluation data is generated. The evaluation data contains the detection results, such as information such as the category, location and severity of the defects. The microphone defect image is compared and analyzed with the image microphone defect evaluation data. The main purpose of this step is to verify and refine the detection results to ensure the accuracy of the model prediction results. By comparing the image and the evaluation data one by one, the system can further correct the model's erroneous predictions and generate image microphone defect comparison data through the comparison results. This data set provides the difference between different model predictions and actual defect locations, which is of great significance for model optimization and error adjustment. Based on the generated image microphone defect comparison data, a defect detection report is constructed. By integrating the evaluation data and the comparison data, the key information in the defect detection process can be systematically displayed, including the defect type, location, detection reliability and other contents, and finally a complete defect detection report is generated.

[0123] In this specification, a microphone defect detection system based on an image model is provided, which is used to perform the above-mentioned microphone defect detection method based on an image model. The microphone defect detection system based on an image model includes:

[0124] The data acquisition and preprocessing module is used to obtain multimodal data sets; perform data preprocessing on the multimodal data sets to generate a spatiotemporally aligned multimodal data stream;

[0125] Multi-material coupling modeling and optimization module, which is used to perform multi-material coupling modeling on the spatiotemporally aligned multimodal data stream using the preset MEMS microphone structure, and perform physical-GAN enhanced optimization to generate a microphone defect coupling model;

[0126] The feature extraction and decoupling network module is used to extract domain-invariant features of the microphone defect coupling model, construct a feature network, and generate a microphone double-branch feature decoupling network; the microphone double-branch feature decoupling network is used to build a meta-learning framework for the microphone defect coupling model, and perform image space interpolation optimization to build an image microphone detection defect model;

[0127] The defect detection and evaluation module is used to perform microphone defect detection evaluation based on the image microphone defect detection model and generate a microphone image model evaluation report.

[0128] The beneficial effect of the present invention is that by acquiring multimodal data such as images, vibration waveforms and sound pressure signals, and preprocessing these data, a time-space aligned data stream is generated. This process ensures the synchronization of data from different sources in time, and provides unified input data for subsequent model training. This measure effectively solves the problem of inaccurate time alignment between multimodal data and reduces noise interference between data. By presetting the MEMS microphone structure, multi-material coupling modeling is performed on the multimodal data stream, and the physical-generative adversarial network (GAN) is applied for optimization, and the generated microphone defect coupling model has high accuracy and reliability. Physical-GAN enhanced optimization not only improves the generalization ability of the model, but also can effectively capture the interaction between different materials, avoiding the errors and deficiencies of the single data source model. By extracting domain-invariant features and constructing feature networks for the defect coupling model, combining the dual-branch feature decoupling network with the meta-learning framework, the distinguishability and stability of the defect features are further enhanced. Image space interpolation optimization improves the quality of image data, ensures that the fusion of images and other modal data is more accurate, and finally constructs an efficient image microphone defect detection model. The detection model is used for evaluation and a microphone image model evaluation report is generated, which provides a quantitative detection effect for practical applications. This process avoids the shortcomings of defect detection models in traditional methods that are difficult to process multimodal data by integrating multiple optimization strategies, and significantly improves the accuracy of detection and the flexibility of application. Therefore, the present invention solves the deficiencies of traditional microphone defect detection methods in data processing, model accuracy and generalization ability through the combination of multimodal data stream fusion, physical optimization and deep learning, and improves the accuracy, efficiency and reliability of defect detection.

[0129] Therefore, the embodiments should be regarded as illustrative and non-restrictive from all points, and the scope of the present invention is limited by the appended claims rather than the above description, and it is therefore intended that all changes falling within the meaning and range of equivalent elements of the application documents are included in the present invention.

[0130] The above description is only a specific embodiment of the present invention, so that those skilled in the art can understand or implement the present invention. Various modifications to these embodiments will be apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to the embodiments shown herein, but should conform to the widest scope consistent with the principles and novel features invented herein.

Claims

1. A microphone defect detection method based on an image model, characterized in that: The following steps are involved: Step S1: Obtain a multimodal dataset; Perform data preprocessing on multimodal datasets to generate spatiotemporally aligned multimodal data streams; Step S2: using a preset MEMS microphone structure to perform multi-material coupling modeling on the spatiotemporally aligned multimodal data stream to generate a MEMS microphone physical property model; performing physical-GAN enhancement optimization on the MEMS microphone physical property model to generate a microphone defect coupling model; Step S3: extracting domain-invariant features of the microphone defect coupling model and constructing a feature network to generate a microphone double-branch feature decoupling network; A meta-learning framework is built for the microphone defect coupling model using the microphone dual-branch feature decoupling network, and image space interpolation optimization is performed to construct an image microphone detection defect model. Step S4: Perform microphone defect detection evaluation based on the image microphone defect detection model to generate microphone image model evaluation data; construct a report on the microphone image model evaluation data to generate a microphone image model evaluation report.

2. The microphone defect detection method based on image model according to claim 1, characterized in that: Step S1 includes the following steps: Step S11: acquiring a multimodal data set, wherein the multimodal data set includes a microphone image sequence, a microphone vibration waveform, and a microphone sound pressure signal; Step S12: calculating the zero-crossing rate of the multimodal data set to obtain a zero-crossing rate signal set; performing nonlinear quantization processing on the zero-crossing rate signal set to obtain microphone multimodal sampling quantization data; Step S13: preprocess the microphone multimodal sampling quantized data to generate a time-space aligned multimodal data stream.

3. The microphone defect detection method based on image model according to claim 2, characterized in that: Data preprocessing of microphone multimodal sampling quantized data includes the following steps: Perform image denoising on the microphone image sequence through a filter to generate microphone image denoising data; perform histogram equalization illumination correction on the microphone image denoising data to generate microphone illumination correction data; perform resolution scale normalization on the microphone illumination correction data to generate microphone image normalization data; Perform vibration noise reduction processing on the microphone vibration waveform by using high-pass filtering to generate microphone vibration waveform noise reduction data; perform Fu value normalization processing on the microphone vibration waveform noise reduction data to generate microphone vibration waveform standardized data; De-noising the microphone sound pressure signal by using wavelet transform to generate microphone de-noised pressure signal data; extracting the frequency component of the microphone de-noised pressure signal data by fast Fourier transform, and performing normalization processing to generate microphone sound pressure standardized data; The microphone image standardized data, the microphone vibration waveform standardized data and the microphone sound pressure standardized data are windowed and timestamp aligned to generate a spatiotemporally aligned multimodal data stream.

4. The microphone defect detection method based on image model according to claim 1, characterized in that: Step S2 includes the following steps: Step S21: performing multimodal data coupling on the spatiotemporally aligned multimodal data stream, and outputting microphone multimodal coupling data; Step S22: using a preset MEMS microphone structure to perform physical modeling on the microphone multi-modal coupling data to generate a MEMS microphone physical characteristic model; Step S23: Perform physical-GAN enhancement optimization based on the MEMS microphone physical characteristic model to generate a microphone defect coupling model.

5. The microphone defect detection method based on image model according to claim 4, characterized in that: Step S21 includes the following steps: Step S211: capturing surface physical defects of the microphone image standardization data to generate microphone surface physical defect data; performing crack pattern analysis on the microphone surface physical defect data, and performing smoothness corrosion scanning to generate microphone physical characteristic data; Step S212: performing mechanical response waveform analysis on the standardized data of microphone vibration waveform to generate microphone mechanical response waveform data; performing spectrum feature analysis on the microphone mechanical response waveform data to generate microphone time domain spectrum feature data; performing physical-time domain feature analysis on the microphone time domain spectrum feature data and the microphone physical characteristic data to integrate them to generate microphone physical-time domain feature data; Step S213: extracting the acoustic signal amplitude from the microphone sound pressure normalization data, and performing two-dimensional drawing to generate a two-dimensional drawing of the microphone acoustic signal; performing peak analysis on the two-dimensional drawing of the microphone acoustic signal to generate microphone acoustic characteristic data; Step S214: performing multimodal acoustic coupling on the microphone acoustic characteristic data and the microphone physical-time domain characteristic data to generate microphone multimodal coupling data.

6. The microphone defect detection method based on image model according to claim 4, characterized in that: Step S23 includes the following steps: Step S231: performing microphone vibration-acoustic response simulation on the microphone acoustic characteristic data and the microphone physical characteristic data to generate microphone vibration-acoustic response constraint items; Step S232: using the microphone vibration-acoustic response constraint item to perform physical constraint processing on the MEMS microphone physical characteristic model to generate a microphone defect model; constructing a GAN discriminator for the microphone defect model to generate a microphone GAN discriminator; Step S233: Perform physical-GAN enhancement optimization on the microphone defect model based on the microphone GAN discriminator to generate a microphone defect coupling model.

7. The microphone defect detection method based on image model according to claim 1, characterized in that: Step S3 includes the following steps: Step S31: extracting the sound-wave coupling shared encoder of the microphone defect coupling model to generate a microphone dual-domain shared encoder; performing domain-specific screening based on the microphone dual-domain shared encoder to generate a microphone private domain encoder; Step S32: performing maximum mean distribution alignment based on the microphone dual-domain shared encoder and the microphone private-domain encoder, and constructing a feature network to generate a microphone dual-branch feature decoupling network; Step S33: Use the microphone dual-branch feature decoupling network to build a meta-learning framework for the microphone defect coupling model, and perform image space interpolation optimization to construct an image microphone detection defect model.

8. The microphone defect detection method based on image model according to claim 7, characterized in that: Step S33 includes the following steps: Step S331: constructing an edge-cloud collaborative architecture using a microphone dual-branch feature decoupling network to generate a microphone edge-cloud collaborative architecture; Step S332: generating decoupled samples for the microphone defect coupling model based on the microphone edge-cloud collaborative architecture to obtain microphone feature space interpolation data; Step S333: performing virtual defect sample space interpolation optimization on the microphone defect coupling model according to the microphone feature space interpolation data to obtain an image microphone detection defect model.

9. The microphone defect detection method based on image model according to claim 1, characterized in that: Step S4 includes the following steps: Step S41: acquiring a microphone defect image; performing microphone defect detection evaluation on an image microphone defect detection model to generate image microphone defect evaluation data; Step S42: comparing the microphone defect image with the image microphone defect evaluation data to generate image microphone defect comparison data; Step S43: constructing a defect detection report based on the image microphone defect comparison data to generate an image microphone defect detection report.

10. A microphone defect detection system based on an image model, characterized in that: For executing the microphone defect detection method based on an image model as claimed in claim 1, the microphone defect detection system based on an image model comprises: The data acquisition and preprocessing module is used to obtain multimodal data sets; perform data preprocessing on the multimodal data sets to generate a spatiotemporally aligned multimodal data stream; Multi-material coupling modeling and optimization module, which is used to perform multi-material coupling modeling on the spatiotemporally aligned multimodal data stream using the preset MEMS microphone structure, and perform physical-GAN enhanced optimization to generate a microphone defect coupling model; The feature extraction and decoupling network module is used to extract domain-invariant features of the microphone defect coupling model, construct a feature network, and generate a microphone double-branch feature decoupling network; the microphone double-branch feature decoupling network is used to build a meta-learning framework for the microphone defect coupling model, and perform image space interpolation optimization to build an image microphone detection defect model; The defect detection and evaluation module is used to perform microphone defect detection evaluation based on the image microphone defect detection model and generate a microphone image model evaluation report.

Citation Information

Cited By

  • Industrial big data analysis method based on intelligent manufacturing

    CN120612329A