Sound equipment manufacturing abnormity intelligent detection method and system

By building an acoustic twin model and a hybrid neural network, and combining acoustic, image, and electrical data for defect identification, the problems of missed detection and information isolation in traditional acoustic testing are solved, and accurate and efficient quality control and production optimization are achieved in the acoustic manufacturing process.

CN120746400APending Publication Date: 2025-10-03CHANGSHA HOTONE AUDIO

Patent Information

Application Number
CN202511189264.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-25
Publication Date
2025-10-03

AI Technical Summary

Technical Problem

Traditional audio detection methods have a high missed detection rate and slow speed, and cannot be matched with automated production control systems. In addition, the information in the detection link is isolated and cannot be comprehensively analyzed and optimized.

Method used

Acoustic data is acquired through a microphone array, sensing equipment detects internal structures and circuits, an acoustic twin model is constructed, and a hybrid neural network is combined for defect identification and linkage analysis, providing real-time feedback on detection results and adjusting production processes.

Benefits of technology

It achieves precise and efficient quality control, improves production efficiency, reduces defect rates, and realizes continuous optimization and intelligent management of the production process through data feedback.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120746400A_ABST
    Figure CN120746400A_ABST
Patent Text Reader

Abstract

The invention discloses a sound equipment manufacturing abnormity intelligent detection method and system, and relates to the technical field of sound equipment detection, and the method comprises the steps: building a twin model of a sound equipment based on feature parameters, dividing the twin model into a plurality of sub-models according to the demands of each link, building a mixed neural network model based on the twin model, and carrying out the detection of the sound equipment manufacturing abnormity. The method comprises the following steps: capturing long-range dependence, identifying sound appearance and structure anomalies, combining and inputting multi-source characteristic parameters into a hybrid neural network model, carrying out defect identification and linkage analysis, carrying out anomaly diagnosis on a detection result after completing multi-link data linkage analysis, identifying various anomaly types existing in the sound, and if anomalies are identified, judging whether the sound is abnormal or not. And a detection result is fed back to a production control end in real time, unqualified sound boxes are automatically screened out, and production process parameters are automatically adjusted according to detected abnormal types. According to the detection system, the production efficiency is effectively improved, the defect rate is reduced, and continuous optimization and intelligent management of the production process are realized through data feedback.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of audio detection, and in particular to an intelligent detection method and system for audio manufacturing anomalies. Background Art

[0002] With the rapid development of technologies such as artificial intelligence, machine learning, and computer vision, anomaly detection in audio manufacturing has gradually become intelligent. Intelligent detection systems can automatically identify potential defects and anomalies by analyzing various data of audio products during the production process, and detect substandard products in advance. They are widely used in the production process of various audio products, such as home theater audio, Bluetooth audio, headphones, car audio, etc., helping manufacturers improve product quality, reduce production costs, and enhance market competitiveness.

[0003] The existing technology has the following defects: 1. Traditional audio testing relies on manual testing or single-sensor testing, which presents a series of problems. First, the missed detection rate is high, especially for minor defects such as diaphragm cracks and magnetic gap offset. These problems are often difficult to detect with the naked eye, resulting in substandard products being missed. Second, traditional testing relies on empirical judgment, is slow, and cannot match the high-speed pace of automated production control systems, thus causing production bottlenecks and affecting overall production efficiency and product quality. 2. Existing audio testing methods still have data silos. The data generated by multiple links such as acoustic testing, appearance testing, and circuit testing of audio often do not form an effective linkage analysis, resulting in isolated information in each testing link and the inability to conduct comprehensive analysis and optimization. This situation not only makes it impossible to accurately discover potential problems in the product, but also limits the comprehensive control and improvement of the production process.

[0004] Based on this, the present invention proposes an intelligent detection method and system for audio manufacturing anomalies, which can provide more accurate, efficient and real-time quality control in the audio manufacturing process, effectively improve production efficiency, reduce defect rates, and realize continuous optimization and intelligent management of the production process through data feedback. Summary of the Invention

[0005] The purpose of the present invention is to provide a method and system for intelligent detection of audio manufacturing anomalies to address the shortcomings of the background technology.

[0006] In order to achieve the above object, the present invention provides the following technical solution: an intelligent detection method for abnormalities in audio manufacturing, the detection method comprising the following steps: Acquire acoustic data through a microphone array, and use sensor equipment to detect the internal structure and circuit of the speaker to obtain internal image data and electrical data; Perform feature extraction on each type of data to refine feature parameters related to defect detection; Build a twin model of the sound system based on characteristic parameters, and divide the twin model into multiple sub-models according to the needs of each link; A hybrid neural network model is constructed based on the twin model to capture long-range dependencies and identify abnormalities in sound appearance and structure; Combine multi-source feature parameters and input them into a hybrid neural network model to perform defect identification and linkage analysis; After completing the multi-link data linkage analysis, the test results are diagnosed for abnormalities to identify various types of abnormalities in the audio system; If an abnormality is identified, the detection results will be fed back to the production control end in real time, unqualified audio will be automatically screened out, and the production process parameters will be automatically adjusted according to the type of abnormality detected.

[0007] In a preferred embodiment, a hybrid neural network model is constructed based on the twin model to capture long-range dependencies and identify abnormalities in sound appearance and structure, including the following steps: Hybrid neural network models include Transformer and CNN; After Transformer and CNN process time series data and image data respectively, they fuse the time series data and image data and conduct joint analysis: Small sample learning technology is introduced to increase the amount of training data for the hybrid neural network model.

[0008] In a preferred embodiment, the time series data and image data are fused and jointly analyzed, including the following steps: The feature vectors from Transformer and CNN are concatenated to form a feature vector. The fused feature vector is used as input and sent to the fully connected layer for classification or regression analysis.

[0009] In a preferred embodiment, the Transformer data processing step includes: Learn the correlation between various time points in the data by calculating the weighted sum of each time step in the input sequence; Add positional encoding to the input data and sense the order of the input data; Learn information from data from multiple perspectives through multi-head attention; The CNN includes convolutional layers, pooling layers and fully connected layers: Convolutional layer: The input image is processed layer by layer through multiple convolutional layers. Each convolution operation extracts low-level features of the image. As the network layer increases, the features extracted by the convolutional layer become increasingly complex. Pooling layer: used to reduce the size of the feature map. Pooling operations include maximum pooling and average pooling. Fully connected layer: After convolution and pooling, the feature map is flattened and input into the fully connected layer for the final classification or regression task. The fully connected layer learns the nonlinear relationship of the feature map.

[0010] In a preferred embodiment, performing abnormality diagnosis on the detection results to identify various abnormal types present in the sound includes the following steps: By analyzing the frequency response curve and distortion acoustic characteristics of the speakers, and comparing the actual collected frequency response data with the expected standard values, frequency response distortion, excessive harmonic distortion, or other sound quality issues can be identified. When the acoustic performance is abnormal, the abnormal location is marked within the time window. The edge detection algorithm extracts details from the surface image of the speaker, identifies scratches or cracks, compares the image with the standard appearance, and detects deviations in the speaker assembly process. Once defects are detected, the location of the defect is marked on the image. By analyzing the waveform of the electrical signal, abnormal fluctuations in current or voltage can be identified, sudden changes in voltage and current can be detected to identify short circuit and overload faults, the time point when the electrical fault occurs can be calibrated, and the cause of the fault can be provided.

[0011] In a preferred embodiment, the detection results are subjected to abnormality diagnosis to identify various types of abnormalities present in the audio system, further comprising the following steps: after completing various types of abnormality detection, the acoustic abnormality, appearance defect, and electrical fault are marked separately, and the location of the defect is recorded; Defect type marking: Each defect is classified according to the preset standards and its severity is marked; Location positioning: For each detected defect, a location tag is provided and a defect report is provided based on these tags.

[0012] In a preferred embodiment, if an abnormality is identified, the detection results are fed back to the production control end in real time, unqualified audio is automatically screened out, and production process parameters are automatically adjusted according to the type of abnormality detected, including the following steps: After the anomaly is identified and located, the detection results are transmitted to the production control system; The sorting system of the production control system automatically screens out unqualified products and removes them from the production control system based on the type and location of the detected anomalies; Automatically adjust production process parameters based on the type of abnormality detected.

[0013] In a preferred embodiment, a twin model of the sound system is constructed based on characteristic parameters, and the twin model is divided into multiple sub-models according to the requirements of each link, including the following steps: The acoustic sub-model of the speaker is constructed using frequency response curves and distortion acoustic data. The appearance quality sub-model includes surface defects and assembly deviations of the speaker. The electrical performance sub-model reflects the health status of the electrical system through current and voltage signals. Each sub-model is trained independently, and the feature outputs of different sub-models are fused to form a comprehensive feature vector, which serves as the input of the twin model; After the twin model is completed, a digital image is generated for each audio product for comparison with the actual product. The digital image includes the acoustic performance, appearance quality and electrical performance of the audio.

[0014] In a preferred embodiment, feature extraction is performed on each type of data to refine feature parameters related to defect detection, including the following steps: De-noising, data cleaning and standardization of the collected raw data; For acoustic data, the time domain signal is converted into the frequency domain signal through Fourier transform, and the response characteristics of the audio equipment at different frequencies are extracted; For image data, edge detection algorithm is used to extract edge information of the acoustic surface; For electrical data, the time series signals of voltage and current are analyzed to extract the signal change patterns.

[0015] The present application also provides an intelligent detection system for audio manufacturing anomalies, including a multi-source feature extraction module, a model building module, and an anomaly identification and automatic adjustment module; Multi-source feature extraction module: This module acquires acoustic data through a microphone array and uses sensor equipment to detect the internal structure and circuits of the speaker, obtaining internal image data and electrical data. It then extracts features from each type of data and refines feature parameters relevant to defect detection. Model building module: This module builds a twin model of the sound system based on characteristic parameters and divides the twin model into multiple sub-models according to the requirements of each link. A hybrid neural network model is then constructed based on the twin model to capture long-range dependencies and identify abnormalities in the sound system's appearance and structure. Abnormal identification and automatic adjustment module: Combines multi-source feature parameters and inputs them into the hybrid neural network model to perform defect identification and linkage analysis, diagnoses abnormalities in the test results, and identifies various abnormal types in the sound. If an abnormality is identified, the test results are fed back to the production control end in real time, automatically screening out unqualified sound, and automatically adjusting the production process parameters according to the detected abnormality type.

[0016] In the above technical solution, the technical effects and advantages provided by the present invention are: The present invention constructs a twin model of an acoustic system based on characteristic parameters. This model is then divided into multiple sub-models based on the requirements of each stage. A hybrid neural network model is then constructed based on the twin model to capture long-range dependencies and identify abnormalities in the acoustic system's appearance and structure. Multi-source characteristic parameters are then combined and input into the hybrid neural network model for defect identification and linkage analysis. After completing the multi-stage linkage analysis, the detection results are analyzed for abnormality, identifying various types of abnormalities in the acoustic system. If an abnormality is identified, the detection results are fed back to the production control end in real time, automatically screening out unqualified acoustic systems and adjusting production process parameters based on the detected abnormality type. This detection system can provide more accurate, efficient, and real-time quality control during the acoustic system manufacturing process, effectively improving production efficiency and reducing defect rates. Furthermore, data feedback enables continuous optimization and intelligent management of the production process. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments described in the present invention. For ordinary technicians in this field, other drawings can also be obtained based on these drawings.

[0018] Figure 1 Flowchart of the present invention.

[0019] Figure 2 This is a diagram of the architecture of the present invention. DETAILED DESCRIPTION

[0020] To make the objectives, technical solutions, and advantages of the embodiments of the present invention more clear, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.

[0021] Example: See Figure 1 As shown, the present embodiment provides an intelligent detection method for abnormalities in audio manufacturing, the detection method comprising the following steps: Comprehensive data collection is performed using a variety of high-precision equipment. A high-precision microphone array captures acoustic data such as the speaker's frequency response curve and distortion, ensuring accurate reflection of the speaker's acoustic performance. Industrial cameras and infrared sensors also conduct a comprehensive inspection of the speaker's exterior, capturing surface defects such as scratches and assembly deviations, and enabling in-depth detection of internal structural anomalies (such as loose coils). Furthermore, current and voltage sensors monitor the operating status of the speaker's circuit modules, capturing electrical data and comprehensively recording electrical performance.

[0022] Collected data often contains noise and incomplete information, necessitating data preprocessing. This stage primarily involves noise removal, data cleaning, and standardization to ensure data accuracy and consistency. Furthermore, feature extraction is performed on each type of data (acoustic, image, and electrical signals) to refine characteristic parameters critical to defect detection, laying a solid foundation for subsequent analysis and modeling.

[0023] A twin model of the speaker is constructed based on characteristic parameters. This model integrates different dimensions of the speaker (acoustic performance, appearance quality, and electrical performance) to form a digital "mirror image" of the speaker. Based on the needs of each link, the model is divided into multiple sub-models, each focusing on a specific area (such as acoustics, appearance, or circuitry), providing accurate baseline data for comprehensive analysis and anomaly detection.

[0024] Based on the acoustic twin model, a hybrid neural network model based on Transformer and CNN was constructed, combining acoustic, image, and electrical signals for feature fusion and joint analysis. The Transformer processes time series data to capture long-range dependencies, while the CNN extracts spatial features from image data to identify anomalies in the appearance and structure of the acoustics. Small-sample learning technology was introduced to address the issue of insufficient initial defect samples, ensuring that the model can accurately detect anomalies even with small data volumes.

[0025] Once data collection and model training are complete, multi-step data linkage analysis is performed. This process combines acoustic, image, and electrical data for comprehensive defect identification and linkage analysis. For example, an acoustic performance anomaly may be closely related to a cosmetic defect or circuit failure. Through this linkage analysis, the system can accurately identify and locate the root cause of the problem, thereby improving the accuracy of defect detection.

[0026] After analyzing multi-step data linkage, the system conducts a comprehensive anomaly diagnosis of the test results. Specifically, the system identifies various anomalies within the audio product, including acoustic performance disturbances (such as frequency response curve distortion), cosmetic defects (such as scratches and assembly deviations), and electrical faults (such as current / voltage anomalies). The type and location of each defect are precisely labeled, providing a basis for subsequent production adjustments.

[0027] If an anomaly is detected, the system immediately transmits the test results to the production control center in real time, automatically triggering the sorting system within the production control system to ensure that unqualified products are promptly removed. Furthermore, the system automatically adjusts production process parameters, such as glue dosage and assembly pressure, based on the type of anomaly detected to prevent similar issues from recurring. This dynamic feedback mechanism not only improves production efficiency but also ensures consistent and stable product quality.

[0028] This application constructs a twin model of the audio system based on characteristic parameters, and divides the twin model into multiple sub-models according to the needs of each link. A hybrid neural network model is constructed based on the twin model to capture long-range dependencies and identify abnormalities in the appearance and structure of the audio system. Multi-source characteristic parameters are combined and input into the hybrid neural network model to perform defect identification and linkage analysis. After completing the multi-link data linkage analysis, the test results are diagnosed for abnormalities and the various types of abnormalities in the audio system are identified. If an abnormality is identified, the test results are fed back to the production control end in real time, automatically screening out unqualified audio systems, and automatically adjusting the production process parameters based on the detected abnormality type. This detection system can provide more accurate, efficient, and real-time quality control in the audio manufacturing process, effectively improving production efficiency, reducing defect rates, and achieving continuous optimization and intelligent management of the production process through data feedback.

[0029] See also Figure 2 As shown, the intelligent detection system for abnormalities in audio manufacturing described in this embodiment includes a multi-source feature extraction module, a model building module, and an abnormality recognition and automatic adjustment module; Multi-source feature extraction module: This module acquires acoustic data through a microphone array and uses sensor equipment to detect the internal structure and circuits of the speaker, obtaining internal image data and electrical data. It then extracts features from each type of data, refining feature parameters related to defect detection and sending these feature parameters to the model building module. Model building module: This module builds a twin model of the sound system based on characteristic parameters and divides the twin model into multiple sub-models according to the requirements of each link. A hybrid neural network model is then constructed based on the twin model to capture long-range dependencies and identify abnormalities in the sound system's appearance and structure. The hybrid neural network model is then sent to the abnormality identification and automatic adjustment module. Abnormal identification and automatic adjustment module: Combines multi-source feature parameters and inputs them into the hybrid neural network model to perform defect identification and linkage analysis, diagnoses abnormalities in the test results, and identifies various abnormal types in the sound. If an abnormality is identified, the test results are fed back to the production control end in real time, automatically screening out unqualified sound, and automatically adjusting the production process parameters according to the detected abnormality type.

[0030] Comprehensive data collection is performed using a variety of high-precision equipment. A high-precision microphone array captures acoustic data such as the speaker's frequency response curve and distortion, ensuring accurate reflection of the speaker's acoustic performance. Industrial cameras and infrared sensors also conduct a comprehensive inspection of the speaker's exterior, capturing surface defects such as scratches and assembly deviations, and enabling in-depth detection of internal structural anomalies (such as loose coils). Furthermore, current and voltage sensors monitor the operating status of the speaker's circuit modules, capturing electrical data and comprehensively recording electrical performance.

[0031] In the intelligent detection of audio manufacturing anomalies, comprehensive data collection using a variety of high-precision equipment is crucial. First, a high-precision microphone array is used to acquire acoustic data, such as the speaker's frequency response curve and distortion. The microphone array configuration captures the speaker's response across different frequency ranges, ensuring accurate reflection of the speaker's acoustic performance. This process utilizes advanced signal processing technology. By capturing high-resolution sound waves, the frequency response of the speaker's output can be precisely measured and compared to standard performance indicators, revealing whether the speaker exhibits frequency response distortion or excessive distortion.

[0032] In terms of visual inspection, the combined use of industrial cameras and infrared sensors enables comprehensive inspection of the speaker's exterior. Industrial cameras capture images of the speaker's exterior, detecting surface defects such as scratches, cracks, and assembly deviations. Infrared sensors, through infrared radiation detection, penetrate deep into the speaker's interior, identifying potential problems such as loose coils and abnormal solder joints. By combining these two devices, the system can acquire high-resolution image data in real time and, combined with image processing algorithms, automatically identify and classify defects.

[0033] Furthermore, electrical performance testing is also a key component of this technical solution. Current and voltage sensors are used to monitor the operating status of the audio circuit module. By recording electrical data in real time, a comprehensive assessment of the audio system's electrical performance can be made. These sensors can accurately monitor current and voltage changes, promptly detecting potential circuit anomalies such as overcurrent, overvoltage, or power fluctuations. This helps determine whether the electrical module is functioning properly and whether there are potential fault hazards.

[0034] Collected data often contains noise and incomplete information, necessitating data preprocessing. This stage primarily involves noise removal, data cleaning, and standardization to ensure data accuracy and consistency. Furthermore, feature extraction is performed on each type of data (acoustic, image, and electrical signals) to refine characteristic parameters critical to defect detection, laying a solid foundation for subsequent analysis and modeling.

[0035] During the data preprocessing phase, the collected raw data must first be denoised, cleaned, and standardized. This process is crucial because raw data often contains external environmental noise, sensor errors, and other unavoidable interference, which can affect the accuracy of subsequent analysis. Therefore, a series of technical methods must be employed to optimize the data.

[0036] Denoising: It is the first step in data preprocessing. Commonly used denoising methods include signal filtering, smoothing, and frequency domain denoising techniques. In acoustic data, commonly used denoising methods are low-pass filters and Kalman filters. Low-pass filters remove noise above a certain set frequency and retain low-frequency signals, which can effectively filter out high-frequency interference components. In Kalman filtering, the system predicts the signal at each moment and compares it with the actual sampling value, continuously adjusting the estimated value to reduce the impact of noise and improve signal accuracy. Denoising in image data usually uses median filtering or Gaussian filtering, which removes noise points and retains edge features by smoothing the local area of ​​the image.

[0037] Data cleaning: It is a key step to remove outliers and missing values ​​in the data. At this stage, the system needs to identify and process abnormal data points. These anomalies may be caused by sensor failure, external interference or operational errors. For example, in audio electrical data, voltage fluctuations caused by short circuits or poor contact may occur, and such fluctuations will cause abnormal electrical signals. To this end, the cleaning process uses rule detection methods to identify these outliers and repair them based on the trends of other data points. For example, if the voltage data fluctuates violently, the system will replace the outlier with the mean of the surrounding data to avoid interference with subsequent model training. For missing values, common processing methods include mean filling, linear interpolation or model-based prediction filling. These methods can infer missing values ​​based on known data, thereby ensuring data integrity.

[0038] Standardization: The purpose is to convert data of different dimensions and scales into a unified standard scale so that subsequent feature extraction and model training can proceed smoothly. A common method of standardization is to convert the data into a standard normal distribution with a mean of 0 and a variance of 1, which helps to eliminate the dimensional differences between different data features. In acoustic data, due to the large differences in the intensity of sound waves at different frequencies, standardization can effectively handle these differences so that the responses of all frequencies can be compared at the same scale. Image data may have different resolutions and lighting conditions. Standardization adjusts the pixel values ​​so that the color distribution and contrast of each image are at the same level, avoiding the impact of scale differences between images on the detection results.

[0039] After data preprocessing, the feature extraction phase begins. The goal of feature extraction is to extract important information related to defect detection from the processed data and provide useful input for subsequent analysis and modeling. Feature extraction methods are tailored to the data type (acoustic, image, or electrical).

[0040] Acoustic Data Feature Extraction: For acoustic data, the goal of feature extraction is to analyze key signal characteristics such as frequency response curves and distortion. Common feature extraction methods include: Spectral analysis: This technique converts time-domain signals into frequency-domain signals through Fourier transforms, extracting the response characteristics of audio equipment at different frequencies. The spectrum can be used to identify whether the audio equipment exhibits frequency response distortion.

[0041] Time domain analysis: Evaluate the dynamic performance of the audio, such as transient response and noise, by calculating the signal's time domain characteristics such as mean, variance, and peak value.

[0042] Image data feature extraction: In feature extraction of image data, the goal is to extract features related to appearance defects, especially surface defects and internal structural anomalies, through image processing techniques.

[0043] Edge detection: Using edge detection algorithms such as the Sobel operator and the Canny operator to extract edge information on the surface of the audio system can help detect surface defects such as scratches and cracks.

[0044] Texture analysis: By extracting the texture features of the audio surface (such as uniformity, roughness, etc.), we can identify whether there are assembly deviations or surface defects.

[0045] Deep learning feature extraction: Use convolutional neural networks (CNNs) to automatically extract complex features from images, which is important for detecting more complex appearance defects or internal structural abnormalities.

[0046] Electrical data feature extraction: For electrical data, feature extraction focuses on analyzing the changing trends in signals such as current and voltage.

[0047] Timing analysis: Analyze the timing signals of voltage and current to extract characteristics such as signal change patterns and fluctuation amplitudes to help detect electrical faults.

[0048] Outlier detection: By setting a reasonable threshold range, it can extract abnormal fluctuation information in electrical signals and quickly locate possible fault points in the circuit.

[0049] After feature extraction, all extracted features are converted into vectors or matrices to facilitate subsequent machine learning model training and analysis. The feature vectors at this point contain key information for each data type, and these features will serve as the core input for defect detection model training and inference.

[0050] A twin model of the speaker is constructed based on characteristic parameters. This model integrates different dimensions of the speaker (acoustic performance, appearance quality, and electrical performance) to form a digital "mirror image" of the speaker. Based on the needs of each link, the model is divided into multiple sub-models, each focusing on a specific area (such as acoustics, appearance, or circuitry), providing accurate baseline data for comprehensive analysis and anomaly detection.

[0051] When building a twin model of a speaker based on characteristic parameters, the core goal is to integrate and model various dimensions of the speaker (including acoustic performance, cosmetic quality, and electrical performance) through digital technology, thereby generating a highly accurate "mirror image" of the speaker. This twin model not only reflects the various performance indicators of the speaker during the production process but also provides precise baseline data for subsequent defect detection and performance evaluation.

[0052] After data integration is complete, the twin model will be used to generate multi-dimensional models based on the speaker's distinct characteristics. Essentially, the twin model is a highly abstract digital representation that integrates the three dimensions of a speaker: acoustic performance, exterior quality, and electrical performance.

[0053] Acoustic Performance Modeling: First, an acoustic sub-model of the speaker is constructed using acoustic data such as frequency response curves and distortion. Acoustic data processing includes methods such as spectral analysis and time-domain analysis. By extracting key signal features (such as frequency response, distortion, and noise level), the speaker's acoustic performance is accurately characterized. In the twin model, these features serve as input to the acoustic sub-model, describing the speaker's basic sound output characteristics.

[0054] Appearance Quality Modeling: Secondly, the appearance quality sub-model focuses on the visual characteristics of the speaker, including surface defects (such as scratches and cracks) and assembly deviations. Image data captured by industrial cameras and infrared sensors can be used to extract key surface and internal features of the speaker using image processing algorithms (such as edge detection and texture analysis). In the twin model, these visual features are modeled using computer vision techniques to ensure an accurate reflection of the speaker's appearance quality.

[0055] Electrical Performance Modeling: Finally, the electrical performance submodel focuses on the operating status of the audio circuits, primarily reflecting the health of the electrical system through current and voltage signals. Time series signal analysis methods, such as waveform analysis and outlier detection, are used to extract key features from the electrical data (such as voltage fluctuations and power variations). Within the twin model, this data helps characterize the operating status of the audio electrical system.

[0056] Based on the twin model, the system fuses the sub-models to achieve comprehensive analysis and anomaly detection across different dimensions. Each sub-model focuses on a specific aspect of the sound, processing the corresponding data type through a specific algorithm. When performing multi-dimensional integration, a weighted fusion strategy is often used, in which the output of each sub-model is weighted according to its importance in the overall model. Sub-model fusion can be performed through the following steps: Sub-model training: Each sub-model is trained independently, using different algorithms for specific data types (such as acoustic data, image data, or electrical data). For example, the acoustic sub-model might be trained using time-frequency analysis combined with machine learning models (such as support vector machines or deep learning models). The image sub-model might use convolutional neural networks (CNNs) for image feature extraction and classification. The electrical sub-model might employ time series data analysis methods, such as LSTMs (long short-term memory networks), to capture temporal dependencies in electrical signals.

[0057] Feature fusion: The feature outputs of different sub-models are fused, typically using feature-level fusion techniques. The feature vectors of acoustic, image, and electrical data are combined according to specific rules (such as concatenation and weighted averaging) to form a comprehensive feature vector, which serves as the input to the final twin model.

[0058] Model Optimization: By optimizing the twin model as a whole, we further enhance the synergy between the sub-models. During the optimization process, we use gradient descent or other optimization algorithms to iteratively optimize the fused feature vectors, ensuring that each sub-model makes accurate predictions within its domain while improving the detection accuracy of the overall model.

[0059] After the twin model is complete, a digital "mirror image" is generated for each audio product for comparison with the actual product. This mirror image includes all baseline data for acoustic performance, appearance quality, and electrical performance, serving as a reference for proper product operation. By comparing this data with real-time data, the twin model can accurately identify anomalies in the audio product and detect defects.

[0060] When an anomaly is detected, the twin model can pinpoint the specific problem, such as acoustic distortion, cosmetic scratches, or electrical faults, and identify the type and location of the problem. Based on this anomaly information, the system provides real-time data feedback to the production side, automatically adjusting the production process or performing product screening to ensure that substandard products are promptly eliminated.

[0061] One advantage of the twin model is that it can be updated and iterated in real time. As data accumulates during the production process, the twin model dynamically adjusts and optimizes based on this new data. For example, as more defect samples are added, the model's detection accuracy continues to improve, enabling it to identify more complex anomalies. To maintain the model's efficiency and accuracy, the system regularly retrains the model to ensure it remains current with the latest production conditions and product requirements.

[0062] Through this series of steps, the twin model of the audio system can play an important role in production, testing, and quality control, providing accurate benchmark data and real-time abnormal feedback, thereby ensuring the high quality and performance of audio products.

[0063] Based on the acoustic twin model, a hybrid neural network model based on Transformer and CNN was constructed, combining acoustic, image, and electrical signals for feature fusion and joint analysis. The Transformer processes time series data to capture long-range dependencies, while the CNN extracts spatial features from image data to identify anomalies in the appearance and structure of the acoustics. Small-sample learning technology was introduced to address the issue of insufficient initial defect samples, ensuring that the model can accurately detect anomalies even with small data volumes.

[0064] Building on the acoustic twin model, a hybrid neural network model based on Transformers and CNNs was constructed to enable feature fusion and joint analysis of multimodal data (acoustics, images, and electrical signals). The core goal of this model is to use deep learning techniques to comprehensively process information from different data sources to improve the accuracy and efficiency of anomaly detection. The following are the key steps in the construction process: First, data from various sensors (such as high-precision microphones, industrial cameras, infrared sensors, current and voltage sensors, etc.) needs to be preprocessed. These data types, including acoustic signals, image data, and electrical signals, correspond to different feature spaces and therefore require targeted processing.

[0065] Acoustic Data: We perform time-frequency analysis on acoustic data, such as using the Short-Time Fourier Transform (STFT) to convert time-domain signals into frequency-domain signals, thereby extracting key information such as frequency response and distortion. These features serve as input to the Transformer model, capturing the performance fluctuations of the speaker at different frequencies.

[0066] Image data: After normalization and data augmentation, the image data is fed into a convolutional neural network (CNN) for feature extraction. CNNs are effective at extracting spatial features from images, such as scratches, cracks, and assembly deviations on the speaker's exterior.

[0067] Electrical signal data: Current and voltage time series data often exhibit long-range dependencies, and therefore require time series analysis. Simple smoothing algorithms or wavelet transforms are typically used to denoise the electrical signals to remove high-frequency noise. After these processes, the data is converted into feature vectors or matrices suitable for neural network model input.

[0068] The data form of the electrical signal (for model input): Assume the sampling frequency is , window length T point, number of channels C, (such as voltage Va / Nb / Nc, current la / lb / lc, etc.); Original sequence (single channel / multi-channel): , N is the number of sampling points; Sliding window segments (commonly used for training), , If there is a label (such as fault type), it can be formed .

[0069] Frequency domain / time-frequency domain tensor: Frequency domain: Perform FFT on each channel .

[0070] Time-frequency domain: STFT / CWT (Can be viewed as a "multi-channel spectrum") Feature matrix (statistics + spectrum): extract the mean, RMS, Kurtosis, harmonic amplitude, etc. for each window to obtain .

[0071] Three-phase synchronization stacking and phasor / sequence components: 、 As 6-channel input, or first do symmetrical component transformation to get .

[0072] The moving average method is used for denoising, and the expression is: , is the window radius (unit: point), is the smoothing window length, is the original signal, is the smoothed signal.

[0073] Transformer networks are particularly adept at capturing long-range dependencies and global features when processing time series data. For acoustic and electrical signal data, Transformer networks use a self-attention mechanism to capture long-term dependencies in the data. This allows them to accurately learn the correlations between different time points, especially when the time series is complex.

[0074] In this multimodal signal processing scheme, both electrical and acoustic signals are input as time series data into the Transformer network for feature extraction. However, before entering the Transformer, each requires independent preprocessing. For the electrical signal, denoising (such as wavelet transform or filtering) is first used to remove high-frequency noise. Subsequently, the signal is windowed and normalized to form fixed-length time slices. For the acoustic signal, pre-emphasis, frame windowing, and time-frequency transformation (such as STFT or Mel spectrum extraction) are performed to obtain a standardized time series feature matrix. After preprocessing, the two types of time series data are typically fed into separate Transformer encoders to learn their internal long-range dependencies and global time series features. This split processing approach avoids modeling interference caused by significant differences in the distribution of features between different modalities. The resulting electrical and acoustic signal feature vectors are concatenated at the Transformer output to form a joint time series feature representation. If the electrical signal and the acoustic signal are strictly aligned on the sampling time axis, they can also be directly spliced ​​in the feature dimension after preprocessing and then input into the same Transformer network for cross-modal temporal dependency modeling.

[0075] In image modality processing, the raw image signal is normalized and resized before being fed into a convolutional neural network (CNN), such as ResNet or EfficientNet, to extract spatial structural features. The global image feature vector output by the CNN is then fused with the temporal feature vector obtained through Transformer processing. This fusion is done at the feature level, concatenating the vectors to form a comprehensive feature representation encompassing the global temporal pattern of the electrical signal, the acoustic feature pattern, and the spatial pattern of the image. Finally, this fused feature vector is fed into a multi-layer perceptron (MLP) or other classification / regression head for task prediction, enabling joint modeling and decision-making of multimodal information.

[0076] In the model, Transformer processes data through the following steps: Self-attention mechanism: This mechanism learns the correlations between different time points in the data by calculating the weighted sum of each time step in the input sequence. For example, in an acoustic signal, the response of a certain frequency may be correlated with changes in other frequencies. The self-attention mechanism can effectively capture these dependencies.

[0077] Positional encoding: Because the Transformer lacks the spatial structure of traditional convolutional neural networks, positional encoding is added to the input data to enable the model to perceive the order of the input data. For time series data, positional encoding helps the Transformer understand the order of the signal in the time dimension.

[0078] Multi-Head Attention Mechanism: Through multi-head attention, the Transformer can learn information from data from multiple perspectives (i.e., multiple attention heads), thereby improving the expressiveness of feature representation. This is particularly important for acoustic and electrical signals, as they may exhibit different change patterns at different time scales.

[0079] CNNs are primarily used for processing image data. In this model, they are responsible for extracting spatial features from the appearance and structure of speakers. Appearance defects such as scratches and cracks can be automatically learned and extracted through the convolutional layers of CNNs. The specific steps include: Convolutional layers: The input image is processed layer by layer through multiple convolutional layers. Each convolution operation extracts low-level features of the image (such as edges, textures, etc.). As the network layers increase, the features extracted by the convolutional layers become increasingly complex and abstract.

[0080] Pooling layer: The pooling layer is used to reduce the size of the feature map, reduce the amount of computation, and enhance the robustness of the model. Commonly used pooling operations include maximum pooling and average pooling, which can effectively preserve the key features in the image.

[0081] Fully connected layer: After convolution and pooling, the feature map is flattened and input into the fully connected layer for the final classification or regression task. Through the fully connected layer, CNN can learn complex nonlinear relationships, thereby improving the accuracy of anomaly detection.

[0082] After the Transformer and CNN process the time series data and image data respectively, these features need to be fused for joint analysis. This process requires the design of an appropriate feature fusion strategy, which can usually be adopted in the following ways: Feature concatenation: The feature vectors from the Transformer and CNN are concatenated to form a large feature vector. These fused features are then fed into the subsequent fully connected layers for classification or regression analysis.

[0083] Weighted-Fusion: This approach weights features from different sources, assigning different weights to features. This approach dynamically adjusts the weight of each data source based on its importance in the detection task, thereby improving model performance.

[0084] Multimodal learning: By designing a joint training framework, the Transformer and CNN share information in the same network for end-to-end learning. This approach enables the model to extract more comprehensive features from multimodal data and improve the final detection accuracy.

[0085] Since there may be a shortage of defective samples in the early stages of production, small sample learning technology is introduced to address the problem of insufficient data. Small sample learning improves the generalization ability of the model through the following methods: Transfer learning: This technique leverages model weights trained on a large dataset to fine-tune a small sample of data, thereby improving model training efficiency and accuracy. For example, a pre-trained Transformer model can be fine-tuned on a smaller dataset of acoustic defects to achieve even better performance.

[0086] Data augmentation: Generate diverse defect samples through data augmentation techniques. Image data can be enhanced through rotation, scaling, and cropping to generate more training data. For acoustic and electrical data, data diversity can be enhanced by adding noise and time delays.

[0087] Model regularization: Regularization techniques (such as Dropout and L2 regularization) are used to prevent overfitting. In particular, regularization can help improve the generalization ability of the model in small sample learning scenarios.

[0088] The trained hybrid neural network model can detect potential anomalies in real time when new data is input. To do this, the model first jointly analyzes the acoustic, visual, and electrical characteristics of the speaker and classifies it based on the fused feature vector. If an anomaly is detected, the system generates an anomaly report and feeds the results back to production control, triggering automated production adjustments or rejection of substandard products.

[0089] Through these steps, the hybrid neural network based on Transformer and CNN can effectively perform multimodal anomaly detection in audio, and can still ensure high-precision detection results even when the sample data is small.

[0090] Once data collection and hybrid neural network model training are complete, multi-step data linkage analysis is performed. This process combines acoustic, image, and electrical data for comprehensive defect identification and linkage analysis. For example, an abnormal acoustic performance may be closely related to a cosmetic defect or circuit failure. Through this linkage analysis of data, the system can accurately identify and locate the root cause of the problem, thereby improving the accuracy of defect detection.

[0091] The hybrid neural network model is not only responsible for extracting high-dimensional features of acoustic, image, and electrical data respectively, but also directly provides unified feature representation and task prediction results for multi-link data linkage analysis at its output. Specifically, after being processed by network structures such as Transformer and CNN, the data of the three modalities are mapped to the same feature space and spliced ​​into a comprehensive feature vector at the fusion layer. This vector contains both the temporal pattern of the acoustic signal, the spatial texture information of the image signal, and the dynamic feature changes of the electrical signal. In the prediction stage of the model, the network not only outputs the final defect category or fault type, but also outputs the contribution weight or attention distribution of each modal feature in the prediction. These output indicators provide an interpretable basis for the subsequent multi-link data linkage analysis.

[0092] During the multi-link data linkage analysis process, the system uses the comprehensive feature vectors and prediction results output by the neural network model to trace back the performance of each modality in anomaly detection. For example, if the model assigns a higher weight to the acoustic feature in a certain prediction, and the electrical feature also shows significant anomalies, it can be inferred that the defect may be caused by mechanical or structural problems caused by abnormal circuit operation; on the contrary, if the image feature weight is prominent and the acoustic signal is normal, it may be more inclined to be a pure appearance defect. This feature contribution analysis based on model output makes the linkage analysis no longer rely on manual experience judgment, but is based on the cross-modal features extracted by the unified model, thereby realizing a closed-loop process from feature extraction, anomaly detection to multi-link correlation diagnosis. In this way, the output of the neural network model is not only the result of defect detection, but also the core driving data source of the linkage analysis.

[0093] Once data collection is complete and the hybrid neural network model is trained, the next step is multi-step data analysis. The core goal of this phase is to integrate and analyze acoustic, image, and electrical data for comprehensive defect identification and problem location. Through multi-dimensional data integration, the system can more comprehensively understand the performance of the audio system, accurately identify and locate the root causes of potential problems, and thus improve the accuracy and efficiency of defect detection. The specific process is as follows: The first step in multi-segment data linkage analysis is to fuse data from different sensors. Acoustic, image, and electrical data typically exist in different dimensions and have a certain degree of time skew, so alignment and synchronization are necessary to ensure that the data can be effectively combined.

[0094] Time Series Data Alignment: Acoustic and electrical data are typically time series data with temporal dependencies. First, align these data using timestamps or interpolation methods to ensure that data at each time step matches. For example, if the acoustic and electrical data have different sampling frequencies, interpolation methods are needed to align them to the same time step.

[0095] Image Data Alignment: Although image data is not time-series data, when performing multi-step analysis, we typically select a time window that corresponds to the acoustic and electrical data to ensure that the appearance of the sound is captured within the same time period. By using feature extraction and time tag binding in image processing, we ensure that each image is aligned with the corresponding acoustic and electrical data.

[0096] By analyzing data from multiple links, the system can more accurately pinpoint the root cause of a problem. For example, at a specific moment, the acoustic data's frequency response curve is distorted, the electrical data exhibits abnormal fluctuations, and the image data also shows obvious scratches on the exterior of the speaker. By analyzing this information together, the system can infer that the distortion may be related to current fluctuations caused by a circuit fault and the scratches on the exterior, thereby more accurately locating the problem.

[0097] To further improve diagnostic accuracy, graph analysis or causal reasoning models can be used to construct functional dependency graphs for various audio components and analyze the causal relationships between different data sources. For example, this allows the system to determine whether an electrical fault has caused a degradation in acoustic performance, or whether a cosmetic defect has affected the operating status of the electrical system. In this way, the system can trace the propagation path of the anomaly and identify its true root cause.

[0098] Once the problem is located, the system will generate real-time feedback based on the analysis results and trigger corresponding production adjustments. Specifically, the system will provide feedback in the following ways: Real-time production feedback: When the system detects an anomaly, it automatically generates an alert and pushes it to the production control system. The system will locate the problem to the specific production link based on the type of defect (such as acoustic distortion, cosmetic defects, or electrical failure) and guide operators to make necessary interventions or adjustments.

[0099] Automated process adjustments: By analyzing the root causes of defects, the system can automatically adjust the production process. For example, if electrical data reveals an overcurrent problem, the system will adjust the current parameters. If an assembly problem is found to have caused a cosmetic defect, the system will automatically adjust the pressure or speed of the assembly equipment to prevent similar problems from occurring again.

[0100] Through the coordinated analysis of these links, the system can achieve more accurate defect identification and root cause location based on multi-dimensional data fusion, thereby significantly improving production efficiency and product quality.

[0101] After analyzing multi-step data linkage, the system conducts a comprehensive anomaly diagnosis of the test results. Specifically, the system identifies various anomalies within the audio product, including acoustic performance disturbances (such as frequency response curve distortion), cosmetic defects (such as scratches and assembly deviations), and electrical faults (such as current / voltage anomalies). The type and location of each defect are precisely labeled, providing a basis for subsequent production adjustments.

[0102] After completing the multi-step data analysis, the system enters the comprehensive anomaly diagnosis phase. In this phase, the system comprehensively analyzes acoustic, visual, and electrical data to identify various anomalies within the audio product. Specifically, the system identifies and locates potential defects based on the analysis results for each data type (acoustic performance, cosmetic defects, and electrical faults). Each defect type is not only accurately identified but also precisely marked at its location, providing precise guidance for subsequent production adjustments and optimization.

[0103] Acoustic performance anomalies are primarily detected by analyzing acoustic characteristics such as the speaker's frequency response curve and distortion. The system compares the actual frequency response data collected with the expected standard values ​​to identify frequency response distortion, excessive harmonic distortion, or other sound quality issues.

[0104] Frequency Response Curve Distortion Detection: This function identifies distortion by calculating the deviation between the frequency response curve and a standard curve. For example, the system can set a tolerance range; if a frequency band on the frequency response curve falls outside this range, it will be marked as abnormal. The system algorithmically calculates the deviation value for each frequency point, and if the deviation exceeds the set threshold, it is marked as abnormal.

[0105] Distortion Detection: This function evaluates the distortion of audio by calculating the total harmonic distortion (THD) at different frequencies. When the distortion exceeds a set threshold, the system identifies it as abnormal. For example, if the distortion exceeds a certain threshold (e.g., 5%), the system triggers an abnormality alarm.

[0106] Abnormality Location: When acoustic performance is abnormal, the system precisely marks the location of the abnormality within a specific frequency range or time window. For example, if frequency response distortion is concentrated in a certain frequency band, the system will mark the problem area within that band and provide the specific location, helping the production control system quickly locate and repair the problem.

[0107] Appearance defect detection is performed using image data. After being processed by a CNN (convolutional neural network) model, the image data can identify defects such as scratches, cracks, and assembly deviations on the speaker's appearance. The specific detection process includes: Scratch and Crack Detection: Using edge detection algorithms (such as Canny edge detection and the Sobel operator), the system can extract details from the audio surface image and identify the presence of scratches or cracks. These algorithms accurately locate defective areas by calculating changes in pixel values ​​within the image.

[0108] Assembly deviation detection: Assembly deviation can be detected by comparing the image to a preset model or standard appearance, thereby detecting deviations during the speaker assembly process. Common algorithms include shape matching and template matching. By comparing the image to a standard template, the system can accurately locate the deviation.

[0109] Defect Marking and Location: After detecting a defect, the system marks its specific location on the image and calculates its size, type, and severity. For example, for a scratch, the system marks its start and end points, as well as its width. For assembly deviations, the system provides the deviation amount and precisely calibrates its location.

[0110] Electrical fault detection relies primarily on time-series data collected by current and voltage sensors. Electrical faults typically manifest as abnormal current fluctuations, overcurrent, overvoltage, or power loss. The system effectively detects these faults by analyzing the time-series of electrical signals.

[0111] Current / voltage fluctuation detection: By analyzing the waveform of electrical signals, the system can identify abnormal current or voltage fluctuations. For example, the system can detect overcurrent or voltage instability by calculating the mean and standard deviation of the electrical signal. If the fluctuation amplitude exceeds the set range, it will be marked as abnormal.

[0112] Short circuit and overload faults: Short circuit and overload faults often cause sudden changes in voltage and current. The system can identify these faults by detecting sudden changes in voltage and current. The system algorithmically detects areas of rapid change in the signal and uses set thresholds to determine whether a short circuit or overload has occurred.

[0113] Fault Location and Reporting: The system accurately pinpoints the time an electrical fault occurs and provides the cause of the fault. For example, if the system detects abnormal current fluctuations and a cosmetic defect occurring simultaneously, this may indicate a connection between the fault and an assembly issue. The system will indicate this relationship in the report and conduct a correlation analysis.

[0114] After detecting various anomalies, the system accurately labels each defect based on its type and location. The system labels acoustic anomalies, cosmetic defects, and electrical faults separately, and records the specific location of the defect.

[0115] Defect Type Labeling: Each defect (such as frequency response distortion, scratches, and overcurrent) is categorized according to pre-set criteria and labeled with its severity. For example, frequency response distortion might be labeled "mild distortion" or "severe distortion," while cosmetic defects might be labeled "scratches" or "cracks," and electrical faults might be labeled "abnormal current" or "overvoltage."

[0116] Location: For each detected defect, the system provides a precise location marker. For example, acoustic distortion will indicate the affected frequency band, cosmetic defects will be marked on the image, and electrical faults will indicate the time of occurrence or the location of the electrical component. The system provides a detailed defect report based on these markers.

[0117] If an anomaly is detected, the system immediately transmits the test results to the production control center in real time, automatically triggering the sorting system within the production control system to ensure that unqualified products are promptly removed. Furthermore, the system automatically adjusts production process parameters, such as glue dosage and assembly pressure, based on the type of anomaly detected to prevent similar issues from recurring. This dynamic feedback mechanism not only improves production efficiency but also ensures consistent and stable product quality.

[0118] Once the system identifies an anomaly and diagnoses it, it transmits the test results to the production control system in real time, automatically triggering the sorting system within the production control system to ensure that substandard products are promptly removed. Furthermore, the system automatically adjusts production process parameters, such as glue dosage and assembly pressure, based on the type of anomaly detected to prevent similar issues from recurring. This dynamic feedback mechanism not only improves production efficiency but also effectively ensures the consistency and stability of product quality.

[0119] Once an anomaly is identified and located, the system promptly transmits the detection results to the production control system. This process is typically achieved through Industrial Internet of Things (IIoT) technology, where the detection equipment and the production control system are connected in real time via a high-speed data communication network.

[0120] Data transmission: Abnormal diagnosis results are sent to the production control end via data interfaces (such as OPC UA and MQTT protocols). The data includes information such as the defect type (such as acoustic performance abnormalities, cosmetic defects, and electrical failures), the specific location of the defect, and the severity of the defect.

[0121] Real-time feedback mechanism: Upon receiving an abnormal diagnosis result, the production control system immediately activates the appropriate response process. For example, if a defective audio product is detected in a production process, the control system will trigger an alarm and issue a handling instruction to the operator. This instruction typically includes the identification number of the problem product, the defect type, and the corresponding remedial measures.

[0122] After the anomaly is identified, the sorting system of the production control system will automatically screen out unqualified products and remove them from the production control system based on the type and location of the detected anomaly.

[0123] Sorting Algorithms: Sorting systems make real-time decisions based on exception information. For example, using a product's unique ID number (such as a barcode or RFID tag), the sorting system can quickly identify the product being inspected and determine whether it has a defect. If the defect is severe and cannot be repaired, the product is automatically removed from the production control system. The algorithmic logic for this process typically uses set thresholds to automatically classify qualified and unqualified products.

[0124] Sorting Mechanism: Based on abnormal feedback, the sorting system can take swift action. If acoustic performance is abnormal (such as a distorted frequency response curve), the system will determine whether the product should continue to flow through the production control system based on the acoustic detection data. If a cosmetic defect (such as a scratch or assembly deviation) is detected, the system will determine whether to conduct further manual intervention or reject the product based on the type of defect. Electrical faults (such as abnormal current or voltage) will also result in the timely identification and isolation of products.

[0125] Automated control: Sorting systems typically perform automated operations through robotic arms, conveyor control systems, or automatic push devices. These devices can quickly isolate abnormal products to prevent unqualified products from entering subsequent production or delivery links.

[0126] In addition to screening out substandard products, the system can also automatically adjust production process parameters based on the type and severity of detected anomalies to ensure similar issues do not recur in subsequent production runs. This process is typically achieved through machine learning or rule-based systems, which make adjustments based on historical data and real-time anomaly information.

[0127] Logic for adjusting process parameters: First, the system analyzes the relationship between the detected anomaly and the production process. For example, if a batch of audio products exhibits distortion in the frequency response curve, the system will further analyze whether this distortion is related to process parameters such as glue dosage and assembly pressure. If the data analysis indicates a correlation, the system will automatically issue instructions to adjust the process parameters.

[0128] Adjustment rules: For example, during the production process, process parameters such as glue amount and assembly pressure may affect the acoustic performance of a speaker. If the system detects excessive distortion, it may infer factors such as uneven glue application or insufficient assembly pressure. The system will automatically adjust the glue application amount or increase the assembly pressure according to pre-set process rules.

[0129] Glue dosage adjustment: If the system detects frequent acoustic distortion associated with excessive or insufficient glue, it automatically adjusts the glue application rate. This adjustment is based on past production data and real-time anomaly feedback. For example, if an audio product exhibits high distortion, the system will automatically reduce the glue application rate to mitigate distortion caused by uneven glue application.

[0130] Assembly pressure adjustment: If appearance defects (such as assembly deviation) occur frequently, the system will analyze whether they are caused by insufficient or uneven assembly pressure. If the system detects deviations during the assembly process, it will automatically adjust the assembly machine pressure to better meet standard process requirements.

[0131] Dynamic Adjustment Feedback: The production control system monitors the effects of adjusted process parameters in real time and optimizes them based on new production data. Through this feedback loop, the system continuously improves the production process and ensures consistent product quality.

[0132] Through continuous dynamic feedback, the production system can self-optimize, improve production efficiency and product quality consistency. This mechanism can continuously improve through the following ways: Real-time monitoring and learning: As the production process progresses, the system accumulates a large amount of real-time data. This data includes test results for each product, adjusted process parameters, and the quality performance of subsequent products. The system analyzes this data and uses the new knowledge to optimize future process adjustment strategies.

[0133] Machine Learning and Prediction: Based on historical data and real-time feedback, the system uses machine learning algorithms to predict the types of defects that may occur in the future and take preventative measures. By continuously optimizing process parameter adjustment strategies, the system can effectively reduce the occurrence of substandard products and improve production efficiency and quality.

[0134] Automation and Human Collaboration: Although the system is highly automated, when abnormal conditions exceed preset parameters, the system notifies human operators through real-time reporting. This collaboration between human and automated systems further improves production flexibility and efficiency.

[0135] Through this dynamic feedback mechanism, the system ensures the consistency and stability of product quality and promptly adjusts production process parameters when problems are discovered, thereby reducing waste and the number of substandard products. This not only improves production efficiency, but also reduces production costs and enhances the market competitiveness of the final product.

[0136] Throughout this specification, references to terms such as "one embodiment," "example," or "specific example" indicate that the specific features, structures, materials, or characteristics described in conjunction with that embodiment or example are included in at least one embodiment or example of the present invention. In this specification, schematic representations of these terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in any one or more embodiments or examples.

[0137] The preferred embodiments of the present invention disclosed above are intended only to help illustrate the present invention. These preferred embodiments do not exhaustively describe all details, nor do they limit the present invention to specific embodiments. Obviously, many modifications and variations are possible based on the contents of this specification. These embodiments are selected and described in detail in this specification to better explain the principles and practical applications of the present invention, thereby enabling those skilled in the art to better understand and utilize the present invention. The present invention is limited only by the claims and their full scope and equivalents.

Claims

1. An intelligent detection method for audio manufacturing anomalies, characterized by: The detection method comprises the following steps: Acquire acoustic data through a microphone array, and use sensor equipment to detect the internal structure and circuit of the speaker to obtain internal image data and electrical data; Perform feature extraction on each type of data to refine feature parameters related to defect detection; Build a twin model of the sound system based on characteristic parameters, and divide the twin model into multiple sub-models according to the needs of each link; A hybrid neural network model is constructed based on the twin model to capture long-range dependencies and identify abnormalities in sound appearance and structure; Combine multi-source feature parameters and input them into a hybrid neural network model to perform defect identification and linkage analysis; After completing the multi-link data linkage analysis, the test results are diagnosed for abnormalities to identify various types of abnormalities in the audio system; If an abnormality is identified, the detection results will be fed back to the production control end in real time, unqualified audio will be automatically screened out, and the production process parameters will be automatically adjusted according to the type of abnormality detected.

2. The intelligent detection method for audio manufacturing anomalies according to claim 1, characterized in that: Building a hybrid neural network model based on the twin model to capture long-range dependencies and identify abnormalities in sound appearance and structure includes the following steps: Hybrid neural network models include Transformer and CNN; After Transformer and CNN process time series data and image data respectively, they fuse the time series data and image data and conduct joint analysis: Small sample learning technology is introduced to increase the amount of training data for the hybrid neural network model.

3. The intelligent detection method for audio manufacturing anomalies according to claim 2, characterized in that: The fusion and joint analysis of time series data and image data includes the following steps: The feature vectors from Transformer and CNN are concatenated to form a feature vector. The fused feature vector is used as input and sent to the fully connected layer for classification or regression analysis.

4. The intelligent detection method for audio manufacturing anomalies according to claim 3, characterized in that: The Transformer data processing steps include: Learn the correlation between various time points in the data by calculating the weighted sum of each time step in the input sequence; Add positional encoding to the input data and sense the order of the input data; Learn information from data from multiple perspectives through multi-head attention; The CNN includes convolutional layers, pooling layers and fully connected layers: Convolutional layer: The input image is processed layer by layer through multiple convolutional layers. Each convolution operation extracts low-level features of the image. As the network layer increases, the features extracted by the convolutional layer become increasingly complex. Pooling layer: used to reduce the size of the feature map. Pooling operations include maximum pooling and average pooling. Fully connected layer: After convolution and pooling, the feature map is flattened and input into the fully connected layer for the final classification or regression task. The fully connected layer learns the nonlinear relationship of the feature map.

5. The intelligent detection method for audio manufacturing anomalies according to claim 4, characterized in that: Perform abnormal diagnosis on the test results to identify various abnormal types in the audio system, including the following steps: By analyzing the frequency response curve and distortion acoustic characteristics of the speakers, and comparing the actual collected frequency response data with the expected standard values, frequency response distortion, excessive harmonic distortion, or other sound quality issues can be identified. When the acoustic performance is abnormal, the abnormal location is marked within the time window. The edge detection algorithm extracts details from the surface image of the speaker, identifies scratches or cracks, compares the image with the standard appearance, and detects deviations in the speaker assembly process. Once defects are detected, the location of the defect is marked on the image. By analyzing the waveform of the electrical signal, abnormal fluctuations in current or voltage can be identified, sudden changes in voltage and current can be detected to identify short circuit and overload faults, the time point when the electrical fault occurs can be calibrated, and the cause of the fault can be provided.

6. The intelligent detection method for audio manufacturing anomalies according to claim 5, characterized in that: Performing abnormality diagnosis on the test results to identify various types of abnormalities in the audio system also includes the following steps: After completing various types of abnormality detection, mark the acoustic abnormalities, appearance defects and electrical faults separately, and record the location of the defects; Defect type marking: Each defect is classified according to the preset standards and its severity is marked; Location positioning: For each detected defect, a location tag is provided and a defect report is provided based on these tags.

7. The intelligent detection method for audio manufacturing anomalies according to claim 6, characterized in that: If an anomaly is identified, the test results are fed back to the production control end in real time, unqualified audio is automatically screened out, and production process parameters are automatically adjusted according to the type of anomaly detected, including the following steps: After the anomaly is identified and located, the detection results are transmitted to the production control system; The sorting system of the production control system automatically screens out unqualified products and removes them from the production control system based on the type and location of the detected anomalies; Automatically adjust production process parameters based on the type of abnormality detected.

8. The intelligent detection method for audio manufacturing anomalies according to claim 7, characterized in that: The twin model of the sound system is constructed based on the characteristic parameters and divided into multiple sub-models according to the requirements of each link. The following steps are included: The acoustic sub-model of the speaker is constructed using frequency response curves and distortion acoustic data. The appearance quality sub-model includes surface defects and assembly deviations of the speaker. The electrical performance sub-model reflects the health status of the electrical system through current and voltage signals. Each sub-model is trained independently, and the feature outputs of different sub-models are fused to form a comprehensive feature vector, which serves as the input of the twin model; After the twin model is completed, a digital image is generated for each audio product for comparison with the actual product. The digital image includes the acoustic performance, appearance quality and electrical performance of the audio.

9. The intelligent detection method for audio manufacturing anomalies according to claim 8, characterized in that: Perform feature extraction on each type of data to refine the feature parameters related to defect detection, including the following steps: De-noising, data cleaning and standardization of the collected raw data; For acoustic data, the time domain signal is converted into the frequency domain signal through Fourier transform, and the response characteristics of the audio equipment at different frequencies are extracted; For image data, edge detection algorithm is used to extract edge information of the acoustic surface; For electrical data, the time series signals of voltage and current are analyzed to extract the signal change patterns.

10. An intelligent detection system for audio manufacturing anomalies, used to implement the detection method according to any one of claims 1 to 9, characterized in that: It includes multi-source feature extraction module, model building module, and anomaly identification and automatic adjustment module; Multi-source feature extraction module: This module acquires acoustic data through a microphone array and uses sensor equipment to detect the internal structure and circuits of the speaker, obtaining internal image data and electrical data. It then extracts features from each type of data and refines feature parameters relevant to defect detection. Model building module: This module builds a twin model of the sound system based on characteristic parameters and divides the twin model into multiple sub-models according to the requirements of each link. A hybrid neural network model is then constructed based on the twin model to capture long-range dependencies and identify abnormalities in the sound system's appearance and structure. Abnormal identification and automatic adjustment module: Combines multi-source feature parameters and inputs them into the hybrid neural network model to perform defect identification and linkage analysis, diagnoses abnormalities in the test results, and identifies various abnormal types in the sound. If an abnormality is identified, the test results are fed back to the production control end in real time, automatically screening out unqualified sound, and automatically adjusting the production process parameters according to the detected abnormality type.

Citation Information

Patent Citations

  • Sound inspection management system

    CN115460530A

  • Digital twinning-based multi-modal anomaly detection system

    CN119128778A

  • Product line full-process intelligent detection equipment and optimization analysis system

    CN120031542A

  • Digital twinborn visual modeling method and system based on neural network

    CN120196672A

Cited By

  • X-ray-based integrated detection method for multi-process defects of special-shaped component

    CN121563915A