A loudspeaker sound vibration closed-loop self-checking method based on built-in sweep signal self-excitation
By incorporating a built-in frequency sweep signal self-excitation method, combined with a fully connected vibration fingerprint encoder and evolution trajectory map, the interpretability and real-time performance issues of speaker fault diagnosis are solved, achieving high accuracy and visual diagnosis of speaker faults, applicable to scenarios such as consumer electronics and car audio systems.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- SHENZHEN FENDA TECH CO LTD
- Filing Date
- 2026-02-25
- Publication Date
- 2026-05-29
AI Technical Summary
Existing speaker fault diagnosis technologies cannot accurately identify structural assembly defects and abnormal phenomena such as material aging. They lack interpretability and real-time performance, and embedded solutions are difficult to achieve full-process localization and visual diagnosis.
The method employs a built-in frequency sweep signal self-excitation method, obtains the time-domain signal through anti-aliasing filtering, performs piecewise windowed short-time Fourier transform to generate a complex time-frequency matrix, constructs a fully connected vibration fingerprint encoder, generates a vibration evolution trajectory map, and outputs interpretable thermal tags through two-dimensional plane projection and color depth mapping.
It achieves highly accurate identification and interpretable diagnosis of speaker faults, supports dynamic trend analysis and visual diagnosis on resource-constrained platforms, and improves early fault detection rate and operation and maintenance efficiency.
Smart Images

Figure CN122120689A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of acoustic signal processing and embedded intelligent diagnostic technology, and in particular to a closed-loop self-testing method for loudspeaker vibration based on self-excitation of built-in sweep frequency signals. Background Technology
[0002] Currently, in the loudspeaker and audio equipment industry, vibration detection and fault diagnosis technologies generally rely on spectrum analysis methods using built-in sweep excitation and microphone feedback signals. Mainstream solutions typically employ digital signal processors (DSPs) or microcontrollers (MCUs) for time-domain acquisition and traditional frequency-domain analysis, extracting the loudspeaker response spectrum data based on algorithms such as Short-Time Fourier Transform (STFT) and Fast Fourier Transform (FFT). In practical engineering, spectrum analysis results are often output in numerical vector form, such as frequency amplitude spectra and harmonic distortion parameters, and are combined with signal thresholds or empirical rules for fault diagnosis. With the widespread adoption of end-side intelligent detection modules in miniaturized audio equipment, the industry is beginning to pursue the popularization of embedded self-testing technology to improve the efficiency and accuracy of full-line inspection and user-end self-diagnosis.
[0003] In existing patents and academic literature, closed-loop frequency sweep detection, spectral energy analysis, and some simple statistical features such as peak amplitude and resonant frequency shift have become representative technical approaches. These methods can achieve initial screening of loudspeaker performance and automatic identification of some typical faults (such as severe voice coil adhesion, open circuits, and short circuits), but deeper structural assembly defects, material aging, nonlinear distortion, and other abnormal phenomena are difficult to accurately determine using a single spectral vector or single threshold rule. Moreover, most embedded end-side solutions only output "PASS / FAIL" conclusions, lacking physical explanations of fault mechanisms or information tracing capabilities, and cannot meet the needs of actual maintenance processes for "fault cause localization" and "diagnostic evidence backtracking."
[0004] Typical application scenarios for this technology include rapid production line inspection, after-sales maintenance screening, and user-end health self-checks. In these scenarios, the system needs to perform frequency response testing on the speaker end-side using self-excited frequency sweeping and automatically output fault judgment results that can be used for quality management or maintenance guidance. However, current industry solutions mostly focus on rapid judgment accuracy. Whether based on traditional digital signal processing parameters (such as distortion ratio and resonant frequency offset) or simplified statistical analysis, they cannot present the physical causal relationship between the frequency domain data and the specific structural or material defects of the speaker. Some high-end solutions attempt to introduce cloud-based deep learning inference, but this leads to practical problems such as "opaque black-box judgment," "inability to handle local computing resources," and "lack of real-time traceability of data uploads," and is also unsuitable for embedded testing scenarios with extremely high requirements for interpretability and engineering verifiability.
[0005] The current state of the industry reflects the following technological shortcomings: Frequency domain analysis results lack interpretability. Existing end-side spectrum analysis often outputs numerical vectors, making it difficult for users or maintenance personnel to understand the cause of the fault, resulting in insufficient reliability of the judgment conclusions.
[0006] Spectral characteristics cannot be directly mapped to physical defects. The relationship between response data and loudspeaker structure / materials / assembly process lacks modeling, failing to support the generation and tracing of structured diagnostic evidence.
[0007] Edge-side intelligent decision-making relies on complex models or cloud-based inference. Many solutions cannot achieve end-to-end localization on resource-constrained embedded MCUs / DSPs, making it difficult to guarantee real-time performance, data security, and cost control.
[0008] The diagnostic output lacks visualization and human-computer interaction support. Users and engineers cannot intuitively view abnormal spectrum patterns and corresponding structural faults, which is not conducive to early warning and cannot form a closed loop for maintenance guidance. Summary of the Invention
[0009] This application provides a speaker vibration closed-loop self-testing method based on built-in sweep frequency signal self-excitation, which aims to solve one of the problems or problems of the prior art mentioned in the background art.
[0010] This application provides a closed-loop self-test method for loudspeaker vibration based on built-in sweep frequency signal self-excitation, specifically including: S1: Acquire the feedback signal of the loudspeaker after being excited by the frequency sweep, and perform anti-aliasing filtering to obtain a time-domain signal sequence with a sampling rate of a preset frequency; S2: Extract stable response segment data from the time-domain signal sequence, extract valid data from the stable response segment data, and form a stable response segment time-domain signal; S3: Perform piecewise windowed short-time Fourier transform processing on the time-domain signal of the stable response segment to generate a complex time-frequency matrix covering a preset frequency band; S4: Determine the interpretable feature vector based on the complex time-frequency matrix; S5: Construct a fully connected vibration fingerprint encoder based on the pre-stored mapping relationship between speaker structure-material-assembly defects and spectral morphology; input the interpretable feature vector into the fully connected vibration fingerprint encoder to obtain the vibration fingerprint code; S6: Perform evolution trajectory construction processing on the vibration fingerprint code sequence obtained by self-inspection, and generate vibration evolution trajectory map through two-dimensional plane projection and color depth mapping; S7: The sound fingerprint code of the sound evolution trajectory map is binarized according to a preset threshold, and matched with the fault-fingerprint-handling suggestion triplet knowledge base to generate interpretable thermal labels and natural language judgment conclusions, and output fault diagnosis evidence.
[0011] The speaker vibration closed-loop self-testing method based on built-in sweep frequency signal self-excitation provided in this application has the following beneficial effects: (1) The embedded end-side interpretable spectrum analysis method provided in this application significantly improves the accuracy and interpretability of fault identification compared with the traditional self-testing mechanism based on overall performance threshold or simple amplitude comparison. In the prior art, the factory testing and self-diagnosis of audio equipment mostly rely on manual listening or only on macroscopic indicators such as total harmonic distortion and frequency response flatness, which makes it difficult to locate specific defect types and is not sensitive to early degradation; more importantly, such methods cannot answer the question of "why it is judged as abnormal", resulting in a lack of handling basis at the repair end. This solution constructs a six-dimensional interpretable feature vector - including resonant peak offset, envelope asymmetry, high-frequency attenuation slope, odd harmonic energy ratio, transient tail length and number of phase change points - starting from the physical acoustic mechanism, accurately characterizing the typical response mode of the loudspeaker under different assembly defects. Each dimension corresponds to a clear mechanical structural abnormality characterization. For example, magnetic circuit eccentricity will lead to symmetry destruction, and voice coil adhesion will cause nonlinear vibration, which manifests as drift. This feature system not only avoids the untraceability of the input-output relationship in black-box models, but also allows each criterion to be directly verified for its rationality by acoustic engineers.
[0012] (2) Furthermore, this solution combines a lightweight "vibration fingerprint encoder" with locally fixed weights with an evolution trajectory map to achieve highly robust dynamic trend analysis and defect clustering capabilities on a resource-constrained embedded platform, effectively overcoming the problem of misjudgment caused by noise interference in single-moment snapshot detection. Unlike conventional deep learning, which requires a large amount of training data and cloud inference support, this solution adopts a feedforward neural network structure without a training process. Its weights are fixed in the MCU after offline calibration and are only used to map the six-dimensional physical features into a 32-dimensional standardized "defect activation intensity" vector. Each dimension corresponds to a predefined typical fault mode, realizing a transparent conversion from signal features to fault semantics. More importantly, by constructing a two-dimensional projection trajectory from multiple consecutive self-test results and superimposing color depth indicators, the system can visualize the evolution path of the equipment's health status, identify slow deterioration trends or periodic fluctuation characteristics, and thus provide early warning of potential failure risks. For example, when the trajectory continues to migrate towards the "dust cap detachment" area, even if the alarm threshold has not been triggered, it can prompt maintenance personnel to pay attention to the sealing degradation. This state evolution-based analysis paradigm significantly improves the early fault detection rate and avoids the problem of poor adaptability of the traditional static threshold method in the face of individual differences.
[0013] The aforementioned technologies work synergistically to construct an intelligent self-inspection system with closed-loop feedback capabilities and full-process visibility and controllability. This system not only meets the real-time requirements of embedded devices, but also supports uploading all intermediate data to the quality traceability system via the low-bandwidth BLE protocol, facilitating batch analysis and process optimization on the production side. The final natural language judgment conclusion, combined with local spectrum highlighting, forms an intuitive and interpretable thermal label, which is simultaneously pushed to an LED screen or mobile app, enabling non-professional users to understand the cause of the fault and troubleshooting suggestions. Therefore, this solution fundamentally solves the trust bottleneck of "judgment without understanding, inspection without application" in traditional audio equipment self-inspection, truly achieving a leap from "whether it can produce sound" to "why the sound is abnormal," and is suitable for application scenarios with high requirements for reliability and operational efficiency, such as consumer electronics, car audio, and industrial broadcasting. Attached Figure Description
[0014] Figure 1 The main flowchart of the speaker vibration closed-loop self-testing method based on built-in sweep frequency signal self-excitation, which is a preferred embodiment of this application, is shown in the following diagram.
[0015] Figure 2 This is a sub-flowchart of a speaker vibration closed-loop self-testing method based on a built-in sweep frequency signal self-excitation, which is a preferred embodiment of this application.
[0016] Figure 3 This is another sub-flowchart of the speaker vibration closed-loop self-testing method based on built-in sweep frequency signal self-excitation, which is a preferred embodiment of this application. Detailed Implementation
[0017] Embodiments of the present invention are described in detail below, examples of which are shown in the accompanying drawings, wherein the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present invention, and should not be construed as limiting the present invention.
[0018] The following disclosure provides many different embodiments or examples for implementing different structures of the invention. To simplify the disclosure, specific examples of components and arrangements are described below. Of course, these are merely examples and are not intended to limit the invention. Furthermore, reference numerals and / or letters may be repeated in different examples; such repetition is for simplification and clarity and does not in itself indicate a relationship between the various embodiments and / or arrangements discussed.
[0019] like Figure 1 As shown, this application provides a speaker vibration closed-loop self-testing method based on built-in sweep frequency signal self-excitation, specifically including: S1: Acquire the feedback signal of the loudspeaker after being excited by the frequency sweep, and perform anti-aliasing filtering to obtain a time-domain signal sequence with a sampling rate of a preset frequency; S2: Extract stable response segment data from the time-domain signal sequence, extract valid data from the stable response segment data, and form a stable response segment time-domain signal; S3: Perform piecewise windowed short-time Fourier transform processing on the time-domain signal of the stable response segment to generate a complex time-frequency matrix covering a preset frequency band; S4: Determine the interpretable feature vector based on the complex time-frequency matrix; S5: Construct a fully connected vibration fingerprint encoder based on the pre-stored mapping relationship between speaker structure-material-assembly defects and spectral morphology; input the interpretable feature vector into the fully connected vibration fingerprint encoder to obtain the vibration fingerprint code; S6: Perform evolution trajectory construction processing on the vibration fingerprint code sequence obtained by self-inspection, and generate vibration evolution trajectory map through two-dimensional plane projection and color depth mapping; S7: The sound fingerprint code of the sound evolution trajectory map is binarized according to a preset threshold, and matched with the fault-fingerprint-handling suggestion triplet knowledge base to generate interpretable thermal labels and natural language judgment conclusions, and output fault diagnosis evidence.
[0020] Step S1: Under the conditions of speaker power-on self-test or periodic detection trigger, the feedback signal of the speaker after being excited by frequency sweep is collected through the built-in microphone, and anti-aliasing filtering is performed to obtain a time-domain signal sequence with a sampling rate of 16kHz. Specifically, this includes: S1.1: Detect and process power-on self-test trigger events or periodic detection trigger events for the audio system's operating status to generate a trigger confirmation signal.
[0021] After the audio system enters the working state, a polling detection method for the running status is adopted (parameters: sampling period 10ms, detection signal source is system control bus) to realize real-time acquisition and logical recognition of power-on events and periodic detection events.
[0022] Based on the collected system control bus data, an event type determination algorithm (parameter: event code matching table includes power-on self-test code 0xA1 and periodic detection code 0xB2) is used to accurately determine the type of the triggered event and obtain the event type identifier data.
[0023] Furthermore, through the event validity verification processing method (parameter: status register comparison mechanism, allowable error window ±2ms), the timing consistency of the determined event category is checked, possible false triggers are eliminated, and a verified event category identification result is generated.
[0024] Furthermore, by triggering the priority arbitration algorithm (parameters: priority weight on startup = 2, periodic detection = 1), the effective event categories are prioritized and sorted to obtain the sorted list of triggered events.
[0025] Furthermore, by parsing the event list and generating a trigger confirmation method (parameter: if the weight of the first element in the list is ≥1, then trigger confirmation is achieved), the highest priority event is extracted from the sorted event list as the current trigger source, and the corresponding trigger confirmation signal data is generated.
[0026] By combining the event type determination algorithm with the trigger confirmation generation method, the event detection result of the previous step is transformed into a trigger confirmation signal that can directly drive the subsequent microphone bias control, thereby ensuring the orderly start of the signal acquisition chain and the synchronization of the frequency sweep detection process.
[0027] For example, in an audio control unit equipped with a Cortex-M4 microcontroller, the sampling period of the polling detection method is set to 10ms, and the status byte length of the system control bus is 4 bytes, where the second byte is the event code field. When the event code matching table of the bus data detects a value of \n When this occurs, it is determined to be a power-on self-test event, and the value matched is \n When this occurs, it is determined to be a periodic detection event. The event validity verification process compares the difference between the timestamp of the currently detected event code and the system clock register to see if it is less than \n. Milliseconds ensure the timing accuracy of event recognition; when the priority arbitration algorithm executes, if a power-on event is detected, a weight is assigned. If periodic detection events exist simultaneously, assign weights. Within milliseconds, and significantly improves the synchronization and data validity of subsequent frequency sweep signal acquisition.
[0028] S1.2: Based on the trigger confirmation signal, apply a bias voltage to the built-in microphone to activate the microphone sensing unit.
[0029] S1.3: The acoustic feedback signal of the speaker after being excited by the frequency sweep is processed by the activated built-in microphone to obtain an analog voltage signal.
[0030] Based on the activated state of the built-in microphone sensing unit output by S1.2, the electret condenser microphone uses the sound-to-electric conversion principle (parameters: sensitivity -42dB, internal resistance 2.2kΩ) to realize the physical conversion of the sound pressure waveform generated by the speaker under frequency sweep excitation from the mechanical diaphragm to the change in the distance between the capacitor plates, and converts it into a charge change signal proportional to the sound pressure.
[0031] Furthermore, a charge amplification circuit (parameters: JFET input operational amplifier, gain set to +6dB, bandwidth 10Hz~20kHz) is used to achieve precise conversion of the charge change signal to the corresponding voltage signal, and a millivolt-level analog voltage waveform is obtained, which maintains the amplitude and phase correspondence with the sound pressure wave.
[0032] Furthermore, a bias network stabilization circuit (parameters: bias resistor 2.2kΩ, power supply voltage 2.0V±1%) is adopted to fix the DC operating point of the microphone output, ensuring that the amplitude linear range of the analog voltage signal does not experience clipping distortion within the measurement bandwidth, and generating a DC bias-stabilized analog waveform output.
[0033] Furthermore, by using a source follower buffer circuit (parameters: high input impedance >1MΩ, low output impedance <100Ω), impedance matching between the microphone analog voltage and the downstream anti-aliasing filter is achieved, reducing reflection and attenuation during signal transmission and maintaining the flatness of amplitude-frequency characteristics.
[0034] The following method for calculating the acoustic-to-electrical conversion gain is used to calculate the output voltage at a specific sound pressure level using a formula: in, This is the output voltage (volts). Sound pressure level (decibels). The value is the microphone sensitivity (volts per Pa). Calculations using this formula verify that the output voltage is approximately 7.94mV under conditions of 94dB SPL and a sensitivity of -42dB, achieving a quantized mapping from sound to electricity.
[0035] Through the aforementioned acoustic-to-electric conversion link, the sound pressure signal is converted into an analog voltage signal with complete amplitude and phase without distortion, thus achieving high-fidelity input conditions for subsequent anti-aliasing filtering processing.
[0036] For example, in a production test scenario, the speaker is driven by the MCU to generate a linear sweep signal of 20Hz~5kHz, maintaining a sound pressure level of 94dB SPL. The built-in microphone sensitivity is -42dB (approximately 7.94mV / 94dB), a JFET input operational amplifier is used with a gain of +6dB, and the bandwidth covers 10Hz~20kHz. The bias resistor is set to 2.2kΩ, the DC supply voltage is 2.0V, and impedance matching with the downlink filter module is achieved through a source follower buffer circuit. Under these conditions, the analog voltage waveform output by the microphone maintains linear amplitude across the entire frequency range, based on the formula... Calculations show that the output amplitude at 20Hz is approximately 7.94mV, and the amplitude attenuation at 5kHz is less than 0.5dB. The phase difference between the waveform and the frequency sweep wave of the sound source is less than 5°, providing a high-fidelity and quantifiable analog signal input for subsequent anti-aliasing low-pass filtering and analog-to-digital conversion, significantly improving the reliability of frequency domain feature analysis.
[0037] S1.4: Perform anti-aliasing low-pass filtering on the analog voltage signal to filter out frequency components higher than 8kHz and generate an anti-aliasing filtered signal.
[0038] S1.5: After anti-aliasing filtering, the signal is processed by analog-to-digital conversion at a sampling rate of 16kHz to obtain a digital time-domain signal sequence with a sampling rate of 16kHz.
[0039] Step S2: Extract stable response segment data from the time-domain signal sequence, and extract 4096 valid data points based on the signal energy stability window determination rule to form the stable response segment time-domain signal to be analyzed. Specifically, this includes: S2.1: Based on the statistical analysis results of historical frequency sweep data, a signal energy variance threshold of 0.5dB and a minimum stable duration of 100ms are set, and a signal energy stability window determination rule is constructed. This rule defines a continuous interval where the energy change rate is lower than the threshold as an effective stability window, and outputs a set of signal energy stability window determination rule parameters to provide a technical basis for locating stable response segments.
[0040] Based on the characteristic distribution of historical frequency sweep data, a statistical analysis method (parameters: sample size N=1000, self-test period T=24h) is used to extract the energy stability characteristics of the loudspeaker under frequency sweep excitation. Furthermore, through the variance calculation method (parameters: root mean square energy sequence length L=4096), the energy fluctuation of each time domain sampling segment is quantified, and the statistical average variance value and extreme value distribution results are obtained.
[0041] Furthermore, based on the statistical results, a threshold setting method (parameters: variance threshold = 0.5dB, minimum stable duration = 100ms) is used to mathematically model the energy stability window determination rules and generate a set of rule parameters including the threshold and duration. Further, a rate of change analysis method (parameters: sliding step size Δt = 5ms) is used to identify time intervals where the energy change rate is below the threshold, and these intervals are used as a candidate set of stable windows. Finally, an interval matching algorithm (parameters: threshold = 0.5dB, duration frame k >= 20) is used to filter the validity of the candidate windows and output a set of stable window position indices.
[0042] By adopting a standardized mapping process, the set of stable window determination rule parameters is transformed into technical indicators that can be called by the subsequent S2.2 steps, thereby achieving the expected technical effect of accurate positioning of the stable response segment.
[0043] For example, a certain model of audio equipment with a rated resonant frequency of 320Hz was selected. Statistical analysis was performed on the frequency sweep self-test data from its production line over 30 consecutive days. The sample size was 720 times, with 4096 sampling points per sample. After calculating the variance of each root mean square energy sequence, the variance ranged from 0.35dB to 1.2dB. The energy stability window determination formula was used: in, Let be the root mean square energy value of the i-th frame. This represents the average energy value of this segment. The total number of frames within the segment is given. The calculated σ is compared with a threshold of 0.5dB to filter out continuous intervals that satisfy σ≤0.5dB, and their duration is checked to ensure it is greater than 100ms. Based on the sample statistics for this device, the average length of the effective stable window is 135ms, covering the steady-state region of the swept frequency signal. After the rule parameter set is output, the subsequent S2.2 step can use this to locate the starting point of the window, achieving precise truncation of the steady-state response segment, significantly improving the stability and interpretability of the frequency domain analysis input data.
[0044] S2.2: Perform sliding window energy calculation processing on the time-domain signal sequence to generate a signal energy sequence; specifically, use a 50ms Hanning window sliding window with a 5ms step size to calculate the root mean square energy value, and output the signal energy sequence composed of the energy of each frame as input data for stable window positioning.
[0045] S2.3: Based on the signal energy stability window determination rule, perform stability window positioning processing on the signal energy sequence to determine the starting point of the stable response segment; by detecting continuous intervals with energy variance below 0.5dB and duration exceeding 100ms, output the starting point index of the stability window to provide a precise location reference for data interception.
[0046] Based on the signal energy sequence output by step S2.2, a continuous interval scanning algorithm (parameters: energy variance threshold 0.5dB, minimum duration 100ms) is used to achieve preliminary identification of candidate intervals for stable windows.
[0047] Furthermore, by using the sliding interval variance calculation method (parameters: window length 100ms, step size 5ms), the interval energy variance vector is generated, and the variance value sequence of each sliding interval is obtained for stability determination.
[0048] Furthermore, a threshold comparison method (parameter: variance ≤ 0.5dB) combined with a time accumulation detection method is used to achieve interval filtering that meets the variance threshold and has a continuous duration ≥ 100ms, and a stable interval index list that meets the conditions is generated.
[0049] Furthermore, by using the stable interval start point positioning algorithm (rule: select the earliest time index that meets the conditions as the start point of the stable segment), the starting point of the stable response segment is accurately located, and the starting point position index is output for subsequent data extraction.
[0050] By combining continuous interval scanning with variance thresholding, the energy sequence in step S2.2 is transformed into the starting point index of the stable response segment, thereby achieving accurate calibration of the steady-state data interval and laying a reliable foundation for the subsequent extraction of 4096 stable data points.
[0051] For example, in a speaker frequency sweep test, a sliding window with a Hanning window length of 50ms and a step size of 5ms is used to process a time-domain signal sequence with a sampling rate of 16kHz. The energy value for each frame is obtained when calculating the energy sequence. Using a detection window of 100ms, the formula is: ,in The frame-by-frame energy value within the detection window. The average energy value. To detect the number of frames in the window, the energy variance vector is obtained. In this embodiment, when the sliding window coverage time is between 0.35s and 0.45s, the variance value is 0.46dB and the duration of this state reaches 105ms, which meets the stable interval determination condition. The starting point index of the stable window is the 70th sampling frame, which is used as the starting reference point for subsequent S2.4 data interception. After performing this processing, the stable segment positioning error is significantly reduced, ensuring that the intercepted 4096 data points accurately reflect the swept frequency steady-state response characteristics.
[0052] S2.4: Based on the starting point index of the stable window, extract 4096 consecutive data points from the time-domain signal sequence to form the time-domain signal of the preliminary stable response segment to be analyzed; this process ensures that the extracted segment covers the steady-state response region of the frequency sweep excitation, and outputs the time-domain signal of the preliminary stable response segment as a verification input.
[0053] The input conditions are the starting point index of the stable window determined by the preceding step S2.3 and the digital time-domain signal sequence with a sampling rate of 16kHz generated by anti-aliasing filtering and analog-to-digital conversion.
[0054] An index positioning method (parameters: stable window start point index, sampling rate = 16kHz) is used to achieve precise positioning of the starting point of the digital time domain signal sequence, anchoring the starting position of the truncation operation at the first sample point that meets the energy stability judgment rule.
[0055] Furthermore, by using a fixed-length data truncation algorithm (parameter: truncation length = 4096 points), the batch extraction function of continuous sampling points is realized, and the time-domain signal of the preliminary stable response segment covering the steady-state response of the frequency sweep excitation is obtained.
[0056] Furthermore, a time coverage verification method (parameter: sampling rate = 16kHz) is adopted to physically quantize the time interval corresponding to the intercepted data, so as to confirm that the time length corresponding to the 4096 data points is 256ms, and ensure that this time length falls completely within the steady-state response time window.
[0057] Furthermore, the sidelobe energy monitoring algorithm (parameter: window function type = rectangular window) is used to detect the continuity of energy edges at both ends of the truncated interval, verifying that there are no significant abrupt changes in signal energy at the beginning and end of the truncated segment, so as to ensure the stability of the data before spectrum analysis.
[0058] By combining fixed-length truncation with time coverage verification, the stable window positioning result of the previous step is transformed into a time-domain signal of a preliminary stable response segment with a length of 4096 points, thus achieving the technical effect of the steady-state signal input required for subsequent frequency domain transformation.
[0059] For example, in a self-test process of a certain model of built-in sweep frequency self-test speaker, the starting point index of the stable window is 3200, corresponding to a time of 0.2s. Under a sampling rate of 16kHz, the 3200th sample of the digital time-domain signal is used as the starting point for truncation using an index positioning method. A fixed-length data truncation algorithm is applied to extract 4096 consecutive data points from the 3200th sample, generating a preliminary stable response segment time-domain signal of 256ms. For the truncated data, a time coverage verification method is used to confirm that its time range is from 0.2s to 0.456s, completely covering the sweep frequency steady-state excitation interval. A sidelobe energy monitoring algorithm is used to detect the energy values at the beginning and end of the truncated segment; the root mean square difference is less than 0.05dB, indicating that the signal is stable at the start and end points. Based on this preliminary stable response segment, subsequent S2.5 energy stability verification is performed to obtain the final stable response segment time-domain signal that meets the input requirements for frequency domain analysis, improving the accuracy and interpretability of spectral feature extraction.
[0060] S2.5: Perform energy stability verification processing on the preliminary stable response segment time domain signal to generate the final stable response segment time domain signal; verify that the energy variance of the truncated data meets the 0.5dB threshold requirement of the signal energy stability window judgment rule, and output the final stable response segment time domain signal that conforms to the 4096-point length specification and has stable energy for subsequent frequency domain analysis.
[0061] like Figure 2 As shown, step S3 involves performing a piecewise windowed short-time Fourier transform on the time-domain signal of the stable response segment, using a Hanning window and setting a 50% overlap rate to generate a 512-point complex time-frequency matrix covering the 20Hz to 5kHz frequency band. Specifically, this includes: S3.1: Perform segmentation processing on the stable response segment time domain signal. Based on the preset segment length of 1024 points and the 50% overlap rate parameter, the 4096-point stable response segment time domain signal is divided into 8 overlapping data segments, and a segmented time domain signal sequence is generated as the input for window function processing.
[0062] The time-domain signal of the final stable response segment is segmented using a fixed-length segmentation algorithm (parameters: segment length 1024 points, overlap rate 50%) to decompose the 4096-point input sequence into several analysis units along the time axis.
[0063] Furthermore, by using a sliding index calculation (parameter: step size 512 points), 50% overlap of sampling points between adjacent data segments is achieved, resulting in an index set covering the entire response segment.
[0064] Furthermore, by extracting regional data from the original sequence through an index set, high-precision extraction of 1024 continuous samples corresponding to each segment is achieved, and a segmented time-domain signal list is generated.
[0065] Furthermore, by normalizing the amplitude of the segmented sampled data (parameter: peak amplitude normalized to ±1), the amplitude range of each segment is made consistent, providing balanced input for subsequent window function weighting.
[0066] By using the segmented algorithm described above, the stable response segment result from the previous step is transformed into a segmented time-domain signal sequence of 8 overlapping data segments, thus preparing the input for time-series decomposition and spectrum calculation.
[0067] For example, in a built-in frequency sweep vibration self-test, the final stable response segment time-domain signal is 4096-point 16-bit PCM data with a sampling rate of 16kHz. Applying a segmentation algorithm with a segment length of 1024 points and a step size of 512 points, the segment starting index sequence can be obtained as: [0, 512, 1024, 1536, 2048, 2560, 3072, 3584], corresponding to segments 1 to 8 respectively. Taking segment 1 as an example, its sample range is 0 to 1023 points. After performing peak amplitude normalization on this data, a fixed-length sample sequence with an amplitude range of ±1 is obtained. This process is repeated for the remaining segments to obtain 8 preprocessed segmented samples. The length of each sample segment conforms to the 1024-point specification, and adjacent segments overlap by 512 points, satisfying the time-frequency balance condition required for short-time Fourier transform analysis. The segmented sequence under the parameter configuration significantly improves the spectral smoothness and frequency resolution when the Hanning window function is applied for weighted and FFT operations, ensuring that the separation between the formant region and the background noise region in the final 512-frequency complex time-frequency matrix is significantly improved.
[0068] S3.2: Perform Hanning window function weighting processing on the segmented time-domain signal sequence, and perform amplitude modulation on each data segment based on the mathematical definition of Hanning window to eliminate the spectral leakage effect and generate a windowed segmented time-domain signal sequence as the input of fast Fourier transform.
[0069] For segmented time-domain signal sequences, the Hanning window function weighting method (parameters: window length 1024 points, window coefficient calculation formula based on mathematical definition) is used to achieve amplitude modulation of each signal segment to reduce spectral leakage effect.
[0070] Furthermore, the formula for generating the Hanning window coefficient is used. (in For the index of sampling points within the window, (where the window length is specified), to generate the coefficient sequence for the entire data segment and obtain the corresponding window function coefficient array as the weighted input.
[0071] Furthermore, by using a point-by-point multiplication method (parameters: window coefficient array and segmented data vector), the window coefficients are multiplied by the amplitude of the segmented data to obtain the windowed segmented time-domain signal sequence.
[0072] Furthermore, by using the endpoint energy smoothing algorithm (parameter: the number of sampling points at the front and back ends of the windowed data is 5% of the window length), the end transition of the windowed signal is made smooth, eliminating the instantaneous jump caused by truncation.
[0073] Through the above-mentioned Hanning window weighting and endpoint smoothing, the segmented time-domain signal sequence output by S3.1 is transformed into a windowed segmented time-domain signal sequence with spectral leakage suppression and energy distribution optimization characteristics, thus achieving the expected technical effect of providing high-quality input for subsequent fast Fourier transform.
[0074] For example, in a self-testing scenario for a production-grade audio system, a segmentation operation is performed on the stable response time-domain signal with a length of 4096 points. Each segment has a length of 1024 points and an overlap rate of 50%, resulting in 8 segmented data segments. For each segment, the Hanning window coefficient is applied... (Window length) The calculations show that each coefficient corresponds to a floating-point double-word data format. Through point-by-point multiplication, the weighted amplitude distribution maintains a maximum value of 1.0 at the center of the segment and smoothly decreases to 0 at the endpoints, effectively reducing spectral leakage. After endpoint smoothing, the amplitude transition between the front and rear signals is smooth, and the maximum amplitude difference of instantaneous amplitude changes at the endpoints is reduced to 1 / 5 of the original data. After this windowed segmented signal is input into the FFT, the background noise spectrum rise is significantly reduced, and the sidelobe energy decreases significantly, effectively improving the calculation stability and interpretability of subsequent frequency domain features such as formant shift and envelope asymmetry.
[0075] S3.3: Perform 1024-point Fast Fourier Transform on the windowed segmented time-domain signal sequence, calculate the frequency domain complex representation of each data segment based on the radix-2 FFT algorithm, and generate an initial complex spectrum matrix as the frequency axis calibration input.
[0076] For windowed segmented time-domain signal sequences, the radix-2 FFT algorithm (parameters: transform length 1024 points, input data type is single-precision floating point) is used to perform discrete Fourier transform operations, realizing the transformation of each segmented time-domain signal into a complex sequence in the frequency domain.
[0077] Furthermore, by using the divide-and-conquer recursive structure in the radix-2 FFT algorithm, the butterfly operation formula is used to calculate each level of frequency domain components, and a complex frequency domain vector of length 1024 is obtained, ensuring that the frequency domain amplitude and phase information are not distorted.
[0078] Furthermore, based on the FFT results, the output of each segment is processed by complex sequence storage, and the real and imaginary parts are stored separately in the row vector of the normalized matrix to form a preliminary segmented spectrum matrix.
[0079] Furthermore, through the amplitude calculation method (formula: ,in For the real part of the complex number, (For the imaginary part of the complex number) Extract the amplitude information of each frequency point, and simultaneously retain the original phase data to support subsequent phase feature analysis.
[0080] Furthermore, by sequentially arranging the complex frequency domain vectors of each windowed segment, an initial complex spectrum matrix is generated. The rows of the matrix correspond to the segment number, and the columns correspond to the frequency point index. This matrix serves as the direct input for frequency axis calibration processing.
[0081] By using the radix-2 FFT algorithm, the windowed segmented time-domain signal from the previous step is transformed into an initial complex spectrum matrix containing amplitude and phase information, thus achieving high-precision fidelity of the full-amplitude features in the frequency domain.
[0082] For example, in one implementation, the time-domain signal of the stable response segment with a length of 4096 points is divided into 8 segments of 1024 points each according to the parameters in S3.1 and S3.2, and processed with Hanning window weights. The radix-2 FFT algorithm is used, and the 1024-point FFT operation is performed on the Cortex-M4 microcontroller by calling the DSP library function arm_cfft_f32. The operation time for each segment is approximately 0.34ms. The butterfly operation uses... As the core of amplitude square calculation, and with The amplitude at each frequency point was obtained. Both the real and imaginary parts were stored in a matrix as floating-point arrays with a matrix dimension of 8×1024. Validation results show that the amplitude spectrum deviates from the theoretical frequency response by less than 0.01dB, and the phase error is less than 0.2°, significantly improving the accuracy of subsequent frequency axis mapping and feature extraction.
[0083] S3.4: Perform frequency axis linear mapping processing on the initial complex spectrum matrix. Based on the sampling rate of 16kHz and 1024-point FFT parameters, convert the frequency point index into physical frequency values and generate a physical frequency calibration matrix of 512 effective frequency points covering 0Hz to 8kHz as the frequency band interception input.
[0084] The initial complex spectrum matrix input is processed by frequency index extraction. The column-by-column index scanning algorithm (parameter: number of columns = 512) is used to obtain the index of each frequency point in the FFT calculation result and obtain the frequency index sequence as the basic data for physical frequency mapping.
[0085] Furthermore, the physical frequency value is calculated using the frequency axis linear mapping formula (parameters: sampling rate fs = 16000Hz, FFT points N = 1024), thus converting the frequency index sequence into corresponding physical frequency spatial coordinates; the formula is expressed as follows: in, For elements in the frequency index sequence, Sampling rate, The number of FFT points.
[0086] Furthermore, the above formulas are processed in batches using a vectorized calculation method (parameter: matrix data type = float32) to achieve efficient conversion of all 512 frequency indices and generate a sequence of physical frequency values; this process ensures the calculation accuracy and execution efficiency of the mapping process.
[0087] Furthermore, an effective frequency point truncation algorithm (parameter: highest frequency = fs / 2 = 8000Hz) is adopted to implement upper and lower limit constraints on the physical frequency value, limiting the frequency value range to 0Hz to 8kHz, resulting in a set of 512 effective frequency points covering 0Hz to 8kHz.
[0088] By constructing a physical frequency calibration matrix (parameter: matrix dimension = 512×K), the frequency value sequence is combined with the corresponding time frame index to form a two-dimensional physical frequency calibration matrix, thereby assigning the physical meaning of the frequency axis and providing a precise coordinate mapping basis for subsequent target frequency band interception processing.
[0089] For example, in the self-test of speaker model 506, the initial complex spectrum matrix has a dimension of 512×8, the FFT point count is configured to 1024, the sampling rate is set to 16000Hz, and the index sequence [0,1,2,...,511] is obtained through an index scanning algorithm. The linear mapping formula is then applied. Calculate the physical frequency value for each index. For example, the frequency corresponding to index 320 is... =5000Hz. Under batch vectorized computation, all 512 physical frequency values are generated in a single function call, and all frequency points in the range of 0Hz to 8kHz are retained by the frequency upper limit constraint. Finally, combined with 8 time frame indices, a 512×8 physical frequency calibration matrix is obtained, ensuring that each frequency point has a unique physical value correspondence with its corresponding time frame. This matrix significantly improves the accuracy of coordinate mapping and the interpretability of spectrum analysis in subsequent target frequency band interception and frequency domain feature calculation.
[0090] S3.5: Perform target frequency band truncation processing on the physical frequency calibration matrix. Based on the frequency range threshold of 20Hz to 5kHz, filter and retain the data in the corresponding frequency point index interval 1 to 320, and generate a 512-frequency point complex time-frequency matrix covering the 20Hz to 5kHz frequency band as the input for frequency domain feature extraction.
[0091] A frequency range threshold filtering algorithm (parameters: lower limit 20Hz, upper limit 5kHz) is executed on the physical frequency calibration matrix input to remove frequency components below 20Hz and above 5kHz.
[0092] Furthermore, by using the frequency point index mapping method (parameters: sampling rate 16kHz, FFT length 1024), a one-to-one correspondence between frequency values and index positions is achieved, and the index interval data sequence corresponding to the target frequency band is obtained.
[0093] Furthermore, by using an index interval truncation algorithm (parameter: retain the index interval from 1 to 320), the row and column filtering operations of the original physical frequency calibration matrix are implemented, and a complex spectrum submatrix containing only the target frequency band is generated.
[0094] Furthermore, a frequency band data alignment processing method (parameter: 512 frequency points) is applied to perform column completion and normalization operations on the truncated complex spectrum submatrix, maintaining the matrix dimension consistent with the definition of the time-frequency analysis link.
[0095] Through the above filtering and alignment process, the physical frequency calibration matrix of the previous step is transformed into a complex time-frequency matrix with 512 frequency points covering 20Hz to 5kHz, thereby providing high-precision complex spectrum data within a limited frequency band for the subsequent frequency domain feature extraction module.
[0096] For example, in a speaker self-test scenario, the physical frequency calibration matrix covers 512 frequency points from 0Hz to 8kHz, with a sampling rate set to 16kHz and an FFT length of 1024 points. A frequency range threshold filtering algorithm is executed to filter the frequency values... Hz and > After removing Hz data, a mapping relationship from index 1 to 320 of the effective frequency range is obtained. An index range truncation algorithm is used to retain the complex spectrum within this range, forming a 320-row × 8-column complex submatrix. To meet the input requirements of the feature extraction module, a frequency band data alignment method is used to interpolate and pad the 320 rows of spectrum to 512 rows, while keeping the corresponding time frame number of each column unchanged. The output of this step is a set of 512-frequency complex time-frequency matrices covering the target frequency band, with its amplitude and phase spectra strictly limited to the range of 20Hz to 5kHz. In the subsequent S4 step, this effectively improves the physical relevance of the calculation of features such as resonant frequency offset and envelope asymmetry, and significantly enhances the reliability and interpretability of fault diagnosis.
[0097] like Figure 3 As shown, step S4 involves calculating a six-dimensional interpretable feature vector based on the complex time-frequency matrix, including the formant offset, envelope asymmetry, high-frequency attenuation slope, odd harmonic energy ratio, transient response tail length, and number of phase abrupt change points, forming a structured feature vector characterizing the sound characteristics. Specifically, this includes: S4.1: Perform loudspeaker model matching processing on the pre-stored nominal parameter library to obtain the nominal resonant frequency parameter; perform bandwidth limiting processing on the amplitude spectrum of the complex time-frequency matrix based on the nominal resonant frequency parameter to focus the resonant region data; perform quadratic interpolation peak search algorithm processing on the limited amplitude spectrum to accurately locate the actual resonant frequency; perform difference calculation processing on the actual resonant frequency and the nominal resonant frequency to generate the resonant frequency offset characteristic parameter; perform temperature coefficient compensation processing on the resonant frequency offset characteristic parameter to output the temperature-corrected standardized frequency deviation value.
[0098] Based on the 512-frequency complex time-frequency matrix covering the 20Hz to 5kHz frequency band generated in the previous step S3.5, the nominal resonant frequency is obtained by using the model matching method (parameters: pre-stored nominal parameter library, speaker model identifier).
[0099] Furthermore, by using a bandwidth limiting algorithm (parameter: nominal resonant frequency ± Δf window width), the resonant region of the amplitude spectrum of the complex time-frequency matrix is focused, and a limited spectrum matrix containing only amplitude data near the resonant frequency is obtained.
[0100] Furthermore, a quadratic interpolation peak search algorithm (parameters: interpolation sampling interval 0.1Hz, search range ±Δf) is adopted to accurately locate the measured resonant frequency within the restricted spectrum matrix and generate the actual resonant frequency value.
[0101] Furthermore, by using the difference calculation method (parameters: actual resonant frequency, nominal resonant frequency), the characteristic parameters of the resonant frequency offset are generated, and the uncorrected frequency deviation data is obtained.
[0102] Furthermore, a temperature coefficient compensation algorithm (parameters: acoustic material temperature coefficient k_T, current cavity temperature T_cur, calibration temperature T_ref) is employed to correct the frequency offset by temperature and generate a corrected, standardized frequency deviation value. The specific compensation calculation formula is as follows: in, This is due to uncorrected frequency deviation. The temperature coefficient of the material. The current test temperature. This is for calibrating the reference temperature.
[0103] By employing frequency matching, bandwidth limiting, peak interpolation, difference calculation, and temperature compensation processing, the complex time-frequency matrix from the previous step is transformed into a temperature-corrected standardized frequency deviation value, thereby achieving precise quantification of the resonant peak offset characteristics.
[0104] For example, during the self-test process of a speaker of model SPK-800, the pre-stored nominal resonant frequency parameter is... Hz, the nominal parameter library directly returns this value through the model matching method. The amplitude spectrum of the complex time-frequency matrix is... Bandwidth limiting was performed within the Hz band to obtain resonant region data for 20 frequency points. A quadratic interpolation peak search algorithm was then used to find the point of maximum amplitude within this region at an interpolation interval of 0.1Hz, corresponding to the frequency [frequency missing]. Hz. Difference calculation method outputs uncorrected frequency deviation. Hz. The temperature sensor measures the current cavity temperature. ℃, the calibration reference temperature is ℃, material temperature coefficient Hz / ℃. The correction value is calculated using the temperature coefficient compensation formula. = Hz. Standardized frequency deviation value Hz is directly used as the resonance peak offset feature input of the vibration fingerprint encoder, which effectively improves the consistency and comparability of detection results under different temperature environments.
[0105] S4.1: Obtain the nominal resonant frequency f0 based on the pre-stored nominal parameter library matching the speaker model, and perform an adaptive peak search algorithm on the amplitude spectrum of the complex time-frequency matrix to locate the actual resonant frequency f_res; calculate the difference between the actual resonant frequency f_res and the nominal resonant frequency f0 to generate the resonant peak offset Δf_res characteristic parameter.
[0106] S4.2: Perform left and right integration region division processing on the amplitude spectrum of the complex time-frequency matrix in the neighborhood of the nominal resonant frequency f0, and calculate the envelope asymmetry α characteristic parameter based on the ratio of the left and right integration values; perform normalization processing on the envelope asymmetry α characteristic parameter to eliminate the influence of dimensions and output the normalized envelope asymmetry characteristic value.
[0107] S4.2: Based on the nominal resonant frequency parameter, the amplitude spectrum of the complex time-frequency matrix is divided into left and right integration regions to define the envelope asymmetry calculation domain; energy integration is performed on the left integration region to obtain the left energy value; energy integration is performed on the right integration region to obtain the right energy value; the ratio of the left energy value to the right energy value is calculated to generate the envelope asymmetry characteristic parameter; dynamic range compression is performed on the envelope asymmetry characteristic parameter to output the standardized envelope asymmetry value.
[0108] Based on the nominal resonant frequency parameter, the left and right integral region partitioning algorithm (parameters: nominal resonant frequency f0, integral width Δf) is used to define the computational domain for spectral symmetry analysis of the amplitude spectrum of the complex time-frequency matrix.
[0109] Furthermore, the energy accumulation calculation of the left integration region is realized through the energy integration algorithm (parameter: left frequency interval [f0-Δf, f0]), and the left energy value E_left is obtained.
[0110] Furthermore, the energy accumulation calculation of the right integration region is realized through the energy integration algorithm (parameter: right frequency interval [f0, f0+Δf]), and the right energy value E_right is obtained.
[0111] Furthermore, the energy ratio operation (parameters: E_left, E_right) is used to calculate the envelope asymmetry characteristic parameter B_env, as shown in the following formula: in, The integral energy value on the left side. This represents the integral energy value on the right side.
[0112] Furthermore, the range of B_env feature parameters is adjusted by using a dynamic range compression algorithm (parameters: compression factor k, threshold range [0.1, 10]), so that it maintains numerical robustness under high dynamic conditions and generates a standardized envelope asymmetry value B_env_norm.
[0113] Through the above computation chain, the energy distribution characteristics of the left and right resonant sides are transformed into a standardized envelope asymmetry index, thereby quantifying the physical correlation between the spectral envelope shape and the speaker structural anomaly.
[0114] For example, in a test scenario for a midrange speaker with a nominal resonant frequency f0 = 320Hz, Δf is set to 40Hz. The left integration interval [f0 - Δf, f0] corresponds to the frequency range of 280Hz to 320Hz, and the right integration interval [f0, f0 + Δf] corresponds to the frequency range of 320Hz to 360Hz. Performing energy integration on the amplitude spectrum yields the left energy value E_left = 2.35 × 10⁻ 3 The energy value on the right side, E_right, is 1.88 × 10⁻ 3 Substitute the results into the formula to calculate the envelope asymmetry characteristic parameters: The calculated value is B≈1.25. Dynamic range compression was performed using a compression factor k=0.8, adjusting B_env=1.25 to B_env_norm=1.12. The standardized envelope asymmetry value showed significantly improved stability in multiple batch tests, clearly distinguishing between left-side energy decay anomalies caused by frame bonding failure and right-side energy decay anomalies caused by dust cap detachment, achieving a highly reliable anomaly type indication.
[0115] S4.3: Perform linear regression fitting on the amplitude spectrum of the frequency band below 8kHz in the complex time-frequency matrix, calculate the slope of the fitted line to generate the high-frequency attenuation slope β characteristic parameter; perform logarithmic transformation on the high-frequency attenuation slope β characteristic parameter to output a physically meaningful quantitative index of attenuation rate.
[0116] S4.3: Perform frequency band truncation below 8kHz on the amplitude spectrum of the complex time-frequency matrix to extract the high-frequency attenuation analysis interval; perform logarithmic transformation on the truncated amplitude spectrum to generate dB domain spectrum data; perform linear regression fitting on the dB domain spectrum to calculate the attenuation slope; perform unit standardization on the attenuation slope to generate high-frequency attenuation slope characteristic parameters; output the high-frequency attenuation slope characteristic parameters as a quantitative index of spectral attenuation characteristics.
[0117] S4.4: Perform odd harmonic energy extraction processing on the amplitude spectrum of the fundamental frequency f0 and its 3rd, 5th and 7th harmonic frequencies in the complex time-frequency matrix. Based on the ratio of the sum of odd harmonic energy to the fundamental frequency energy, generate the odd harmonic energy ratio γ characteristic parameter. Perform threshold verification processing on the odd harmonic energy ratio γ characteristic parameter to output the harmonic distortion quantization index.
[0118] S4.4: Perform fundamental frequency localization processing on the amplitude spectrum of the complex time-frequency matrix to obtain the fundamental frequency energy value; perform 3x, 5x, and 7x frequency point extraction processing based on the fundamental frequency localization result to determine the odd harmonic frequency point set; perform energy summation processing on the odd harmonic frequency point set to obtain the total odd harmonic energy; perform ratio calculation processing based on the total odd harmonic energy and the fundamental frequency energy to generate the odd harmonic energy proportion characteristic parameter; perform threshold verification processing on the odd harmonic energy proportion characteristic parameter to output the harmonic distortion quantification index.
[0119] For the amplitude spectrum data of complex time-frequency matrices, a fundamental frequency localization method (parameters: frequency search range 20Hz to 200Hz, peak search accuracy 0.1Hz) is adopted to identify the position of the main peak in the amplitude spectrum, determine the fundamental frequency f0 in a physical sense, and obtain the fundamental frequency energy value E0 based on it.
[0120] Furthermore, by using the odd harmonic frequency extraction method (parameters: target harmonic coefficient set {3,5,7}, frequency tolerance ±0.2Hz), the integer harmonic index based on f0 is calculated, and the frequency points at 3f0, 5f0, and 7f0 in the amplitude spectrum are located to obtain the odd harmonic frequency set H={h1,h2,h3}.
[0121] Furthermore, an energy summation algorithm (parameters: weighting factor w=1, energy unit based on amplitude squared) is used to perform a square transformation on the amplitude of each frequency point in set H and then sum them up to obtain the total energy E_h of the odd harmonics: ,in This refers to the amplitude.
[0122] Furthermore, using an energy proportion calculation method, the ratio of the total energy E_h of odd harmonics to the fundamental frequency energy E0 is calculated to obtain the characteristic parameter R_h of the energy proportion of odd harmonics: .in This represents the total energy of odd harmonics.
[0123] Furthermore, a threshold verification method (parameter: preset threshold T_h=0.15) is adopted to perform a logical comparison between R_h and T_h, generate a harmonic distortion quantization index D_h, and output an abnormal flag when R_h>T_h, otherwise output a normal flag.
[0124] The algorithm described above transforms the fundamental frequency and harmonic amplitude data from the previous step into odd harmonic energy proportion characteristic parameters and their anomaly identification labels, enabling interpretable quantitative analysis of the nonlinear distortion components of the loudspeaker.
[0125] For example, in a loudspeaker testing scenario with a rated fundamental frequency of 80Hz, the amplitude spectrum is located at f0=80.1Hz, E0=0.85 (amplitude squared units); the amplitude squared values at 3f0=240.3Hz, 5f0=400.5Hz, and 7f0=560.7Hz are calculated to be 0.06, 0.04, and 0.03 respectively. After summing the energy, E_h=0.13 is obtained; R_h is calculated using the above formula: ≈0.153. Based on a threshold T_h=0.15, the verification showed R_h slightly higher than the threshold, generating a distortion quantification index D_h=abnormal, indicating that the loudspeaker exhibits detectable odd-order harmonic distortion. In production line verification, this result was consistent with manually detected voice coil adhesion defects, validating the diagnostic effectiveness and interpretability of this characteristic parameter.
[0126] S4.5: Perform -40dB energy threshold determination processing on the envelope signal of the complex time-frequency matrix, and generate the transient response tail length τ feature parameter based on the counting operation of the number of frames required for the envelope to fall to the threshold; perform first-order difference calculation processing on the phase spectrum of the complex time-frequency matrix, and generate the number of phase change points η feature parameter based on the statistical operation of the absolute value of the difference exceeding the threshold; merge the transient response tail length τ and the number of phase change points η feature parameters into a time-domain-phase composite feature vector.
[0127] S4.5: Perform -40dB energy threshold determination processing on the envelope signal of the complex time-frequency matrix to identify the tail start point; perform tail length counting processing based on the tail start point to generate transient response tail length feature parameters; perform first-order difference calculation processing on the phase spectrum of the complex time-frequency matrix to obtain the phase change rate sequence; perform threshold comparison processing based on the phase change rate sequence to filter out points exceeding the threshold; perform statistical processing on the number of points exceeding the threshold to generate phase change point quantity feature parameters; output the transient response tail length and phase change point quantity feature parameters as a time-domain-phase composite feature vector.
[0128] The energy threshold determination method (parameter: threshold - 40dB) is used to identify the tail start point of the envelope signal of the input complex time-frequency matrix.
[0129] Furthermore, the tail length characteristic parameter calculation function is realized by using the tail length counting method (parameters: starting point position index, amplitude attenuation threshold -40dB), and the transient response time length data is obtained.
[0130] Furthermore, a first-order difference calculation method (parameter: full-band phase spectrum data) is adopted to obtain the phase change rate sequence and generate a phase change rate vector for each frequency point.
[0131] Furthermore, a threshold comparison method (parameter: phase change rate threshold π / 4) is used to implement the filtering function of points exceeding the threshold, and to obtain the set of points exceeding the threshold and their frequency position index data.
[0132] Furthermore, a point count method (parameter: set of points exceeding the threshold) is adopted to generate the characteristic parameter of the number of phase change points and output the integrated count index of phase change events.
[0133] By constructing a time-phase composite feature vector, the tail length and the number of phase abrupt change points obtained in the previous step are transformed into data indicators that jointly represent the time and phase, thereby achieving a comprehensive quantitative effect on the transient characteristics and phase stability of the vibration.
[0134] For example, after performing a short-time Fourier transform on the speaker feedback signal, the corresponding envelope signal is extracted and the root mean square energy is calculated. The energy value for each time frame is converted to a dB scale, and a threshold of -40 dB is set. The time point at which the energy curve first falls below this threshold is taken as the tail start point. Assuming the tail start point is at frame 40, with a sampling rate of 16kHz and a frame length of 1024 points, the tail length is calculated to be (8 frames × 64 points) / 16kHz ≈ 32ms. A first-order difference formula is used. in, For phase spectrum, For frequency indexing, compare the phase change rate across the entire frequency band with a set threshold. The frequency points that meet the criteria are compared and selected, and their number is counted. For example, the number of phase change points is 5. The tail length of 32ms and the number of phase change points of 5 are combined into a feature vector [32ms,5] and input into the subsequent vibration fingerprint encoding module. It has been verified that this feature vector is highly correlated with the typical spectral shape of speaker edge ring aging, which can significantly improve the accuracy of early fault identification.
[0135] Step S5: Based on the pre-stored mapping relationship between loudspeaker structure-material-assembly defects and spectral morphology, a three-layer fully connected vibration fingerprint encoder with fixed weights is constructed; the six-dimensional interpretable feature vector is input into the encoder to obtain a 32-dimensional vibration fingerprint code, where each dimension corresponds to the activation intensity of a typical assembly defect pattern. Specifically, this includes: S5.1: Based on the pre-stored dictionary of speaker structure-material-assembly defects and spectral morphology mapping, obtain the construction instructions of a three-layer fully connected vibration fingerprint encoder containing fixed weight parameters; use the fixed-point arithmetic unit of the Cortex-M4 microcontroller to perform the weight matrix initialization operation to generate a neural network topology with 128 nodes in the input layer, 64 nodes in the hidden layer, and 32 nodes in the output layer. The weight parameters are fixedly stored in read-only memory according to historical calibration data to ensure that the encoder has no training process and meets the real-time requirements of the edge side.
[0136] Based on a pre-stored dictionary mapping loudspeaker structure-material-assembly defects to spectral morphology, a dictionary retrieval method (parameters: physical defect type index, spectral feature pattern index) is used to realize the fixed mapping relationship between defect patterns and corresponding spectral morphologies.
[0137] Furthermore, by using a structured instruction parsing algorithm (parameters: number of nodes in a three-layer fully connected network, weight initialization rules), the construction instructions of the vibration fingerprint encoder containing fixed weight parameters are parsed, and a network structure definition dataset is obtained.
[0138] Furthermore, the fixed-point arithmetic unit of the Cortex-M4 microcontroller is used to perform the weight matrix initialization operation (parameters: fixed weight storage address, Q15 fixed-point format) to realize the input layer weight matrix. Hidden layer weight matrix and output layer weight matrix The data is loaded and a fixed-point weight form adapted to the computing power of the microcontroller is obtained.
[0139] Furthermore, a network topology generation algorithm (parameters: 128 nodes in the input layer, 64 nodes in the hidden layer, and 32 nodes in the output layer) is adopted to instantiate the node and connection structure of the three-layer fully connected vibration fingerprint encoder and generate a network topology description file containing a node index table and a connection mapping table.
[0140] Furthermore, by utilizing the permanent storage function of read-only memory (parameters: historical calibration dataset, weight calibration coefficients), the weight parameters are permanently stored, ensuring that the encoder has no training process during operation, thus meeting the real-time and low-power requirements of the edge side.
[0141] Through the above-mentioned weight matrix initialization and network topology construction processing, the mapping dictionary retrieval results of the previous step are transformed into an executable vibration fingerprint encoder instance, realizing the expected technical effect of obtaining stable and physically interpretable defect mode activation intensity after inputting the six-dimensional interpretable feature vector into the encoder.
[0142] For example, in a specific loudspeaker quality testing scenario, the pre-stored mapping dictionary includes four defect patterns: magnetic circuit eccentricity, voice coil adhesion, frame looseness, and dust cap detachment. Each mapping corresponds to a six-dimensional feature template and weight correction coefficients. The three-layer fully connected structure has 128 nodes in the input layer, 64 nodes in the hidden layer, and 32 nodes in the output layer. The weight parameters are stored in ROM address 0x08020000 using Q15 fixed-point format. The fixed-point arithmetic unit will... A 128×6 matrix The 64×128 matrix and The 32×64 matrix is loaded into the on-chip SRAM, and weight calibration coefficients are adjusted. These calibration coefficients are calculated from historical frequency sweep calibration data. When the six-dimensional feature vector [Δf_res=6Hz, α=0.45, β=-3.2dB / kHz, γ=0.12, τ=5 frames, η=3] is input to the encoder, the hardware-accelerated matrix multiplication unit completes forward propagation within 4.2ms. The activation intensity of the 5th dimension (magnetic circuit eccentricity mode) in the output 32-dimensional fingerprint code is 0.82, and the activation intensity of the 11th dimension (frame looseness mode) is 0.76, significantly higher than the values of other dimensions. This result generates a structured diagnostic vector after the subsequent S5.2 input mapping rules, directly indicating that the speaker has physical defects of both magnetic circuit eccentricity and frame looseness. The entire computation chain has a memory footprint of 112KB under Cortex-M4@120MHz conditions, meeting the requirements for real-time self-testing on the embedded end side.
[0143] S5.2: Based on the vibration fingerprint encoder topology constructed in S5.1, the six-dimensional interpretable feature vector output from step S4 is obtained as input data; the input layer feature mapping process is performed using a hardware-accelerated matrix multiplication unit to convert the six-dimensional interpretable feature vector into a 128-dimensional high-order feature space representation, which retains the physical correlation of the original features such as formant offset and envelope asymmetry.
[0144] S5.3: Based on the 128-dimensional high-order feature space representation generated by S5.2, perform nonlinear activation processing on the hidden layer; use the ReLU activation function with pre-fixed threshold to perform element-wise operation on the 64-dimensional hidden layer nodes to extract the coupling feature relationship between envelope asymmetry and high-frequency attenuation slope, and generate a 64-dimensional intermediate feature vector, which represents the implicit correlation mode of loudspeaker structural defects.
[0145] S5.4: Based on the 64-dimensional intermediate feature vector output by S5.3, perform linear transformation processing of the output layer; use the solidified weight matrix to perform 32-dimensional vector projection operation to generate a 32-dimensional acoustic fingerprint code, where each dimension corresponds to the activation intensity value of typical assembly defect modes such as magnetic circuit eccentricity and voice coil adhesion, and the activation intensity value is limited to the range of 0 to 1.
[0146] S5.5: Based on the 32-dimensional acoustic fingerprint code generated by S5.4, perform defect mode activation intensity standardization processing; according to the preset physical meaning mapping rules, perform threshold quantization operation on the values of each dimension to output a structured diagnostic vector that represents the activation intensity distribution of typical assembly defect modes. This vector is directly related to the physical causes of defects such as magnetic circuit eccentricity and frame loosening and supports the subsequent generation of interpretable thermal labels.
[0147] Based on the 32-dimensional acoustic fingerprint generated by S5.4, a standardized mapping method (parameter: physical meaning mapping rule set) is used to achieve a unified measurement of the activation intensity of defect modes. Furthermore, a threshold quantization algorithm (parameter: upper and lower limits of activation intensity corresponding to each defect mode) is used to convert continuous activation values into discrete quantized labels, and the defect existence determination results in each dimension are obtained.
[0148] Furthermore, a logical decision function (parameter: threshold judgment conditions for each dimension) is used to perform Booleanization processing on the quantized labels and generate a binary diagnostic label vector, where each element corresponds to the existence status identifier of a typical assembly defect mode. Further, a structured mapping algorithm (parameter: defect mode physical cause dictionary) is used to convert the binary diagnostic label vector into a structured diagnostic vector, generating diagnostic data directly associated with typical defect modes such as magnetic circuit eccentricity, frame loosening, and voice coil adhesion. Further, a causal chain association method (parameter: defect mode-physical cause mapping linked list) is used to bind the interface between the diagnostic data and the spectrum visualization thermal label generation module, obtaining a standardized diagnostic vector that satisfies the requirements for subsequent interpretable label generation.
[0149] By standardizing the activation intensity of the defect pattern and combining it with the physical meaning mapping rule, the results of step S5.4 are transformed into a structured diagnostic vector with direct verifiability, thus achieving a closed-loop connection between the defect pattern and the intermediate evidence required for the generation of interpretable thermal labels.
[0150] For example, in a speaker vibration self-test, the 32-dimensional vibration fingerprint code output by step S5.4 is [0.83, 0.12, 0.56, 0.91, ...], where the first dimension corresponds to magnetic circuit eccentricity and the fourth dimension corresponds to frame looseness. The standardized mapping method employs a maximum value normalization strategy, uniformly measuring the value of each dimension according to a preset physical meaning mapping rule. For example, the maximum allowed activation value for the first dimension is set to... The normalized first dimension value is still Threshold quantization algorithms will select activation values greater than or equal to... The dimension is marked as "existent", and the corresponding binarized label is... The rest are marked as The logical decision function generates a binary diagnostic identifier vector [1,0,0,1,......] according to this rule, indicating that magnetic circuit eccentricity and frame looseness coexist. The structured mapping algorithm calls the defect pattern physical cause dictionary, mapping the first dimension "magnetic circuit eccentricity" to "abnormal magnet assembly gap" and the fourth dimension "frame looseness" to "outer frame bonding failure", generating a structured diagnostic vector [(magnetic circuit eccentricity, abnormal magnet assembly gap), (frame looseness, outer frame bonding failure)]. The causal chain association method binds this structured diagnostic vector to the interpretable thermal label generation module interface, enabling the thermal label to highlight the double-peak split region at 320Hz (frame looseness feature) and the resonance abnormality spike region at 480Hz (magnetic circuit eccentricity feature) in the spectrum. In this embodiment, the final output structured diagnostic vector directly drives the interpretable thermal label module to complete the visualization highlighting of the local spectrum, significantly improving the transparency and reliability of fault determination.
[0151] Step S6: The seismic fingerprint sequence obtained from 10 consecutive self-tests is processed to construct an evolutionary trajectory. A seismic evolutionary trajectory map is generated through two-dimensional planar projection and color depth mapping to characterize the temporal evolution of seismic features. Specifically, this includes: S6.1: Based on the 32-dimensional acoustic fingerprint code sequence output from the previous step S5, obtain the acoustic fingerprint code sequence dataset of 10 consecutive self-checks to provide basic input data for time series analysis.
[0152] S6.2: Perform feature extraction processing on the vibration fingerprint code sequence dataset, extract the first two-dimensional features of each vibration fingerprint code, and generate a two-dimensional feature vector sequence to determine the coordinate basis of the plane projection.
[0153] Based on the dataset of 10 consecutive self-tested acoustic fingerprint codes output from step S6.1, a dimension index filtering method (parameters: sequence length M=10, target dimension index D1=1, D2=2) is used to separate the first two dimensions of each acoustic fingerprint code, ensuring that the input is a structured set of two-dimensional components.
[0154] Furthermore, by using a vectorized mapping method (parameters: component precision is floating-point 32-bit, range normalized to the [0,1] interval), the numerical normalization of the first two dimensions of features is achieved, and a normalized two-dimensional feature data matrix is obtained to eliminate the amplitude scale differences between different self-inspection batches.
[0155] Furthermore, by using a matrix reconstruction method (parameters: row index corresponds to self-inspection batch, column index corresponds to feature component), the structure of the two-dimensional feature data matrix is rearranged to generate a two-dimensional feature vector sequence arranged in chronological order, while retaining the batch index association information with the original sequence.
[0156] Furthermore, a feature consistency detection algorithm (parameter: Euclidean distance tolerance threshold ε=0.005) is adopted to monitor the stability of the two-dimensional feature vector sequence and filter out data points that do not meet the consistency requirements, so as to ensure the smoothness of the trajectory in the subsequent planar projection calculation.
[0157] Through the above feature extraction and processing methods, the 32-dimensional vibration fingerprint code sequence from the previous step is transformed into a two-dimensional feature vector sequence that can be directly used for planar projection calculation, thereby realizing the accurate construction of the planar trajectory space coordinate basis in the main step.
[0158] For example, in an actual production line end-side inspection scenario, the length of the vibration fingerprint code sequence for 10 consecutive self-inspections is M=10, and each fingerprint code is a 32-dimensional floating-point numerical vector. Selecting dimension indices 1 and 2 yields the original two-dimensional component set, with a numerical range of [0.1345, 0.9852]. A range normalization method is then applied. in, =0.9852, =0.1345. After normalization, the resulting two-dimensional feature matrix has a shape of 10×2, with row indices 0 to 9 corresponding to the batch order. The Euclidean distance consistency detection formula is applied to this matrix: in and These are the normalized differences between the first and second dimensions of adjacent batches of two-dimensional features. When d > 0.005, the batch is marked as unstable data and removed. After this processing, nine valid two-dimensional feature vectors are finally obtained, significantly improving the accuracy of the planar coordinates and meeting the technical requirements for data continuity and stability in subsequent trajectory construction.
[0159] S6.3: Perform two-dimensional plane projection processing based on the two-dimensional feature vector sequence, calculate the projection coordinates of each feature vector, and generate a trajectory point coordinate sequence to construct the spatial location framework of the trajectory.
[0160] S6.4: Perform trajectory connection processing on the coordinate sequence of the trajectory points to connect adjacent trajectory points to form a motion trajectory, so as to characterize the continuous evolution path of the vibration characteristics.
[0161] S6.5: Based on the motion trajectory and the third-dimensional intensity value of each vibration fingerprint code, perform color depth mapping processing to map the third-dimensional intensity into a color depth indicator, and generate a vibration evolution trajectory map to visualize the temporal evolution law of vibration features.
[0162] Based on the generated sequence of motion trajectory point coordinates and their corresponding third-dimensional activation intensity values, a color mapping algorithm (parameters: color space mode is HSV, intensity range is set to 0 to 1) is used to convert the third-dimensional activation intensity values into color depth values for trajectory visualization rendering. Furthermore, a normalization algorithm (parameters: minimum value mapped to V=0.3, maximum value mapped to V=1.0) is used to compress the dynamic range of the intensity values in the color space, resulting in a normalized color depth dataset.
[0163] Furthermore, by using a color index generation algorithm (parameters: Hue is fixed at 220° blue, and Saturation is fixed at 0.9), the normalized color depth dataset is combined with the trajectory point coordinate sequence to generate a trajectory point structure array with color attributes, and a color encoding vector is generated for map drawing.
[0164] Furthermore, through a trajectory rendering algorithm (parameters: anti-aliasing mode enabled, line width set to 2px), trajectory points with color attributes are connected according to the motion trajectory. Simultaneously, color gradient rendering rules are applied to ensure a smooth and continuous transition of color depth with changes in the third dimension intensity. The following mapping formula is used for color depth calculation: in, This represents the third-dimensional activation intensity value of the current trajectory point. This represents the minimum activation strength. The maximum activation strength, This is the luminance component used for color mapping after normalization.
[0165] Furthermore, by using a map generation engine (parameters: output resolution of 800×600 pixels, background transparency set to 0.0), the drawn color depth trajectory is superimposed onto a two-dimensional plane coordinate frame to generate a complete sound evolution trajectory map file (format: PNG or BMP) to visualize the temporal evolution of sound characteristics.
[0166] By using color mapping and trajectory rendering, the motion trajectory and third-dimensional intensity result from the previous step are transformed into color depth visualization data with intensity information, thereby enabling intuitive representation and trend analysis visualization of the evolution process of vibration characteristics.
[0167] For example, in 10 consecutive self-inspection data sets of a certain loudspeaker production batch, the activation intensity value of the third dimension of the trajectory point ranges from 0.25 to 0.85. The color mapping algorithm described above is used to set... =0.25、 =0.85, after normalization, the luminance component is obtained. The range is 0.0 to 1.0. The color space parameter Hue is fixed at 220° corresponding to blue, and the saturation is set to 0.9. The intermediate intensity value of 0.55 is converted into a luminance component of 0.5 through this mapping formula, corresponding to the medium brightness of visual blue. The trajectory rendering adopts an anti-aliasing mode and gradient lines, with a smooth color transition at each trajectory point, finally generating an 800×600 pixel vibration evolution trajectory map file, which visually presents a continuous trend of darker colors in areas of lower intensity and brighter colors in areas of higher intensity. When this map is called by the production line quality inspection system, the trajectory moves from the lower left corner to the upper right corner and the color gradually brightens, verifying the gradual deterioration trend of a certain structural defect in the speaker during use, which helps to implement maintenance or replacement strategies in advance.
[0168] Step S7: Based on the locally stored fault-fingerprint-handling suggestion triplet knowledge base, the endpoint vibration fingerprint code of the vibration evolution trajectory spectrum is binarized according to a preset threshold, and matched with the knowledge base to generate interpretable thermal tags and natural language judgment conclusions, outputting fault diagnosis evidence containing locally highlighted areas of the spectrum. Specifically, this includes: S7.1: Perform preset threshold binarization processing on the endpoint vibration fingerprint code of the vibration evolution trajectory map, and make logical judgments on each dimension of the 32-dimensional vibration fingerprint code based on the solidified threshold matrix. Convert the continuous activation intensity value into a binary feature vector to generate discrete feature identifiers that characterize the existence of typical assembly defect modes.
[0169] S7.2: Based on the binary feature vector, perform a precise matching operation of the fault-fingerprint-handling suggestion triplet knowledge base. According to the pre-stored logical rule set, compare the binary feature vector with the fault feature template in the knowledge base to locate the set of fault entries with the highest matching degree, so as to obtain the fault type identifier and associated handling suggestion code of the corresponding assembly defect mode.
[0170] The input object is the binarized feature vector output from step S7.1, and the execution object is the locally stored fault-fingerprint-disposal suggestion triplet knowledge base.
[0171] A precise matching retrieval method is adopted (parameters: the matching rule set is derived from the number of historical calibration samples ≥ 500, and the matching threshold is set to be completely consistent) to realize the one-to-one logical comparison function between the binary feature vector and each fault feature template in the knowledge base.
[0172] Furthermore, by using a logical pattern matching algorithm (parameters: supporting AND / OR / XOR three types of logical operations, and the bit order is fixed according to the vibration fingerprint code dimension), the binary feature vector and the knowledge base template are matched by Boolean operations in the bitmap space, and the matching Boolean matrix data results are obtained.
[0173] Furthermore, by using a matching degree calculation method (parameter: the matching degree formula is based on Hamming distance normalization), the matching degree of the Boolean matrix data is quantified, and a matching degree value vector is generated; the matching degree calculation formula is: in, The total number of bits in the feature code. The Hamming distance is the distance between the binarized feature vector and the template.
[0174] Furthermore, the matching degree value vector is sorted using a matching degree value sorting method (parameter: descending order, take the first N=3 items), and a set of fault entries with the highest matching degree is generated.
[0175] By using a fault entry set index retrieval method, the sorting results of the previous step are mapped to the knowledge base storage space, and the fault type identifiers and associated handling suggestion codes of the corresponding assembly defect patterns are extracted to realize the physical association confirmation of the diagnostic target.
[0176] For example, in a two-way speaker system with a rated power of 50W, the binarized feature vector is 32 bits long, with the 3rd, 7th, and 15th bits being active. Using the Hamming distance matching algorithm, the corresponding bits of the knowledge base template T1 are activated at the 3rd and 7th bits, and deactivated at the 15th bit. The Hamming distance is then calculated. =1, matching degree is =0.96875, ranking first in descending order of matching degree. Mapping to the knowledge base yields the fault type identifier "bath stand loose" and the treatment suggestion code "CHK-FRM-GLUE". This result is used in subsequent steps to generate spectral highlight regions and output natural language judgments, verifying that the abnormal splitting of the mid-frequency resonance peak is highly consistent with the physical cause of the basin stand bonding failure, and the matching effect significantly improves the reliability of the fault judgment.
[0177] S7.3: Perform interpretable thermal label generation processing on the matched fault type identifier, extract the coordinate parameters of key spectral regions based on the fault feature mapping dictionary, and call the spectral highlight rendering algorithm in combination with the fault type identifier to generate a two-dimensional thermal distribution layer covering the original spectral data, so as to visually present the fault-related local highlight areas of the spectrum.
[0178] S7.4: Based on the fault type identifier and handling suggestion code, the natural language judgment conclusion generation process is executed. The preset fault semantic template library is called for parameterized filling, the fault type identifier is converted into a structured diagnostic statement, and a natural language judgment conclusion text containing fault phenomenon description, physical cause explanation and maintenance guidance information is generated.
[0179] Based on the fault type identifier and handling suggestion code generated in the preceding step S7.3, the parameterized filling function of the diagnostic statement is realized by using the preset fault semantic template calling method (parameters: fault type identifier, handling suggestion code).
[0180] Furthermore, through a fault characteristic mapping algorithm (parameters: fault type identifier, physical cause library index), the structural defect patterns corresponding to the fault type identifier are matched one by one with the pre-stored physical cause descriptions, and structured physical cause description text data is obtained.
[0181] Furthermore, through the maintenance guidance information combination algorithm (parameters: handling suggestion code, guidance template library index), the handling suggestion code is associated with the preset maintenance guidance template by index, and a detailed text result of maintenance operation steps is generated.
[0182] Furthermore, a multi-segment diagnostic statement splicing processing method is adopted (parameters: diagnostic phenomenon description text, physical cause description text, and maintenance guidance text) to achieve sequential splicing of three independent diagnostic elements and generate a composite text containing information on fault phenomena, physical causes, and maintenance guidance.
[0183] By using a natural language structured rendering algorithm, the text results from the previous step are transformed into standardized diagnostic conclusion data, achieving semantically consistent and user-understandable natural language judgment output.
[0184] For example, for the input fault type identifier "magnetic circuit eccentricity" and the handling suggestion code "inspect magnet assembly gap", the preset fault semantic template T15 is invoked. The template content is "[fault phenomenon] detected, which is caused by [physical cause], and [repair guidance] is recommended". Through the fault characteristic mapping algorithm, "magnetic circuit eccentricity" is matched to the physical cause library entry "magnet position displacement causes magnetic field inhomogeneity", and the corresponding physical cause description is generated. "Inspect magnet assembly gap" is matched to the operation step "disassemble speaker housing, check the gap between magnet and voice coil, use a precision positioning fixture to re-center, reset and fix" through the guidance template library. The phenomenon description "double peak splitting occurs in a specific frequency band of the speaker" is concatenated with the above physical cause description and repair guidance to generate a diagnostic statement, and finally a natural language judgment conclusion is generated: "Double peak splitting occurs in a specific frequency band of the speaker. This phenomenon is caused by magnetic field inhomogeneity due to magnet position displacement. It is recommended to disassemble speaker housing, check the gap between magnet and voice coil, use a precision positioning fixture to re-center, reset and fix". In this scenario, diagnostic results can be directly displayed on the audio terminal screen and simultaneously saved in the data backtracking package for production line quality traceability, significantly improving the interpretability and user trust of frequency domain analysis results.
[0185] S7.5: Perform fault diagnosis evidence synthesis processing on interpretable thermal labels and natural language judgment conclusions, integrate thermal distribution layers and diagnostic conclusion text to form a multimodal diagnostic evidence package, and use coordinate alignment algorithm to accurately associate the spectral highlight areas with the fault descriptions in the conclusion text to output complete fault diagnosis evidence with spectral local highlight area annotations.
[0186] After step S7, the method further includes: S8: Perform low-bandwidth compression processing on the interpretable feature vector, vibration fingerprint code, vibration evolution trajectory map and interpretable thermal tag to generate a data backtracking packet that conforms to the Bluetooth BLE protocol specification.
[0187] Step S8 specifically includes: S8.1: Based on the dynamic range characteristics of the six-dimensional interpretable feature vector, a quantization compression table is constructed; based on the binarization distribution characteristics of the acoustic fingerprint code, a bitmap compression mapping is generated; combining the above, a low-bandwidth compression algorithm is constructed to minimize the feature data volume.
[0188] S8.2: Perform set operations on the six-dimensional interpretable feature vector, sonic fingerprint code, sonic evolution trajectory map, and interpretable thermal labels to generate the original dataset, which serves as the input object for the low-bandwidth compression algorithm.
[0189] S8.3: Apply a low-bandwidth compression algorithm to the original dataset, perform quantization and bitmap encoding operations, and generate a compressed data stream to reduce data transmission bandwidth requirements.
[0190] S8.4: Perform Bluetooth BLE protocol encapsulation processing on the compressed data stream, add service UUID and feature descriptor, generate data backtracking packets that conform to the specifications, and ensure protocol compatibility with the production line quality traceability system.
[0191] S8.5: Perform CRC check processing on the data traceability packet to verify data integrity; output the verified data traceability packet to the Bluetooth module to support the production line quality traceability system to access and verify it.
[0192] For those skilled in the art, various other corresponding changes and modifications can be made based on the technical solutions and concepts described above, and all such changes and modifications should fall within the protection scope of the claims of this invention.
[0193] Unless otherwise defined, the technical or scientific terms used herein shall have the ordinary meaning as understood by one of ordinary skill in the art to which this application pertains. The terms “first,” “second,” “third,” and similar terms used in this patent application specification and claims do not indicate any order, quantity, or importance, but are merely used to distinguish different components. Similarly, the terms “an” or “a” and similar terms do not indicate a quantity limitation, but rather indicate the presence of at least one. The terms “comprising” or “including” and similar terms mean that the elements or objects preceding “comprising” or “including” encompass the elements or objects listed following “comprising” or “including” and their equivalents, and do not exclude other elements or objects. The “multiple” mentioned in the embodiments of this application refers to two or more. A and / or B indicate three possibilities: A; B; and A and B.
[0194] The above description is merely an exemplary embodiment of this application, but the scope of protection of this application is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in this application, and such modifications or substitutions should all be covered within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
Claims
1. A closed-loop self-testing method for loudspeaker vibration based on built-in sweep frequency signal self-excitation, specifically including: S1: Acquire the feedback signal of the loudspeaker after being excited by the frequency sweep, and perform anti-aliasing filtering to obtain a time-domain signal sequence with a sampling rate of a preset frequency; S2: Extract stable response segment data from the time-domain signal sequence, extract valid data from the stable response segment data, and form a stable response segment time-domain signal; S3: Perform piecewise windowed short-time Fourier transform processing on the time-domain signal of the stable response segment to generate a complex time-frequency matrix covering a preset frequency band; S4: Determine the interpretable feature vector based on the complex time-frequency matrix; S5: Construct a fully connected vibration fingerprint encoder based on the pre-stored mapping relationship between speaker structure-material-assembly defects and spectral morphology; input the interpretable feature vector into the fully connected vibration fingerprint encoder to obtain the vibration fingerprint code; S6: Perform evolution trajectory construction processing on the vibration fingerprint code sequence obtained by self-inspection, and generate vibration evolution trajectory map through two-dimensional plane projection and color depth mapping; S7: The sound fingerprint code of the sound evolution trajectory map is binarized according to a preset threshold, and matched with the fault-fingerprint-handling suggestion triplet knowledge base to generate interpretable thermal labels and natural language judgment conclusions, and output fault diagnosis evidence.
2. The loudspeaker vibration closed-loop self-testing method based on built-in sweep frequency signal self-excitation according to claim 1, characterized in that, After step S7, the method further includes: S8: Perform low-bandwidth compression processing on the interpretable feature vector, vibration fingerprint code, vibration evolution trajectory map and interpretable thermal tag to generate a data backtracking packet that conforms to the Bluetooth BLE protocol specification.
3. The loudspeaker vibration closed-loop self-testing method based on built-in sweep frequency signal self-excitation according to claim 1, characterized in that, Step S1 specifically includes: S1.1: Detect and process the power-on self-test trigger event or periodic detection trigger event of the audio system operation status to generate a trigger confirmation signal; S1.2: Apply bias voltage to the built-in microphone based on the trigger confirmation signal; S1.3: The built-in microphone, after being processed by bias voltage application, performs acoustic-to-electric conversion on the acoustic feedback signal of the speaker after frequency sweep excitation to obtain an analog voltage signal; S1.4: Perform anti-aliasing low-pass filtering on the analog voltage signal to generate an anti-aliasing filtered signal; S1.5: After anti-aliasing filtering, the signal is processed by analog-to-digital conversion at a preset sampling rate to obtain a digital time-domain signal sequence.
4. The loudspeaker vibration closed-loop self-testing method based on built-in sweep frequency signal self-excitation according to claim 1, characterized in that, Step S2 specifically includes: S2.1: Construct a signal energy stability window determination rule based on the statistical analysis results of historical frequency sweep data, and output the parameter set of the signal energy stability window determination rule; S2.2: Perform sliding window energy calculation processing on the time-domain signal sequence to generate a signal energy sequence; S2.3: Based on the signal energy stability window determination rule, perform stability window positioning processing on the signal energy sequence to determine the starting point of the stable response segment; by detecting continuous intervals with energy variance less than 0.5dB and duration exceeding 100ms, output the starting point index of the stability window; S2.4: Based on the starting point index of the stable window, extract 4096 consecutive data points from the time-domain signal sequence to form a preliminary stable response segment time-domain signal, and output the preliminary stable response segment time-domain signal; S2.5: Perform energy stability verification processing on the time-domain signal of the preliminary stable response segment to generate the time-domain signal of the final stable response segment, and output the time-domain signal of the final stable response segment.
5. The loudspeaker vibration closed-loop self-testing method based on built-in sweep frequency signal self-excitation according to claim 1, characterized in that, In S2.2, a sliding window energy calculation process is performed on the time-domain signal sequence, specifically as follows: The root mean square energy value is calculated using a 50ms Hanning window sliding window with a 5ms step size, and the signal energy sequence composed of the energy of each frame is output.
6. The loudspeaker vibration closed-loop self-testing method based on built-in sweep frequency signal self-excitation according to claim 1, characterized in that, In S3, during the segmented windowed short-time Fourier transform process, a Hanning window is used and a preset overlap rate is set.
7. The loudspeaker vibration closed-loop self-testing method based on built-in sweep frequency signal self-excitation according to claim 1, characterized in that, Step S3 specifically includes: S3.1: Perform segmentation processing on the stable response segment time domain signal. Based on the preset segmentation length of 1024 points and the 50% overlap rate parameter, the 4096-point stable response segment time domain signal is divided into 8 overlapping data segments to generate a segmented time domain signal sequence. S3.2: Perform Hanning window function weighting processing on the segmented time-domain signal sequence, and perform amplitude modulation on each data segment based on the mathematical definition of Hanning window to generate a windowed segmented time-domain signal sequence; S3.3: Perform Fast Fourier Transform (FFT) processing on the windowed segmented time-domain signal sequence, calculate the frequency domain complex representation of each data segment based on the radix-2 FFT algorithm, and generate an initial complex spectrum matrix; S3.4: Perform frequency axis linear mapping processing on the initial complex spectrum matrix. Based on the sampling rate of 16kHz and 1024-point FFT parameters, convert the frequency index into physical frequency values and generate a calibration matrix covering the physical frequency. S3.5: Perform target frequency band truncation processing on the physical frequency calibration matrix. Based on the frequency range threshold of 20Hz to 5kHz, filter and retain the data in the corresponding frequency point index interval 1 to 320 to generate a complex time-frequency matrix of 512 frequency points covering the 20Hz to 5kHz frequency band.
8. The loudspeaker vibration closed-loop self-testing method based on built-in sweep frequency signal self-excitation according to claim 1, characterized in that, Step S4 specifically includes: S4.1: Perform loudspeaker model matching processing on the pre-stored nominal parameter library to obtain the nominal resonant frequency parameter; perform bandwidth limiting processing on the amplitude spectrum of the complex time-frequency matrix based on the nominal resonant frequency parameter; perform quadratic interpolation peak search algorithm processing on the limited amplitude spectrum; perform difference calculation processing on the actual resonant frequency and the nominal resonant frequency to generate the resonant frequency offset characteristic parameter; perform temperature coefficient compensation processing on the resonant frequency offset characteristic parameter. S4.2: Perform left and right integration region division processing on the amplitude spectrum of the complex time-frequency matrix based on the nominal resonant frequency parameter; perform ratio calculation processing based on the energy value on the left and the energy value on the right to generate the envelope asymmetry characteristic parameter; perform dynamic range compression processing on the envelope asymmetry characteristic parameter to output the standardized envelope asymmetry value; S4.3: Perform frequency band truncation below 8kHz on the amplitude spectrum of the complex time-frequency matrix; perform logarithmic transformation on the truncated amplitude spectrum to generate spectral data; perform linear regression fitting based on the spectrum to calculate the attenuation slope; perform unit standardization on the attenuation slope to generate high-frequency attenuation slope characteristic parameters. S4.4: Perform fundamental frequency localization processing on the amplitude spectrum of the complex time-frequency matrix to obtain the fundamental frequency energy value; perform frequency point extraction processing based on the fundamental frequency localization result to determine the set of odd harmonic frequency points; perform energy summation processing on the set of odd harmonic frequency points to obtain the total energy of odd harmonics; perform ratio calculation processing based on the total energy of odd harmonics and the fundamental frequency energy to generate the characteristic parameter of odd harmonic energy proportion; perform threshold verification processing on the characteristic parameter of odd harmonic energy proportion to output the harmonic distortion quantification index. S4.5: Perform a preset energy threshold determination process on the envelope signal of the complex time-frequency matrix to identify the tail start point; perform tail length counting processing based on the tail start point to generate transient response tail length feature parameters; perform first-order difference calculation processing on the phase spectrum of the complex time-frequency matrix to obtain the phase change rate sequence; perform threshold comparison processing based on the phase change rate sequence to filter out points exceeding the threshold; perform statistical processing on the number of points exceeding the threshold to generate phase change point number feature parameters; output transient response tail length and phase change point number feature parameters.
9. The loudspeaker vibration closed-loop self-testing method based on built-in sweep frequency signal self-excitation according to claim 1, characterized in that, Step S5 specifically includes: S5.1: Based on the pre-stored dictionary of loudspeaker structure-material-assembly defects and spectral morphology mapping, obtain the construction instructions for the fully connected vibration fingerprint encoder; perform weight matrix initialization to generate the neural network topology. S5.2: Based on the vibration fingerprint encoder topology, obtain interpretable feature vectors; use hardware-accelerated matrix multiplication units to perform input layer feature mapping processing to convert the interpretable feature vectors into a 128-dimensional high-order feature space representation. S5.3: Based on the 128-dimensional high-order feature space representation, perform nonlinear activation processing on the hidden layer; use the ReLU activation function with a pre-fixed threshold to perform element-wise operations on the 64-dimensional hidden layer nodes to extract the coupling feature relationship between envelope asymmetry and high-frequency attenuation slope, and generate a 64-dimensional intermediate feature vector. S5.4: Based on the 64-dimensional intermediate feature vector, perform linear transformation processing on the output layer; use the solidified weight matrix to perform 32-dimensional vector projection operation to generate a 32-dimensional vibration fingerprint code; S5.5: Based on the 32-dimensional acoustic fingerprint code, perform defect mode activation intensity standardization processing; according to the preset physical meaning mapping rules, perform threshold quantization operation on each dimension value to output a structured diagnostic vector representing the activation intensity distribution of typical assembly defect modes.
10. A speaker vibration closed-loop self-testing method based on built-in sweep frequency signal self-excitation according to claim 9, characterized in that, In S5.1, the weight matrix initialization operation is performed, specifically by using the fixed-point arithmetic unit of the Cortex-M4 microcontroller to perform the weight matrix initialization operation.