An AI-based Bluetooth earphone voice control method, system, and device

By using Fourier transform and multi-layer convolution operations, combined with cluster analysis and speech signal preprocessing, an optimal tuning parameter set is generated, solving the problems of sound source identification and sound quality optimization in complex audio scenarios for Bluetooth headphones, and achieving precise dynamic sound quality adjustment and improved user experience.

CN121148388BActive Publication Date: 2026-05-29SHENZHEN SHINETEK TECH CO LTD

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
SHENZHEN SHINETEK TECH CO LTD
Filing Date
2025-10-20
Publication Date
2026-05-29

AI Technical Summary

Technical Problem

Existing Bluetooth headset voice control methods struggle to achieve accurate sound source identification and dynamic sound quality optimization in complex audio scenarios, resulting in a lack of targeted sound quality adjustment and a poor user experience.

Method used

By acquiring real-time audio and speech signals, Fourier transform is performed to generate frequency distribution features. Multi-layer convolution operations are used to extract resonance feature vectors, cluster analysis is used to determine the sound source groups, the optimal tuning parameter set is generated, and smoothing and speech signal preprocessing are performed to achieve dynamic tuning control.

Benefits of technology

It achieves accurate sound source identification and dynamic sound quality optimization in complex audio scenarios, improves the targeting of sound quality adjustment and user experience, simplifies the activation process of the tuning scheme, reduces the risk of accidental triggering, and ensures the real-time performance and scene adaptability of audio signals.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121148388B_ABST
    Figure CN121148388B_ABST
Patent Text Reader

Abstract

The application relates to the technical field of Bluetooth earphone control, and discloses a Bluetooth earphone voice control method, system and equipment based on AI, the method comprising the following steps: acquiring real-time audio and a voice signal, and obtaining frequency distribution characteristics through Fourier transformation; a matched peak position set is obtained through multi-layer convolution operation and peak value detection; a single sound source feature subset is extracted and clustered according to the peak position set, so that a classified sound source group is obtained; a harmonic resonance feature vector is fused and analyzed, a preliminary tuning parameter set and a candidate list are generated, and an optimal set is determined; a dynamic tuning scheme is obtained through smoothing processing and weight adjustment; after the voice signal instruction is matched, the scheme is activated, the real-time audio is adjusted, and an optimized signal is output. The method can realize intelligent real-time audio optimization and multi-frequency adaptive adjustment based on a voice instruction, and improves the sound quality performance and auditory experience of the Bluetooth earphone in different scenes.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of Bluetooth headset control technology, and in particular to an AI-based Bluetooth headset voice control method, system, and device. Background Technology

[0002] Currently, with the widespread application of smart audio devices, Bluetooth headsets, as a core representative of portable audio devices, are increasingly highlighting the importance of their voice control functions. Intelligent voice control not only significantly enhances the user experience but also promotes the application of audio devices in diverse scenarios, such as music enjoyment, voice calls, and immersive entertainment. Researching how to optimize the audio processing and control of Bluetooth headsets through artificial intelligence technology has become an important topic in the field of smart hardware. This technology not only needs to meet users' personalized needs for sound quality but also needs to achieve precise real-time adjustment in complex and ever-changing audio environments to adapt to the needs of different users and scenarios.

[0003] In existing technologies, voice control methods for Bluetooth headsets primarily rely on preset equalizer (EQ) parameters or signal processing strategies based on simple audio classification algorithms. These methods achieve a certain degree of sound quality optimization by adjusting the gain or attenuation of audio signals within fixed frequency bands, or by switching processing modes according to limited scene categories. However, existing technologies suffer from insufficient audio signal recognition and dynamic adjustment capabilities. Specifically, since audio signals typically contain multiple sound sources (such as human voices, instrument sounds, and environmental noise), their frequency components are complex and overlapping, making it difficult for existing methods to accurately distinguish the resonant characteristics of different sound sources, resulting in a lack of targeted sound quality adjustment. For example, in complex audio scenarios such as symphonic music, existing technologies cannot effectively identify the fundamental frequency and harmonic structure of different instruments, thus failing to achieve fine-tuning for specific sound sources. Furthermore, these methods are mostly based on static preset logic and cannot generate adaptive tuning schemes according to real-time changes in audio content, resulting in abrupt sound quality adjustments and high latency when switching between different music styles, severely impacting the user experience.

[0004] Therefore, existing technologies struggle to achieve accurate sound source identification and dynamic sound quality optimization in complex audio scenarios. Summary of the Invention

[0005] This invention provides an AI-based Bluetooth headset voice control method, system, and device to solve the problem of difficulty in achieving accurate sound source identification and dynamic sound quality optimization in complex audio scenarios.

[0006] Firstly, in order to solve the above-mentioned technical problems, the present invention provides an AI-based Bluetooth headset voice control method, comprising:

[0007] Acquire real-time audio and speech signals, and perform Fourier transform on the real-time audio signals to obtain frequency distribution characteristics;

[0008] Multi-layer convolution operation is performed on the frequency distribution features to generate resonance feature vectors of different sound sources, and peak positions are extracted from the resonance feature vectors. It is then detected whether the peak positions match a preset peak threshold. If they match, the initial matching degree is calculated, and a set of matched peak positions is obtained.

[0009] Based on the matched set of peak positions, feature subsets of a single sound source are extracted, and the similarity matrix between feature subsets is determined by cluster analysis to obtain the classified sound source groups.

[0010] Based on the classified sound source groups, the resonant feature vectors are fused to generate a preliminary tuning parameter set. The distribution characteristics of the preliminary tuning parameter set are analyzed through the similarity matrix to generate a diverse candidate list and obtain the optimal tuning parameter set.

[0011] If the low-frequency gain in the optimal tuning parameter set exceeds the high-frequency gain, smoothing is performed to generate a smoothing parameter set. At the same time, multiple candidate tuning schemes are obtained from the candidate list by combining the smoothing parameter set. After low-frequency gain check and parameter weight adjustment, the final dynamic tuning scheme is obtained by sorting them by priority.

[0012] The voice signal is preprocessed and matched with a preset command. The dynamic tone-tuning scheme is activated based on the matching result to obtain an activated tone-tuning control mode.

[0013] Based on the activated tuning control mode, the real-time audio signal is adjusted in real time, and combined with low-frequency dominant scene recognition and smoothing parameter verification, an optimized audio signal is obtained.

[0014] Preferably, the step of acquiring real-time audio signals and speech signals, and performing a Fourier transform on the real-time audio signals to obtain frequency distribution characteristics, includes:

[0015] Acquire real-time audio signals;

[0016] Based on the real-time audio signal, signal conversion is performed to obtain a digital signal sequence;

[0017] Based on the digital signal sequence, a spectrum conversion is performed to obtain spectrum data;

[0018] Based on the spectrum data, the fundamental frequency value is extracted to obtain the position of the fundamental frequency peak.

[0019] Based on the fundamental frequency peak position, harmonic peaks are identified. If the peak intensity of a harmonic peak position is greater than a preset peak threshold, the harmonic peak positions are sorted to obtain an ordered harmonic position sequence.

[0020] Based on the ordered harmonic position sequence, the frequency interval between the fundamental frequency and the harmonics is calculated to obtain the frequency distribution characteristics.

[0021] Preferably, the step of performing multi-layer convolution operations on the frequency distribution features to generate resonant feature vectors for different sound sources, extracting peak positions from the resonant feature vectors, detecting whether the peak positions match a preset peak threshold, and if they match, calculating an initial matching degree and obtaining a set of matched peak positions, includes:

[0022] Perform multi-layer convolution operations on the frequency distribution features to obtain the resonance feature vector;

[0023] Based on the resonant feature vector, an initial set of peak positions is extracted, and the corresponding peak intensities are calculated to obtain a set of peak intensities;

[0024] If the peak intensity in the set of peak intensities is greater than the preset low-frequency gain threshold, then the resonant feature vector corresponding to the peak intensity is retained; otherwise, it is discarded to obtain a subset of resonant feature vectors that meet the conditions.

[0025] The high-frequency peak positions are extracted from the subset of resonant feature vectors. If the high-frequency peak positions are within a preset peak threshold range, the initial matching degree is calculated, and a set of matched peak positions is obtained.

[0026] Preferably, the step of extracting a feature subset of a single sound source based on the matched peak position set, and determining the similarity matrix between the feature subsets through cluster analysis to obtain the classified sound source group includes:

[0027] Based on the set of matched peak positions, the corresponding feature values ​​are extracted from the subset of resonant feature vectors, and the feature values ​​are arranged in chronological order to obtain the sound source feature sequence;

[0028] Based on the sound source feature sequence, calculate the eigenvalues ​​and eigenvectors of its covariance matrix, and then filter them to obtain the dimensionality-reduced eigenvectors.

[0029] Based on the reduced-dimensional feature vectors, the distance from each feature vector to the preset cluster center is calculated, and a similarity matrix between feature subsets is constructed to obtain the initial sound source group;

[0030] The matched set of peak positions is refined by interpolation to obtain an optimized set of peak positions;

[0031] Based on the optimized set of peak positions, the matching degree between the optimized peak positions and the preset position threshold is recalculated to obtain a new matching degree;

[0032] If the new matching degree is greater than the initial matching degree, then weighted fusion is initiated, and the difference between the new matching degree and the initial matching degree is used as the weight to correct the elements in the similarity matrix, thereby obtaining a corrected similarity matrix;

[0033] Based on the corrected similarity matrix, the initial sound source groups are reclassified to obtain the classified sound source groups.

[0034] Preferably, in an optional embodiment, the step of fusing the resonant feature vectors according to the classified sound source groups to generate a preliminary tuning parameter set, analyzing the distribution characteristics of the preliminary tuning parameter set through the similarity matrix, generating a diverse candidate list, and obtaining the optimal tuning parameter set includes:

[0035] Based on the classified sound source groups, the resonance feature vectors of each group are extracted to obtain a preliminary tuning parameter set;

[0036] The preliminary tuning parameter set is standardized, and the distribution characteristics of the preliminary tuning parameter set are analyzed through the similarity matrix to generate a diversified tuning scheme and obtain a diversified scheme list.

[0037] If the number of solutions in the diversified solution list is less than the preset solution number threshold, the solutions are regenerated; if the threshold is reached, the solutions are output directly to obtain a diversified alternative list.

[0038] Based on the diverse alternative list, the matching degree with the classified sound source groups is calculated to obtain the optimal tuning parameter set.

[0039] Preferably, if the low-frequency gain in the optimal tuning parameter set exceeds the high-frequency gain, smoothing processing is performed to generate a smoothing parameter set. Simultaneously, multiple candidate tuning schemes are obtained from the candidate list based on the smoothing parameter set. After low-frequency gain checking and parameter weight adjustment, the final dynamic tuning scheme is obtained by prioritizing the candidates, including:

[0040] If the low-frequency gain value in the optimal tuning parameter set is higher than the high-frequency gain value, then by comparing the difference between the two, it is determined whether smoothing processing is triggered. If it is triggered, the gain difference judgment result is obtained.

[0041] Based on the gain difference judgment result, the optimal tuning parameter set is smoothed by a preset weight allocation value to obtain a smoothed parameter set;

[0042] Multiple tuning schemes are obtained from the candidate list. Combined with the smoothing parameter set, the tuning scheme most similar to the smoothing parameter set is selected to obtain a candidate tuning scheme set.

[0043] The candidate tuning scheme set is subjected to low-frequency gain check to determine whether the low-frequency gain value in each scheme meets the preset gain threshold. If it does, a subset of tuning schemes that meet the conditions is obtained.

[0044] The subset of tuning schemes is processed a second time to adjust the weight distribution of parameters in each scheme and generate an optimized tuning scheme set.

[0045] The optimized tuning scheme set is sorted, and the final dynamic tuning scheme is obtained according to the priority rules of the final mode control.

[0046] Preferably, the step of preprocessing the speech signal and matching it with a preset instruction, and activating the dynamic tuning scheme based on the matching result to obtain an activated tuning control mode, includes:

[0047] Based on the speech signal, an analog-to-digital converter is used to generate a digital speech signal to obtain the original speech data;

[0048] The original speech data is preprocessed and background noise is removed to obtain clear speech data;

[0049] Feature extraction is performed on the clear speech data, and instruction parsing is performed to obtain the speech instruction content;

[0050] If the content of the voice command matches the preset command library, the corresponding tuning scheme is retrieved from the dynamic tuning scheme to obtain the target tuning scheme;

[0051] Based on the target tuning scheme, it is determined whether the voice signal is a single sound source. If it is determined to be a single sound source, the target tuning scheme is activated to obtain the activated tuning control mode.

[0052] Preferably, the step of performing real-time equalization adjustment on the real-time audio signal according to the activated tuning control mode, combined with low-frequency dominant scene recognition and smoothing parameter verification, to obtain an optimized audio signal includes:

[0053] According to the activated tuning control mode, the real-time audio signal is adjusted in real time to obtain the adjusted audio data;

[0054] Low-frequency features are extracted from the adjusted audio data to determine the scene classification result, and abnormal parameters are removed to obtain optimized audio parameters.

[0055] Based on the optimized audio parameters and the classified sound source groups, adjustments and optimizations are performed to obtain the optimized audio signal.

[0056] Secondly, the present invention provides an AI-based Bluetooth headset voice control system, comprising:

[0057] The real-time signal acquisition and frequency domain conversion module is used to acquire real-time audio and speech signals, and perform Fourier transform on the real-time audio signals to obtain frequency distribution characteristics.

[0058] The resonance feature extraction and single sound source determination module is used to perform multi-layer convolution operation on the frequency distribution features to generate resonance feature vectors of different sound sources, extract peak positions from the resonance feature vectors, detect whether the peak positions match a preset peak threshold, and if they match, calculate the initial matching degree and obtain the set of matched peak positions.

[0059] The sound source feature clustering and group classification module is used to extract feature subsets of a single sound source based on the matched peak position set, and determine the similarity matrix between feature subsets through cluster analysis to obtain the classified sound source groups;

[0060] The tuning parameter generation and optimal selection module is used to fuse the resonance feature vectors according to the classified sound source groups to generate a preliminary tuning parameter set, analyze the distribution characteristics of the preliminary tuning parameter set through the similarity matrix, generate a diverse candidate list, and obtain the optimal tuning parameter set.

[0061] The dynamic tuning scheme optimization module is used to perform smoothing processing if the low-frequency gain in the optimal tuning parameter set exceeds the high-frequency gain, generate a smoothing parameter set, and obtain multiple candidate tuning schemes from the candidate list in combination with the smoothing parameter set. After low-frequency gain check and parameter weight adjustment, the final dynamic tuning scheme is obtained by sorting them by priority.

[0062] The voice command recognition and tone-tuning mode activation module is used to preprocess the voice signal and match it with a preset command, and activate the dynamic tone-tuning scheme according to the matching result to obtain an activated tone-tuning control mode.

[0063] The real-time audio equalization adjustment and optimization output module is used to perform real-time equalization adjustment on the real-time audio signal according to the activated tuning control mode, and obtain the optimized audio signal by combining low-frequency dominant scene recognition and smoothing parameter verification.

[0064] Thirdly, the present invention also provides an electronic device, including a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor, wherein the processor executes the computer program to implement an AI-based Bluetooth headset voice control method as described in any one of the above.

[0065] Fourthly, the present invention also provides a computer-readable storage medium comprising a stored computer program, wherein, when the computer program is executed, it controls the device where the computer-readable storage medium is located to execute any one of the above-described AI-based Bluetooth headset voice control methods.

[0066] Compared with the prior art, the present invention has the following beneficial effects:

[0067] (1) This invention acquires real-time audio signals and speech signals, and performs Fourier transform on the real-time audio signals to obtain frequency distribution characteristics. By synchronously capturing the instantaneous signals of dynamic audio, it avoids feature deviations caused by signal delay. Furthermore, the Fourier transform can accurately convert the time-domain audio signals into frequency distribution characteristics in the frequency domain, fully presenting the signal strength of different frequency bands, laying the foundation for subsequent sound source identification. In addition, the synchronous acquisition of speech signals has reserved a command interaction interface in advance.

[0068] (2) This invention performs multi-layer convolution operations on the frequency distribution features to generate resonant feature vectors for different sound sources, extracts peak positions from the resonant feature vectors, and detects whether the peak positions match a preset peak threshold. If they match, the initial matching degree is calculated, and a set of matched peak positions is obtained. Through multi-layer convolution operations, the frequency distribution features can be refined layer by layer, highlighting the unique resonant characteristics of different sound sources and generating resonant feature vectors with strong discriminative power. This not only filters out peaks that meet preset conditions but also measures the correlation between peaks and sound source features through matching, reducing the risk of misjudgment from a single threshold.

[0069] (3) Based on the matched set of peak positions, this invention extracts feature subsets of a single sound source and determines the similarity matrix between feature subsets through cluster analysis to obtain the classified sound source groups. This invention first accurately extracts feature subsets corresponding to a single sound source based on the matched set of peak positions, ensuring that each subset reflects only the characteristics of a single sound source and avoiding cross-sound source interference. Then, it calculates the similarity matrix between feature subsets through cluster analysis to quantify the correlation between different subsets, thereby achieving clear grouping of sound sources. This invention solves the problems of feature overlap and group ambiguity in sound source classification. Through the extraction of feature subsets of a single sound source and the quantitative analysis of the similarity matrix, it significantly improves the accuracy and clarity of sound source grouping, providing a clear basis for the subsequent generation of targeted tuning parameters.

[0070] (4) Based on the classified sound source groups, this invention fuses the resonant feature vectors to generate a preliminary tuning parameter set. The distribution characteristics of the preliminary tuning parameter set are analyzed using the similarity matrix to generate a diverse candidate list, ultimately yielding the optimal tuning parameter set. First, the resonant feature vectors are fused according to the classified sound source groups to ensure the preliminary parameter set covers the tuning needs of all sound sources, avoiding parameter bias. Then, the parameter distribution characteristics are analyzed using the similarity matrix to select differentiated parameter combinations to form a candidate list. Finally, the optimal parameter set is selected from the candidate list to ensure that it conforms to the characteristics of each sound source and is compatible with the overall audio effect. This achieves diversified generation and precise optimization of the tuning parameter set, solving the problems of single parameters and poor adaptability. The generated optimal tuning parameter set can satisfy the personalized tuning needs of different sound sources while ensuring the coordination of the overall audio effect, thus improving the adaptability and parameter quality of the tuning system.

[0071] (5) If the low-frequency gain in the optimal tuning parameter set exceeds the high-frequency gain, the present invention performs smoothing processing to generate a smoothing parameter set. Simultaneously, multiple candidate tuning schemes are obtained from the candidate list based on the smoothing parameter set. After low-frequency gain checks and parameter weight adjustments, the final dynamic tuning scheme is obtained by prioritizing the selected schemes. By judging the low-frequency gain threshold, imbalanced parameters are smoothed to avoid sound quality imbalance. At the same time, multiple schemes are selected from the candidate list, and after a second low-frequency gain check and parameter weight adjustment, it is ensured that the schemes not only solve the gain imbalance problem but also adapt to different sound quality preferences. Finally, prioritizing the schemes further ensures their optimality and practicality. This solves the problems of low-frequency and high-frequency gain imbalance and insufficient scheme verification in tuning schemes. Through smoothing processing and multiple rounds of optimization and screening, the sound quality balance and reliability of the final dynamic tuning scheme are ensured.

[0072] (6) The present invention preprocesses the speech signal and matches it with a preset command, and activates the dynamic tone-tuning scheme according to the matching result to obtain an activated tone-tuning control mode. First, the speech signal is preprocessed (e.g., noise reduction and interference removal) to improve the clarity of the speech signal and lay the foundation for accurate command matching; then, by matching with the preset command, the dynamic tone-tuning scheme is automatically activated without manual operation, and the preprocessing step significantly reduces the command recognition error. The activation process of the tone-tuning scheme is simplified, solving the problems of cumbersome manual operation and low speech recognition accuracy, realizing convenient and automated activation of the tone-tuning scheme, improving operation efficiency and command recognition accuracy, and reducing the risk of false triggering.

[0073] (7) According to the activated tuning control mode, the present invention performs real-time equalization adjustment on the real-time audio signal, and obtains an optimized audio signal by combining low-frequency dominant scene recognition and smoothing parameter verification. Based on the activated tuning control mode, the real-time audio signal is dynamically equalized to ensure that the adjustment is synchronized with the signal change; at the same time, through low-frequency dominant scene recognition, it adapts to the low-frequency requirements of different scenes (such as movies and music), and then through smoothing parameter verification, it ensures the rationality of the adjustment parameters, and finally generates an optimized audio signal. It realizes the real-time performance and scene adaptability of audio adjustment, solves the problems of poor effect of fixed parameter adjustment and low scene matching, and further ensures the sound quality optimization effect through scene recognition and parameter verification. The generated optimized audio signal can better meet the user's auditory needs in different scenes. Attached Figure Description

[0074] Figure 1 This is a schematic flowchart of the AI-based Bluetooth headset voice control method provided in the first embodiment of the present invention;

[0075] Figure 2 This is a schematic diagram of the structure of the AI-based Bluetooth headset voice control system provided in the second embodiment of the present invention. Detailed Implementation

[0076] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0077] Reference Figure 1 The first embodiment of the present invention provides an AI-based Bluetooth headset voice control method, comprising the following steps:

[0078] S11: Acquire real-time audio and speech signals, and perform Fourier transform on the real-time audio signals to obtain frequency distribution characteristics;

[0079] S12, perform multi-layer convolution operation on the frequency distribution features to generate resonance feature vectors of different sound sources, extract peak positions from the resonance feature vectors, detect whether the peak positions match a preset peak threshold, if they match, calculate the initial matching degree, and obtain the set of matched peak positions.

[0080] S13, Based on the matched set of peak positions, extract the feature subsets of a single sound source, and determine the similarity matrix between the feature subsets through cluster analysis to obtain the classified sound source groups;

[0081] S14. Based on the classified sound source groups, the resonant feature vectors are fused to generate a preliminary tuning parameter set. The distribution characteristics of the preliminary tuning parameter set are analyzed through the similarity matrix to generate a diverse candidate list and obtain the optimal tuning parameter set.

[0082] S15, if the low-frequency gain in the optimal tuning parameter set exceeds the high-frequency gain, then smoothing is performed to generate a smoothing parameter set. At the same time, multiple candidate tuning schemes are obtained from the candidate list in combination with the smoothing parameter set. After low-frequency gain check and parameter weight adjustment, the final dynamic tuning scheme is obtained by sorting them by priority.

[0083] S16, preprocess the voice signal and match it with a preset instruction, and activate the dynamic tone control scheme according to the matching result to obtain the activated tone control mode.

[0084] S17, according to the activated tone control mode, the real-time audio signal is adjusted in real time, and the optimized audio signal is obtained by combining low-frequency dominant scene recognition and smoothing parameter verification.

[0085] In step S11, it is necessary to acquire real-time audio signals and speech signals, and perform Fourier transform on the real-time audio signals to obtain frequency distribution characteristics.

[0086] In one implementation, it is necessary to acquire real-time audio and speech signals, and perform a Fourier transform on the real-time audio signal to obtain frequency distribution characteristics, including:

[0087] Acquire real-time audio signals;

[0088] Based on the real-time audio signal, signal conversion is performed to obtain a digital signal sequence;

[0089] Based on the digital signal sequence, a spectrum conversion is performed to obtain spectrum data;

[0090] Based on the spectrum data, the fundamental frequency value is extracted to obtain the position of the fundamental frequency peak.

[0091] Based on the fundamental frequency peak position, harmonic peaks are identified. If the peak intensity of a harmonic peak position is greater than a preset peak threshold, the harmonic peak positions are sorted to obtain an ordered harmonic position sequence.

[0092] Based on the ordered harmonic position sequence, the frequency interval between the fundamental frequency and the harmonics is calculated to obtain the frequency distribution characteristics.

[0093] It should be noted that real-time audio signals are digital sound wave signals that are dynamically generated over time and are not stored or directly processed (including music, noise, etc.). Speech signals are a subset of these, specifically referring to signals containing semantic information produced by human vocal organs (possessing both periodicity and randomness). Acquiring both involves first using the built-in microphone of a Bluetooth headset to collect human voice signals at a sampling rate of 44.1kHz (human voice frequencies are typically between 300Hz and 3400Hz; a sufficiently high sampling rate ensures signal fidelity), converting analog audio into a digital sequence to provide raw data for subsequent processing. Next, a Fast Fourier Transform (FFT) is used for spectrum conversion; for example, dividing one second of audio into multiple 512-point frames yields a spectrum with a frequency resolution of 86Hz. The spectrum data shows the energy distribution of different frequencies of the signal, allowing the extraction of the fundamental frequency. By analyzing the spectrum, the strongest peak energy within the 100Hz to 500Hz human voice fundamental frequency range (e.g., 200Hz) is found to confirm the fundamental frequency. Subsequently, based on the fundamental frequency... Peak frequency identification identifies harmonic peaks. For example, when the fundamental frequency is 200Hz, harmonic peaks may appear at multiples of 400Hz, 600Hz, etc. Harmonic peaks with intensity greater than a preset threshold (determined based on background noise levels, such as exceeding twice the background noise) need to be selected to reduce interference. Then, the positions of the harmonic peaks are sorted (e.g., 400Hz, 600Hz, 800Hz in ascending order of frequency), and the frequency interval between the fundamental frequency and harmonics is calculated (e.g., a 200Hz interval between 200Hz and 400Hz, and a 400Hz interval between 200Hz and 600Hz). This interval reflects the periodicity of the signal and can be used for tone analysis or speech recognition (e.g., distinguishing different phonemes to improve speech recognition accuracy). Frequency distribution features are a comprehensive feature that reflects the pitch and timbre of audio, including the position, intensity, and interval patterns of frequency components, formed by combining the extracted fundamental frequency peak positions, the selected ordered harmonic position sequence that meets the intensity criteria, and the calculated fundamental frequency and harmonic frequency intervals.

[0094] In step S12, multi-layer convolution operation is performed on the frequency distribution features to generate resonance feature vectors of different sound sources, and peak positions are extracted from the resonance feature vectors. It is then detected whether the peak positions match a preset peak threshold. If they match, the initial matching degree is calculated, and a set of matched peak positions is obtained.

[0095] In one implementation, multi-layer convolution operations are performed on the frequency distribution features to generate resonant feature vectors for different sound sources. Peak positions are extracted from the resonant feature vectors, and it is detected whether the peak positions match a preset peak threshold. If they match, an initial matching degree is calculated, and a set of matched peak positions is obtained, including:

[0096] Perform multi-layer convolution operations on the frequency distribution features to obtain the resonance feature vector;

[0097] Based on the resonant feature vector, an initial set of peak positions is extracted, and the corresponding peak intensities are calculated to obtain a set of peak intensities;

[0098] If the peak intensity in the set of peak intensities is greater than the preset low-frequency gain threshold, then the resonant feature vector corresponding to the peak intensity is retained; otherwise, it is discarded to obtain a subset of resonant feature vectors that meet the conditions.

[0099] The high-frequency peak positions are extracted from the subset of resonant feature vectors. If the high-frequency peak positions are within a preset peak threshold range, the initial matching degree is calculated, and a set of matched peak positions is obtained.

[0100] It should be noted that, firstly, multi-layer convolutional operations are used to obtain the resonance feature vector: The input consists of the converted spectral data of the audio signal (such as a time-frequency matrix, which can be generated based on 44.1kHz sampling rate audio collected by a Bluetooth headset). For example, processing approximately 6 seconds of human voice signal can generate a 512x512 dimensional spectrum, clearly showing the frequency distribution of the audio signal at different time points. The convolutional kernel parameters of the convolutional neural network (3x3 size) are configured, and multi-layer convolutional operations are used to scan and extract the frequency distribution features in the spectral data, capturing frequency pattern information layer by layer, and finally outputting a vector containing resonance features, i.e., the resonance feature vector.

[0101] It is worth noting that the resonant feature extraction for a 512×512 single-channel audio spectrogram (generated from audio at a sampling rate of 44.1kHz) uses a convolutional network consisting of four convolutional layers (three layers followed by pooling layers), specifically: Layer 1 (input 512×512×1, 3×3 convolutional kernel, 64 output channels, ReLU activation, 2×2 max pooling, convolution stride 1×1, pooling stride 2×2, output 256×256×64); Layer 2 (input 256×256×64, 3×3 convolutional kernel, 128 output channels, ReLU activation, 2×2 max pooling, convolution stride 1×1, pooling stride 2×2, output 256×256×64); The system consists of four layers: a max pooling layer (1×1 stride for convolution, 2×2 stride for pooling, output 128×128×128), a third layer (128×128×128 input, 3×3 convolution kernel, 256 output channels, ReLU activation, 2×2 max pooling, 1×1 stride for convolution, 2×2 stride for pooling, output 64×64×256), and a fourth layer (64×64×256 input, 3×3 convolution kernel, 512 output channels, ReLU activation, no pooling, 1×1 stride for convolution, output 64×64×512). Finally, global average pooling is used to output a 512-dimensional resonant feature vector.

[0102] Secondly, the initial peak positions are extracted and peak intensities are calculated: Based on the resonant feature vector, a peak detection algorithm (such as traversing the vector amplitude to find local maximum points) is used to locate and extract the initial peak positions, forming a set of initial peak positions. Simultaneously, the amplitude corresponding to each initial peak position in the resonant feature vector is read, and these amplitudes are used as peak intensities to obtain a set of peak intensities. Specifically, a local range is defined using a symmetrical sliding window (the window size is determined by the frequency resolution Δf and the expected minimum peak interval; for example, when Δf = 10Hz and the minimum interval is 50Hz, the window radius is set to 5, and the total size is 11). The vector amplitude is traversed to find local maximum points. For flat peaks, the center point of a continuous equal amplitude region is selected, and it must meet dual threshold screening: the amplitude is not lower than the height threshold (such as 5%–10% of the maximum amplitude), and the prominence (the difference between the peak value and the maximum value of the first valley on the left and right) is not lower than the prominence threshold (such as 1.5–2 times the height threshold). For example, the amplitude corresponding to the 200Hz peak position is 0.8, and the amplitude corresponding to the 400Hz peak position is 0.6, etc., quantitatively reflecting the energy strength of each peak.

[0103] Subsequently, a subset of resonant feature vectors meeting the low-frequency gain threshold is selected: a low-frequency gain threshold is preset (e.g., 1.5 times the background noise intensity), and the relationship between each peak intensity in the peak intensity set and this threshold is compared one by one. If a peak intensity is greater than the preset low-frequency gain threshold, the resonant feature vector corresponding to that peak intensity is retained. If the peak intensity is less than or equal to the threshold, the corresponding vector is discarded, and finally, a subset of resonant feature vectors meeting the conditions is obtained.

[0104] Finally, high-frequency peak matching and initial matching degree calculation: Based on the resonant feature vector subset, a high-frequency range is defined (e.g., set to above 2000Hz according to the audio signal type; if processing human voice signals, it can also be adjusted to the high-frequency harmonic range above 400Hz according to the actual harmonic distribution). Peak positions within this frequency band are extracted to form a set of high-frequency peak positions. For example, peak positions corresponding to high-frequency harmonics such as 400Hz, 600Hz, and 800Hz are extracted. Then, a preset peak threshold range is set (e.g., set to 0.7 times the fundamental frequency peak intensity). For a frequency range, if the peak intensity at the fundamental frequency of 200Hz is 0.8, then the threshold range corresponds to the frequency range with an amplitude ≥ 0.56. The system determines whether the high-frequency peak position falls within this range. If the amplitude corresponding to a high-frequency peak position (e.g., 400Hz) is 0.6 ≥ 0.56, meeting the threshold requirement, then the initial matching degree is obtained by calculating the similarity between this peak and the reference feature (e.g., the peak at the fundamental frequency of 200Hz). Specifically, this includes first converting the frequency interval into a frequency matching degree (calculating the multiple offset between the high-frequency peak and the fundamental frequency). ,pass This yields the frequency fit from 0 to 1, such as k=1.0 for 400Hz. ,in As the reference frequency, Let be the frequency value of a peak value in the high-frequency band, and k be the multiple offset between the high-frequency peak value and the fundamental frequency. For frequency fit, round(k) is the nearest integer value for k, and amplitude similarity is also calculated. , The amplitude of the fundamental frequency peak value, The amplitude of a certain peak value in the high-frequency band. The amplitude similarity is represented by 0.6 / 0.8 = 0.75; then, a weighted sum with preset weights (e.g., frequency weight 0.6, amplitude weight 0.4) is used to fuse the results into an initial matching degree (M = 0.6 × 1.0 + 0.4 × 0.75 = 0.9). Simultaneously, the peak position of this high-frequency band is included in the set of matched peak positions, ultimately forming a complete set of matched peak positions (e.g., a set containing the 200Hz fundamental frequency, the 400Hz first harmonic, and the 600Hz second harmonic).

[0105] In step S13, based on the matched peak position set, a feature subset of a single sound source is extracted, and the similarity matrix between the feature subsets is determined by cluster analysis to obtain the classified sound source group.

[0106] In one implementation, based on the matched set of peak positions, a feature subset of a single sound source is extracted, and a similarity matrix between the feature subsets is determined through cluster analysis to obtain the classified sound source groups, including:

[0107] Based on the set of matched peak positions, the corresponding feature values ​​are extracted from the subset of resonant feature vectors, and the feature values ​​are arranged in chronological order to obtain the sound source feature sequence;

[0108] Based on the sound source feature sequence, calculate the eigenvalues ​​and eigenvectors of its covariance matrix, and then filter them to obtain the dimensionality-reduced eigenvectors.

[0109] Based on the reduced-dimensional feature vectors, the distance from each feature vector to the preset cluster center is calculated, and a similarity matrix between feature subsets is constructed to obtain the initial sound source group;

[0110] The matched set of peak positions is refined by interpolation to obtain an optimized set of peak positions;

[0111] Based on the optimized set of peak positions, the matching degree between the optimized peak positions and the preset position threshold is recalculated to obtain a new matching degree;

[0112] If the new matching degree is greater than the initial matching degree, then weighted fusion is initiated, and the difference between the new matching degree and the initial matching degree is used as the weight to correct the elements in the similarity matrix, thereby obtaining a corrected similarity matrix;

[0113] Based on the corrected similarity matrix, the initial sound source groups are reclassified to obtain the classified sound source groups.

[0114] It should be noted that, firstly, based on the matched set of peak positions (including the frequency positions of the 200Hz fundamental frequency, the 400Hz first harmonic, and the 600Hz second harmonic), feature values ​​for the corresponding positions are extracted from the filtered subset of resonant feature vectors. For example, key feature parameters such as amplitude of 0.8 at the 200Hz position, amplitude of 0.6 at the 400Hz position, and amplitude of 0.5 at the 600Hz position are extracted. According to the temporal order of the audio signal acquisition (e.g., sorted by time frames within a 1-second sampling period, from frame 1 to frame N), the extracted feature values ​​are arranged sequentially to form a sound source feature sequence that reflects the change of the sound source over time.

[0115] Based on the generated sound source feature sequence, the covariance matrix of the sequence is first calculated. The covariance matrix quantifies the linear correlation between different feature dimensions in the sequence (such as the eigenvalues ​​corresponding to 200Hz, 400Hz, and 600Hz). For example, the value of the element (200Hz, 400Hz) in the matrix can reflect the correlation between the amplitude changes of the two. Next, the eigenvalues ​​and eigenvectors of the covariance matrix are solved using an eigenvalue decomposition algorithm, where the eigenvalue represents the "importance" of the corresponding eigenvector (i.e., the proportion of effective information contained in that dimension). Then, a filtering process is performed: usually, the top K eigenvectors with larger eigenvalues ​​are retained (for example, if the original sequence is 3-dimensional, the first two vectors with the largest eigenvalues ​​are retained after filtering). The sound source feature sequence is projected into a low-dimensional space composed of these K eigenvectors, and finally, the dimensionality-reduced eigenvectors are obtained. For example, the original 3-dimensional sequence is reduced to a 2-dimensional vector "(0.85, 0.62), (0.87, 0.60), ..., (0.83, 0.63)", which reduces data redundancy while retaining the core frequency distribution information of the sequence.

[0116] First, determine the cluster centers. These centers are based on historical sound data. For example, after feature extraction and dimensionality reduction on a representative, labeled reference dataset, calculate the arithmetic mean of all sample points for each category (e.g., "male bass" and "female soprano") to obtain the preset centers. For example, the center for "male bass" is (0.82, 0.57), and the center for "female soprano" is (0.48, 0.91). For each dimensionality-reduced feature vector, calculate its distance to each cluster center using a distance calculation formula (e.g., Euclidean distance). For example, the distance from a dimensionality-reduced vector (0.85, 0.62) to the center for "male bass" is √[(0.85-0.8)²+(0.62-0.6)²]≈0.054, and the distance to the center for "female soprano" is √[(0.85-0.5)²+(0.62-0.9)²]≈0.448. Next, a similarity matrix is ​​constructed among the feature subsets: the values ​​of the matrix elements are inversely proportional to the distance, such as similarity = 1 / (1 + distance). The smaller the distance, the higher the similarity. For example, the similarity between the vector above and the center of "male bass voice" is approximately 1 / (1 + 0.054) ≈ 0.949, and the similarity with the center of "female high voice" is approximately 1 / (1 + 0.448) ≈ 0.691. Finally, according to the principle of "feature vectors belonging to the cluster center with the highest similarity", all feature vectors are grouped to obtain the initial sound source groups.

[0117] Based on the original set of matched peak positions (e.g., 200Hz, 400Hz, 600Hz), a frequency domain refinement interpolation algorithm (e.g., parabolic interpolation) is used to improve position accuracy. Parabolic interpolation can obtain the precise frequency by fitting the frequency-amplitude curves of three discrete points near the peak as a quadratic function and solving for the coordinates of the function's vertex. The size of the interpolation window is usually determined based on the sharpness of the peak and the spectral resolution: for sharper peaks, the window can be selected to include the peak point and its 1-2 adjacent frequency points (3-5 points in total), which can ensure fitting accuracy while avoiding the introduction of irrelevant frequency components; for peaks with slightly wider bandwidths, the window can be appropriately expanded to 5-7 points to balance fitting stability. For example, the 600Hz peak in the original set is obtained based on discrete frequency point detection and may have an error of ±5Hz. Parabolic interpolation performs a second fitting of the frequency-amplitude curves of its three nearby frequency points and calculates the precise frequency value corresponding to the curve vertex. Assuming the fitted value is 602Hz. This process is performed on all peak positions in the set one by one to obtain an optimized set of peak positions (such as 201Hz, 400Hz, 602Hz), reducing the impact of frequency positioning errors on subsequent matching degree calculations.

[0118] Based on the optimized set of peak positions, a preset position threshold is first determined (for example, based on the stability of the sound source's fundamental frequency, the fundamental frequency is allowed to fluctuate by ±3Hz, and the harmonics are allowed to fluctuate by ±5Hz; for example, the position threshold range corresponding to a 200Hz fundamental frequency is 197Hz~203Hz, and the threshold range corresponding to a 400Hz harmonic is 395Hz~405Hz). Then, it is determined whether each optimized peak position falls within the corresponding threshold range: for example, 201Hz falls within 197Hz~203Hz, 400Hz falls within 395Hz~405Hz, and 602Hz also meets the requirements if the corresponding threshold range is 595Hz~605Hz. Then, based on indicators such as the percentage of peaks that meet the threshold requirements and the degree of deviation between the peak position and the threshold center, the new matching degree is recalculated. The formula for calculating the new matching degree is: New matching degree = (Number of peaks that meet the threshold requirements / Total number of peaks) × [1 - (Sum of absolute values ​​of position deviations of all peaks that meet the requirements / Sum of total widths of corresponding thresholds)]. Wherein, the absolute value of position deviation refers to the absolute value of the difference between a single peak and the threshold center (e.g., the deviation between the center of 201Hz and 200Hz is 1Hz), and the total width of the threshold is the sum of the widths of all threshold ranges involved in the calculation (e.g., the width of the ±3Hz threshold is 6Hz). For example, if the original matching degree is 0.7, and all three peaks meet the threshold requirements (accounting for 100%), the total absolute value of the deviation is 1Hz (201Hz) + 0Hz (400Hz) + 2Hz (602Hz) = 3Hz. The corresponding total threshold width is 6Hz (fundamental frequency) + 10Hz (400Hz harmonic) + 10Hz (600Hz harmonic) = 26Hz. Then the new matching degree = 1 × [1 - (3 / 26)] ≈ 0.88, which is improved because the peak position is closer to the center of the threshold.

[0119] First, compare the new matching degree with the initial matching degree: if the new matching degree (e.g., 0.85) is greater than the initial matching degree (e.g., 0.7), then the weighted fusion mechanism is activated. The difference between the two is calculated as the weight, i.e., weight = new matching degree - initial matching degree = 0.15. Then, this weight is used to correct the elements in the constructed similarity matrix: the correction formula can be set as "corrected similarity = original similarity × (1 + weight)", for example, the similarity of a vector in the original "male bass voice" group is 0.949, and after correction it is 0.949 × (1 + 0.15) ≈ 1.091 (if it exceeds 1, it can be normalized to 1); for elements with low similarity (e.g., 0.691 in the "female high voice" group), after correction it is 0.691 × (1 + 0.15) ≈ 0.795. Through this correction method, the association between high matching degree features and corresponding cluster centers is strengthened, and the association between low matching degree features is weakened, finally obtaining the corrected similarity matrix.

[0120] Based on the corrected similarity matrix, the initial sound source groups are reclassified according to the principle of assigning feature vectors to the cluster centers with the highest similarity. For example, before correction, a vector with a similarity of 0.949 to "male bass" and 0.691 to "female soprano" is classified as "male bass." After correction, the similarities are 1.091 and 0.795 respectively, still belonging to "male bass." If a vector with a similarity of 0.72 and 0.70 to the two centers before correction is classified as "male bass," after correction, due to weighting, the similarity increases to 0.828 for "male bass" and 0.805 for "female soprano," maintaining its original classification. However, if a vector with a similarity of 0.71 and 0.70 before correction becomes 0.8165 and 0.805 after correction, it still belongs to "male bass." Finally, through re-judgment of all vectors, the classified sound source groups are obtained, ensuring that the grouping more closely matches the true characteristics of the sound sources.

[0121] In step S14, the resonant feature vectors are fused according to the classified sound source groups to generate a preliminary tuning parameter set. The distribution characteristics of the preliminary tuning parameter set are analyzed through the similarity matrix to generate a diverse candidate list and obtain the optimal tuning parameter set.

[0122] In one implementation, the resonant feature vectors are fused according to the classified sound source groups to generate a preliminary tuning parameter set. The distribution characteristics of the preliminary tuning parameter set are analyzed using the similarity matrix to generate a diverse candidate list, resulting in an optimal tuning parameter set, including:

[0123] Based on the classified sound source groups, the resonance feature vectors of each group are extracted to obtain a preliminary tuning parameter set;

[0124] The preliminary tuning parameter set is standardized, and the distribution characteristics of the preliminary tuning parameter set are analyzed through the similarity matrix to generate a diversified tuning scheme and obtain a diversified scheme list.

[0125] If the number of solutions in the diversified solution list is less than the preset solution number threshold, the solutions are regenerated; if the threshold is reached, the solutions are output directly to obtain a diversified alternative list.

[0126] Based on the diverse alternative list, the matching degree with the classified sound source groups is calculated to obtain the optimal tuning parameter set.

[0127] It should be noted that, firstly, when extracting the initial tuning parameter set, the sound sources are categorized into three groups. For example, the sound sources around 300Hz are grouped into the "mid-range group," and those around 600Hz are grouped into the "treble group." The similarity matrix shows that the Euclidean distance of the vectors within the mid-range group is less than 0.15, indicating compact grouping. Next, the core acoustic features of each group of sound sources are identified, and resonance feature vectors are extracted from them. For example, in the mid-range group, [0.8, 0.3], the first component is the normalized value of the resonance peak amplitude (unitless, originally in sound pressure level dBSPL), and the second component is the normalized value of the resonance frequency offset (unitless, originally in frequency difference Hz). Then, dynamic gain adjustment is used to generate preliminary tuning parameters: its inputs include core inputs (resonance feature vectors of each group containing the original acoustic parameters, such as mid-range group [0.8,0.3], high-range group [0.6,0.7]; preset target acoustic indicators, such as target resonant peak amplitude of mid-range group -4dBSPL, frequency offset 0Hz, target resonant peak amplitude of high-range group -5dBSPL, frequency offset 0Hz) and auxiliary inputs (group compactness represented by the mean Euclidean distance of vectors within the group, such as mean of 0.12 for mid-range group). First, calculate the base gain ΔG1 (target resonant peak amplitude minus actual resonant peak amplitude; for example, the midrange group's target is -4dB, and the actual value is -6dB, resulting in ΔG1 = +2dB; the treble group's target is -5dB, and the actual value is -4dB, resulting in ΔG1 = -1dB). Then, use a correction coefficient k (1 minus 0.5 multiplied by the ratio of the average Euclidean distance within the group to the maximum allowable distance; for example, for the midrange group, k = 1 - 0.5 × (0.12 / 0.15) = 0.6; the key is that the more compact the grouping, the closer k is to 1). The final gain ΔG = ΔG1 × k. After correction, the midrange group remains +2dB, and the treble group remains -1dB. The output is the gain adjustment value (in dB) corresponding to the center frequency of each group (e.g., +2dB at 300Hz, -1dB at 600Hz).

[0128] Secondly, in generating diverse tuning schemes and scheme lists, the initial tuning parameter set is first standardized to eliminate the influence of differences in parameter units and ranges. Then, using a similarity matrix obtained from the sound source feature vectors (such as a matrix generated through K-means clustering that reflects the acoustic similarity of sound source groups), the sound sources are clustered, grouping sound sources with similar acoustic features into one category. Next, a baseline value for historically adapted tuning parameters is calculated for each cluster group, and the parameter fluctuation range is set based on the variance of the acoustic similarity of sound sources within the group (the higher the similarity, the smaller the fluctuation range, ensuring parameter consistency within the group). Simultaneously, the baseline parameters of different cluster groups are forced to maintain significant differences (the lower the acoustic similarity between groups, the greater the parameter difference, strengthening the inter-group differentiation). This grouping mapping ensures that the distribution of tuning parameters matches the distribution of sound source acoustic characteristics. Finally, diverse schemes are generated based on the parameter ranges of each group. For example, after obtaining sound source groups such as "mid-range group" and "treble group" through acoustic similarity matrix clustering, the preliminary tuning parameter sets within each group (such as the mid-frequency EQ gain parameters of the mid-range group and the high-frequency EQ gain parameters of the treble group) can be further subdivided into parameter subclusters with different adjustment tendencies through hierarchical clustering. For example, the mid-range group is clustered to obtain subclusters such as "+3 to +5dB strong gain" and "-1 to +1dB neutral gain", and the treble group is clustered to obtain subclusters such as "+2 to +4dB boost" and "-2 to 0dB attenuation". Then, the center parameter of each subcluster is extracted as a benchmark, and combined with the fluctuation range of the parameter range of the group, it is converted into a specific gain adjustment amplitude (such as the mid-range strong gain scheme is set to +4dB±0.5dB, and the treble attenuation scheme is set to -1dB±0.3dB). Finally, a list of diverse schemes with significant differences in gain adjustment amplitude is formed. Subsequently, when obtaining a diverse candidate list, a preset threshold for the number of schemes is first set (e.g., 5). Then, the total number of schemes in the diverse scheme list is counted. If the number of schemes is less than the preset threshold (e.g., only 3 are generated), the process returns to the previous step, and the clustering parameters are adjusted (e.g., increasing the number of clusters in K-means clustering from 3 to 4) to regenerate tuning schemes and add them to the list. If the number of schemes reaches or exceeds the preset threshold (e.g., 5 or more are generated), duplicate or unreasonable schemes with extreme parameters are directly filtered out, and schemes that meet the requirements are retained, ultimately resulting in a diverse candidate list.

[0129] Finally, when calculating the optimal tuning parameter set, the matching degree calculation indicators and weights are set based on the essential acoustic characteristics of the sound source group, the core experience goals of the application scenario, and the sensitivity of the human ear to audio parameters (e.g., 300-3000Hz frequency band adaptation (weight 60%), dynamic range matching degree (weight 30%), and noise reduction tolerance (weight 10%)). Next, for each scheme in the diverse candidate list, the matching degree with the sound source category is calculated. The matching degree calculation process first establishes quantitative scoring rules for each preset matching degree indicator (e.g., for frequency band adaptation, 100 points are awarded if the scheme gain falls within the required range of the group, 20 points are deducted for every 1dB deviation, and 0 points are awarded if it is more than 5dB below the lower limit of the range; for noise reduction tolerance, 100 points are awarded if the noise reduction effect covers 80% of the noise frequency band within the group, 50 points are awarded for covering 50%, and so on). Then, for each scheme in the diverse candidate list, the matching degree is compared with the core characteristics of the corresponding sound source category. The system evaluates each feature individually, scores each indicator, and then multiplies the score of each indicator by its preset weight and sums them to obtain a specific numerical value of the matching degree between the scheme and the sound source classification (e.g., 85%). For example, in the Bluetooth headphone audio scenario, the "midrange group" tuning scheme increases the gain by 300Hz, achieving a matching degree of 85% with the features of this group, while the "treble group" scheme decreases the gain by 600Hz, achieving a matching degree of 80%. Finally, by comparing the matching degrees of all schemes, the scheme with the highest matching degree (e.g., the scheme with a matching degree of 85% for the "midrange group") is selected as the optimal tuning parameter set.

[0130] In step S15, if the low-frequency gain in the optimal tuning parameter set exceeds the high-frequency gain, a smoothing process is performed to generate a smoothing parameter set. At the same time, multiple candidate tuning schemes are obtained from the candidate list by combining the smoothing parameter set. After low-frequency gain check and parameter weight adjustment, the final dynamic tuning scheme is obtained by sorting them by priority.

[0131] In one implementation, if the low-frequency gain in the optimal tuning parameter set exceeds the high-frequency gain, smoothing is performed to generate a smoothing parameter set. Simultaneously, multiple candidate tuning schemes are obtained from the candidate list based on the smoothing parameter set. After low-frequency gain checking and parameter weight adjustment, the final dynamic tuning scheme is obtained by prioritizing the selected schemes, including:

[0132] If the low-frequency gain value in the optimal tuning parameter set is higher than the high-frequency gain value, then by comparing the difference between the two, it is determined whether smoothing processing is triggered. If it is triggered, the gain difference judgment result is obtained.

[0133] Based on the gain difference judgment result, the optimal tuning parameter set is smoothed by a preset weight allocation value to obtain a smoothed parameter set;

[0134] Multiple tuning schemes are obtained from the candidate list. Combined with the smoothing parameter set, the tuning scheme most similar to the smoothing parameter set is selected to obtain a candidate tuning scheme set.

[0135] The candidate tuning scheme set is subjected to low-frequency gain check to determine whether the low-frequency gain value in each scheme meets the preset gain threshold. If it does, a subset of tuning schemes that meet the conditions is obtained.

[0136] The subset of tuning schemes is processed a second time to adjust the weight distribution of parameters in each scheme and generate an optimized tuning scheme set.

[0137] The optimized tuning scheme set is sorted, and the final dynamic tuning scheme is obtained according to the priority rules of the final mode control.

[0138] It should be noted that when performing gain difference judgment and smoothing trigger, the low-frequency gain value and high-frequency gain value in the optimal tuning parameter set are first read (for example, in the Bluetooth headphone scenario, the optimal parameter set may be "low frequency +3dB, high frequency +1dB"), and the two are compared: if the low-frequency gain value (+3dB) is higher than the high-frequency gain value (+1dB), the difference between the two is calculated (3dB-1dB=2dB), and compared with the preset smoothing trigger threshold (this threshold setting needs to refer to the physical limits such as the maximum displacement of the headphone speaker unit to avoid the risk of distortion, combined with the human ear's perception of the frequency band imbalance threshold, and adapted to the needs of the Bluetooth headphone usage scenario, such as setting it to 1.5dB). Since 2dB>1.5dB, the trigger condition is met, and it is determined that smoothing processing needs to be started, and finally the gain difference judgment result of "smoothing processing required" is obtained; if the difference does not exceed the threshold (such as the difference 1dB<1.5dB), it is determined that no trigger is triggered, and the result of "no smoothing processing required" is obtained.

[0139] When generating the smoothing parameter set, based on the gain difference judgment result of "needing smoothing processing", the preset weight allocation value is called (e.g., low-frequency gain weight 0.4, high-frequency gain weight 0.6, taking into account both low-frequency emphasis and high-frequency balance) to perform smoothing calculation on the optimal tuning parameter set: taking "low frequency +3dB, high frequency +1dB" as an example, the smoothed low-frequency gain value = original low-frequency gain × low-frequency weight + original high-frequency gain × (1 - low-frequency weight) = 3dB × 0.4 + 1dB ×0.6=1.2dB+0.6dB=1.8dB, the smoothed high-frequency gain value = original high-frequency gain × high-frequency weight + original low-frequency gain × (1-high-frequency weight) = 1dB×0.6+3dB×0.4=0.6dB+1.2dB=1.8dB, while retaining the stable values ​​of other parameters (such as intermediate frequency +2dB), and finally integrating to obtain the smoothed parameter set (such as "low frequency +1.8dB, intermediate frequency +2dB, high frequency +1.8dB").

[0140] When selecting candidate tuning schemes, all tuning schemes are first extracted from a diverse list of options (e.g., the list includes scheme A "low frequency +2dB, mid frequency +2dB, high frequency +1.5dB", scheme B "low frequency +1.5dB, mid frequency +2dB, high frequency +2dB", scheme C "low frequency +3dB, mid frequency +2dB, high frequency +1dB", etc.). Then, each scheme is compared with a smoothing parameter set ("low frequency +1.8dB, mid frequency +2dB, high frequency +1"). The similarity of schemes A and B (if using the Euclidean distance formula, the smaller the distance, the higher the similarity) is as follows: Scheme A distance = √[(2-1.8)²+(1.5-1.8)²]≈0.36, Scheme B distance = √[(1.5-1.8)²+(2-1.8)²]≈0.36, Scheme C distance = √[(3-1.8)²+(1-1.8)²]=1.44. Schemes A and B with the highest similarity are selected to form a candidate tuning scheme set.

[0141] When a subset of tuning schemes that meet the criteria is obtained, a low-frequency gain threshold is first preset (e.g., based on the low-frequency output capability of Bluetooth headphones, it is set to "≥1.2dB and ≤2.2dB"). The low-frequency gain of each scheme in the candidate tuning scheme set is checked: the low-frequency gain of scheme A is +2dB, which is within the range of 1.2dB-2.2dB and meets the threshold requirement; the low-frequency gain of scheme B is +1.5dB, which also meets the requirement; if there is a scheme D in the candidate set with "low frequency +1dB, mid frequency +2dB, high frequency +2dB", its low-frequency gain 1dB < 1.2dB does not meet the threshold and is therefore eliminated. Finally, schemes A and B are retained to obtain a subset of tuning schemes that meet the criteria.

[0142] When generating the optimized tuning scheme set, the subset of tuning schemes (Scheme A and Scheme B) undergoes secondary processing to adjust the weight distribution of parameters in each scheme. For example, based on the "voice priority" scenario requirement of Bluetooth headphones, the weight of the mid-frequency (1kHz-3kHz) parameter is increased (from 0.3 to 0.5), while the weight of the high and low frequencies is decreased (from 0.35 to 0.25 each): Scheme A, after adjustment, becomes "low frequency +2dB × 0.25, mid frequency +2dB". The values ​​for Scheme B are: 1.5dB × 0.25 for low frequency, 1dB × 1 for mid frequency, and 0.375dB × 1 for high frequency. Scheme B, after adjustment, has the values ​​for low frequency, mid frequency, and high frequency, 1dB × 2 for high frequency, and 0.25 × 1 for mid frequency. These values, after quantization, are: 0.375dB × 1 for low frequency, 0.375dB × 1 for mid frequency, and 0.5dB × 1 for high frequency. The contribution values ​​reflect the relative importance of the gain in each frequency band after scene weighting and are a normalized indicator used for horizontal comparison and scheme evaluation. When determining the final dynamic tuning scheme, first set the priority rules for the final mode control (e.g., in the Bluetooth headphone scenario, "vocal clarity > bass fullness > treble smoothness"). Then, sort the optimized tuning scheme set (adjusted schemes A and B): compare the mid-frequency contribution values ​​of the two (both are +1dB, with comparable vocal clarity), then look at the low-frequency contribution values ​​(scheme A +0.5dB > scheme B +0.375dB, with better bass fullness), and finally look at the high-frequency contribution values ​​(scheme B +0.5dB > scheme A +0.375dB, but with the lowest priority). Therefore, scheme A is ranked higher than scheme B, and scheme A (adjusted: low-frequency +0.5dB contribution value, mid-frequency +1dB contribution value, and high-frequency +0.375dB contribution value, corresponding to the actual parameters "low-frequency +2dB, mid-frequency +2dB, and high-frequency +1.5dB") is selected as the final dynamic tuning scheme.

[0143] In step S16, the voice signal is preprocessed and matched with a preset instruction. The dynamic tuning scheme is activated according to the matching result to obtain the activated tuning control mode.

[0144] In one implementation, the speech signal is preprocessed and matched with a preset instruction. Based on the matching result, the dynamic tuning scheme is activated to obtain an activated tuning control mode, including:

[0145] Based on the speech signal, an analog-to-digital converter is used to generate a digital speech signal to obtain the original speech data;

[0146] The original speech data is preprocessed and background noise is removed to obtain clear speech data;

[0147] Feature extraction is performed on the clear speech data, and instruction parsing is performed to obtain the speech instruction content;

[0148] If the content of the voice command matches the preset command library, the corresponding tuning scheme is retrieved from the dynamic tuning scheme to obtain the target tuning scheme;

[0149] Based on the target tuning scheme, it is determined whether the voice signal is a single sound source. If it is determined to be a single sound source, the target tuning scheme is activated to obtain the activated tuning control mode.

[0150] It should be noted that when generating raw voice data, the user's voice signal is first collected through the microphone of the Bluetooth headset (such as the user saying "increase the volume"). To ensure the signal quality of subsequent processing, the analog voice signal is converted into a digital signal using an analog-to-digital converter (ADC) module. The sampling rate is usually set to 16kHz, which can preserve voice details (such as intonation and voiceless consonant features in human voice) and balance the amount of data, ultimately generating raw voice data containing time series and amplitude information.

[0151] In obtaining clear speech data, the raw speech data is first preprocessed, and then background noise is removed using noise reduction algorithms. Spectral subtraction is a commonly used noise reduction method: this algorithm first analyzes the spectrum of the raw speech data, identifies non-speech frequency components (such as noise), and then reduces noise by subtracting these components. For example, for speech data containing background noise from a coffee shop, spectral subtraction will first identify low-frequency noise below 50Hz (current interference in the coffee shop, low-frequency sounds of tables and chairs rubbing), suppress this type of noise through spectrum reduction, and at the same time retain the core speech frequency range of 200Hz to 4kHz (the main distribution range of human voices), ultimately generating clear speech data without significant environmental noise.

[0152] When extracting voice commands, the system first extracts features from clear speech data, focusing on Mel-frequency cepstral coefficients (MFCCs). These coefficients accurately capture the timbre and intonation features of speech (such as the tone change when a user says "raise," or the spectral density of the word "volume"). The extracted MFCC feature sequence is then input into the speech parser, where a Hidden Markov Model (HMM) performs command parsing. The model compares the current feature sequence with command templates in a pre-defined command library (such as "raise volume," "lower volume," and "switch to mid-range mode") one by one, determining the command content by calculating feature similarity. For example, when a user says "raise volume," the HMM model finds that the similarity between the current MFCC sequence and the "raise volume" template exceeds a preset threshold (e.g., 90%), ultimately parsing the voice command content of "raise volume."

[0153] It is worth noting that the templates in the speech command recognition model's instruction library are based on Hidden Markov Model (HMM) parameters trained on a large number of real user speech samples. The construction and recognition process is as follows: First, speech samples covering different genders, ages, accents, and with text labels are collected. After pre-emphasis, framing, and windowing preprocessing to eliminate noise and nonlinear effects, Mel-frequency cepstral coefficients (MFCCs) and their first and second-order differences are extracted as feature vectors. Next, for each instruction to be recognized, the MFCC feature sequence of the instruction's speech sample is used as input, and the Baum-Welch algorithm is used to iteratively optimize and obtain the HMM model parameters (state transition probabilities, observation probability distributions, etc.), forming an instruction "template" which is stored in the library. In the recognition stage, the MFCC feature sequence of the input speech is extracted first, and then the Viterbi algorithm is used to calculate its matching probability with each HMM template. The instruction corresponding to the template with the highest probability (e.g., exceeding the 90% threshold) is selected as the recognition result. For example, when a user says "increase the volume," its feature sequence has the highest matching probability with the corresponding template, thus successfully parsing the instruction. When searching for a target tuning scheme, the system first checks the matching of the parsed voice command with the preset command library. If the command library contains the command "increase volume" (and the match is successful), the system retrieves the corresponding tuning scheme from the previously built dynamic tuning scheme database. For example, the volume enhancement scheme is preset to increase the overall audio gain by 3dB (ensuring that the user can clearly perceive the volume change). The volume enhancement scheme is then determined as the target tuning scheme. If the command does not match (e.g., the user says "increase the volume" which is not included in the database and is not associated with "increase volume"), the user needs to be prompted to re-enter the command.

[0154] During the activation of the tone control mode, it is first necessary to determine whether the input voice signal is a single sound source. This avoids false activation caused by non-user voice (such as the voice of someone else speaking). Independent Component Analysis (ICA) is commonly used to achieve this: This algorithm can separate mixed signals (such as a mixture of user voice and the voice of someone else speaking). By analyzing the time-frequency characteristics of the signal, it determines whether the main voice source comes from the target user. For example, in a noisy subway environment, ICA will first separate the user's voice and the voices of surrounding passengers. Then, by comparing the energy ratio and pronunciation rhythm of the two, it confirms that the current main voice source is the user and not someone else, thus determining it to be a single sound source. If the determination result is a single sound source, the system immediately activates the target tone scheme (such as the "volume boost" scheme), generates a tone control signal containing specific parameters (such as a parameter command to increase gain by 3dB), and outputs this signal to the audio processing unit of the Bluetooth headset. After receiving the control signal, the audio processing unit will adjust the output parameters in real time. For example, if the target solution is "low frequency enhancement", which requires a 2dB increase in low frequency gain and no change in high frequency, the audio processing unit will adjust the low frequency band setting of the headphone equalizer to ensure clear sound quality and compliance with user instructions.

[0155] In step S17, the real-time audio signal is adjusted in real time according to the activated tuning control mode, and the optimized audio signal is obtained by combining low-frequency dominant scene recognition and smoothing parameter verification.

[0156] In one implementation, the real-time audio signal is subjected to real-time equalization adjustment according to the activated tuning control mode, and an optimized audio signal is obtained by combining low-frequency dominant scene recognition and smoothing parameter verification, including:

[0157] According to the activated tuning control mode, the real-time audio signal is adjusted in real time to obtain the adjusted audio data;

[0158] Low-frequency features are extracted from the adjusted audio data to determine the scene classification result, and abnormal parameters are removed to obtain optimized audio parameters.

[0159] Based on the optimized audio parameters and the classified sound source groups, adjustments and optimizations are performed to obtain the optimized audio signal.

[0160] It should be noted that, firstly, when performing real-time equalization to obtain adjusted audio data, the activated tuning control mode (such as "Midrange Optimization" mode with parameters of "Low Frequency +2dB, Mid Frequency +2dB, High Frequency +1.5dB", or "Volume Boost" mode with parameters of "Overall Gain +3dB") is read first, and the equalization adjustment parameters corresponding to this mode are loaded into the real-time audio processing of the Bluetooth headset. When a real-time audio signal (such as music played by the user or voice call) is input, different frequency bands are dynamically adjusted according to the parameters: for example, for "Midrange Optimization" mode, the low frequency band of 200-500Hz is increased by 2dB, the mid frequency band of 500Hz-3kHz is increased by 2dB, and the high frequency band above 3kHz is increased by 1.5dB, ensuring that the signal of each frequency band is enhanced or attenuated according to the preset parameters; during the adjustment process, the signal amplitude is monitored in real time to avoid distortion due to excessive gain (e.g., the gain is automatically limited when the amplitude exceeds the maximum quantization value), and finally, the adjusted audio data after frequency band equalization is output.

[0161] Secondly, when extracting low-frequency features and obtaining optimized audio parameters, the adjusted audio data is first subjected to low-frequency feature extraction: the low-frequency band of 20Hz to 200Hz is selected, and the energy distribution, peak frequency, amplitude change rate, and other features of this frequency band are analyzed by short-time Fourier transform (STFT). For example, the extracted low-frequency energy is concentrated in the 80Hz to 120Hz range, with a peak amplitude of 0.7 and an amplitude change rate of less than 0.1dB / frame. Next, the scene classification result is determined based on the low-frequency features: a preset scene feature library is used (e.g., "music scene" low-frequency energy ratio ≥30%, "call scene" low-frequency energy ratio ≤20%, "ambient sound scene" low-frequency energy fluctuation is large). The currently extracted low-frequency features are compared with the scene features in the library. If the energy ratio of 80Hz to 120Hz is 35%, it is determined to be a "music scene". Then, abnormal parameters are removed: a threshold for low-frequency parameters is set (e.g., peak amplitude is within the normal range of 0.3-0.9, and the rate of change is ≤0.2dB / frame). If the low-frequency peak amplitude of a certain frame is 1.1 (exceeding the threshold), it is determined to be an abnormal parameter. The average value of the preceding and following frames (e.g., 0.68) is used to replace the abnormal value. Finally, the corrected parameters of each frequency band are integrated to obtain optimized audio parameters (e.g., "gain +2dB for low frequency 80-120Hz, gain +2dB for mid frequency 500Hz-3kHz, and gain +1.5dB for high frequency above 3kHz").

[0162] Finally, when the optimized audio signal is obtained by combining the optimized audio parameters and the sound source group, the sound source group information classified above is first called (e.g., the current audio signal belongs to the "midrange group", characterized by a fundamental frequency of around 300Hz and prominent mid-frequency energy). The optimized audio parameters are then adapted and adjusted to match the acoustic characteristics of this group. For example, the "midrange group" needs to enhance the mid-frequency details corresponding to the 300Hz fundamental frequency. Based on the optimized audio parameters, the gain of the 500Hz-1kHz frequency band can be increased by 0.5dB (from +2dB to +2.5dB). Then, the AGC subsystem is started, and the parameter configuration corresponding to the "midrange group" is loaded (start-up time 15ms, release time 200ms, target level -15dBFS, compression ratio 3:1). The signal amplitude changes are monitored in real time. When the signal amplitude weakens (e.g., the volume of human voice decreases), the AGC automatically increases the gain (maximum gain limited to within 6dB) to maintain the target level. When the signal amplitude suddenly increases (e.g., a sudden plosive sound), the AGC quickly compresses the gain (response time ≤10ms) to avoid overload. The frequency domain signal processed by AGC is converted back to the time domain through inverse Fourier transform. At the same time, the overlapping addition method is used to eliminate phase distortion in the conversion process, ensuring that the output signal has no phase distortion and no frequency band breaks, and finally obtains an optimized audio signal with clear sound quality and adapted to the characteristics of the sound source.

[0163] It is worth noting that the coordination between AGC and equalization adjustment is achieved through a cascaded process of "equalization first, AGC later". First, equalization adjustment optimizes the energy distribution of the frequency band, and then AGC stabilizes the overall loudness to avoid mutual interference. When equalization makes a significant gain adjustment to a certain frequency band, AGC synchronously adapts the parameters (such as boosting the frequency band to lower the local threshold, weakening the frequency band to extend the release time), and combines the parameter binding strategy with the characteristics of the sound source group (such as using a medium compression ratio for the midrange group and a short start time for the treble group). At the same time, it sets the gain limit of key frequency bands (within ±2dB) to prevent AGC from destroying the frequency band balance set by equalization, and finally achieves the synergistic effect of spectrum optimization and loudness stabilization.

[0164] Reference Figure 2 The second embodiment of the present invention provides an AI-based Bluetooth headset voice control system, comprising:

[0165] The real-time signal acquisition and frequency domain conversion module is used to acquire real-time audio and speech signals, and perform Fourier transform on the real-time audio signals to obtain frequency distribution characteristics.

[0166] The resonance feature extraction and single sound source determination module is used to perform multi-layer convolution operation on the frequency distribution features to generate resonance feature vectors of different sound sources, extract peak positions from the resonance feature vectors, detect whether the peak positions match a preset peak threshold, and if they match, calculate the initial matching degree and obtain the set of matched peak positions.

[0167] The sound source feature clustering and group classification module is used to extract feature subsets of a single sound source based on the matched peak position set, and determine the similarity matrix between feature subsets through cluster analysis to obtain the classified sound source groups;

[0168] The tuning parameter generation and optimal selection module is used to fuse the resonance feature vectors according to the classified sound source groups to generate a preliminary tuning parameter set, analyze the distribution characteristics of the preliminary tuning parameter set through the similarity matrix, generate a diverse candidate list, and obtain the optimal tuning parameter set.

[0169] The dynamic tuning scheme optimization module is used to perform smoothing processing if the low-frequency gain in the optimal tuning parameter set exceeds the high-frequency gain, generate a smoothing parameter set, and obtain multiple candidate tuning schemes from the candidate list in combination with the smoothing parameter set. After low-frequency gain check and parameter weight adjustment, the final dynamic tuning scheme is obtained by sorting them by priority.

[0170] The voice command recognition and tone-tuning mode activation module is used to preprocess the voice signal and match it with a preset command, and activate the dynamic tone-tuning scheme according to the matching result to obtain an activated tone-tuning control mode.

[0171] The real-time audio equalization adjustment and optimization output module is used to perform real-time equalization adjustment on the real-time audio signal according to the activated tuning control mode, and obtain the optimized audio signal by combining low-frequency dominant scene recognition and smoothing parameter verification.

[0172] It should be noted that the AI-based Bluetooth headset voice control system provided in this embodiment of the invention is used to execute all the process steps of the AI-based Bluetooth headset voice control method in the above embodiment. The working principles and beneficial effects of the two are one-to-one, so they will not be described again.

[0173] This invention also provides an electronic device. The electronic device includes a processor, a memory, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the steps described in the various embodiments of the AI-based Bluetooth headset voice control method, for example... Figure 1 The step S11 shown. Alternatively, when the processor executes the computer program, it implements the functions of each module / unit in the above-described device embodiments.

[0174] For example, the computer program may be divided into one or more modules / units, which are stored in the memory and executed by the processor to complete the present invention. The one or more modules / units may be a series of computer program instruction segments capable of performing a specific function, which describe the execution process of the computer program in the electronic device.

[0175] The electronic device may be a desktop computer, laptop, handheld computer, or smart tablet, etc. The electronic device may include, but is not limited to, a processor and memory. Those skilled in the art will understand that the above components are merely examples of electronic devices and do not constitute a limitation on the electronic device. It may include more or fewer components than described above, or combine certain components, or different components. For example, the electronic device may also include input / output devices, network access devices, buses, etc.

[0176] The processor can be a Central Processing Unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor can be a microprocessor or any conventional processor. The processor is the control center of the electronic device, connecting all parts of the electronic device via various interfaces and lines.

[0177] The memory can be used to store the computer programs and / or modules. The processor implements various functions of the electronic device by running or executing the computer programs and / or modules stored in the memory and by calling data stored in the memory. The memory may mainly include a program storage area and a data storage area. The program storage area may store the operating system, at least one application program required for a function (such as sound playback function, image playback function, etc.), etc.; the data storage area may store data created according to the use of the mobile phone (such as audio data, phonebook, etc.). In addition, the memory may include high-speed random access memory, and may also include non-volatile memory, such as hard disk, memory, plug-in hard disk, smart media card (SMC), secure digital card (SD card), flash card, at least one disk storage device, flash memory device, or other volatile solid-state storage device.

[0178] If the modules / units integrated in the electronic device are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, all or part of the processes in the methods of the above embodiments can also be implemented by a computer program instructing related hardware. The computer program can be stored in a computer-readable storage medium, and when executed by a processor, it can implement the steps of the various method embodiments described above. The computer program includes computer program code, which can be in the form of source code, object code, executable files, or certain intermediate forms. The computer-readable medium can include: any entity or device capable of carrying the computer program code, recording media, USB flash drives, portable hard drives, magnetic disks, optical disks, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signals, telecommunication signals, and software distribution media, etc. It should be noted that the content included in the computer-readable medium can be appropriately added or removed according to the requirements of legislation and patent practice in the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, computer-readable media do not include electrical carrier signals and telecommunication signals.

[0179] It should be noted that the device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Furthermore, in the accompanying drawings of the device embodiments provided by this invention, the connection relationships between modules indicate that they have communication connections, which can be specifically implemented as one or more communication buses or signal lines. Those skilled in the art can understand and implement this without any creative effort.

[0180] The specific embodiments described above further illustrate the purpose, technical solution, and beneficial effects of the present invention. It should be understood that the above descriptions are merely specific embodiments of the present invention and are not intended to limit the scope of protection of the present invention. In particular, it should be noted that any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention for those skilled in the art.

Claims

1. A voice control method for Bluetooth headsets based on AI, characterized in that, include: Acquire real-time audio and speech signals, and perform Fourier transform on the real-time audio signals to obtain frequency distribution characteristics; Multi-layer convolution operation is performed on the frequency distribution features to generate resonance feature vectors of different sound sources, and peak positions are extracted from the resonance feature vectors. It is then detected whether the peak positions match a preset peak threshold. If they match, the initial matching degree is calculated, and a set of matched peak positions is obtained. Based on the matched set of peak positions, feature subsets of a single sound source are extracted, and the similarity matrix between feature subsets is determined by cluster analysis to obtain the classified sound source groups. Based on the classified sound source groups, the resonant feature vectors are fused to generate a preliminary tuning parameter set. The distribution characteristics of the preliminary tuning parameter set are analyzed through the similarity matrix to generate a diverse candidate list and obtain the optimal tuning parameter set. If the low-frequency gain in the optimal tuning parameter set exceeds the high-frequency gain, smoothing is performed to generate a smoothing parameter set. At the same time, multiple candidate tuning schemes are obtained from the candidate list by combining the smoothing parameter set. After low-frequency gain check and parameter weight adjustment, the final dynamic tuning scheme is obtained by sorting them by priority. The voice signal is preprocessed and matched with a preset command. The dynamic tone-tuning scheme is activated based on the matching result to obtain an activated tone-tuning control mode. Based on the activated tuning control mode, the real-time audio signal is adjusted in real time, and combined with low-frequency dominant scene recognition and smoothing parameter verification, an optimized audio signal is obtained.

2. The AI-based Bluetooth headset voice control method according to claim 1, characterized in that, The process of acquiring real-time audio and speech signals, and performing Fourier transform on the real-time audio signals to obtain frequency distribution characteristics, includes: Acquire real-time audio signals; Based on the real-time audio signal, signal conversion is performed to obtain a digital signal sequence; Based on the digital signal sequence, a spectrum conversion is performed to obtain spectrum data; Based on the spectrum data, the fundamental frequency value is extracted to obtain the position of the fundamental frequency peak. Based on the fundamental frequency peak position, harmonic peaks are identified. If the peak intensity of a harmonic peak position is greater than a preset peak threshold, the harmonic peak positions are sorted to obtain an ordered harmonic position sequence. Based on the ordered harmonic position sequence, the frequency interval between the fundamental frequency and the harmonics is calculated to obtain the frequency distribution characteristics.

3. The AI-based Bluetooth headset voice control method according to claim 1, characterized in that, The process involves performing multi-layer convolution operations on the frequency distribution features to generate resonant feature vectors for different sound sources, extracting peak positions from these resonant feature vectors, detecting whether the peak positions match a preset peak threshold, and if they match, calculating an initial matching degree and obtaining a set of matched peak positions, including: Perform multi-layer convolution operations on the frequency distribution features to obtain the resonance feature vector; Based on the resonant feature vector, an initial set of peak positions is extracted, and the corresponding peak intensities are calculated to obtain a set of peak intensities; If the peak intensity in the set of peak intensities is greater than the preset low-frequency gain threshold, then the resonant feature vector corresponding to the peak intensity is retained; otherwise, it is discarded to obtain a subset of resonant feature vectors that meet the conditions. The high-frequency peak positions are extracted from the subset of resonant feature vectors. If the high-frequency peak positions are within a preset peak threshold range, the initial matching degree is calculated, and a set of matched peak positions is obtained.

4. The AI-based Bluetooth headset voice control method according to claim 3, characterized in that, The step involves extracting a feature subset of a single sound source based on the matched peak position set, and determining the similarity matrix between the feature subsets through cluster analysis to obtain the classified sound source groups, including: Based on the set of matched peak positions, the corresponding feature values ​​are extracted from the subset of resonant feature vectors, and the feature values ​​are arranged in chronological order to obtain the sound source feature sequence; Based on the sound source feature sequence, calculate the eigenvalues ​​and eigenvectors of its covariance matrix, and then filter them to obtain the dimensionality-reduced eigenvectors. Based on the reduced-dimensional feature vectors, the distance from each feature vector to the preset cluster center is calculated, and a similarity matrix between feature subsets is constructed to obtain the initial sound source group; The matched set of peak positions is refined by interpolation to obtain an optimized set of peak positions; Based on the optimized set of peak positions, the matching degree between the optimized peak positions and the preset position threshold is recalculated to obtain a new matching degree; If the new matching degree is greater than the initial matching degree, then weighted fusion is initiated, and the difference between the new matching degree and the initial matching degree is used as the weight to correct the elements in the similarity matrix, thereby obtaining a corrected similarity matrix; Based on the corrected similarity matrix, the initial sound source groups are reclassified to obtain the classified sound source groups.

5. The AI-based Bluetooth headset voice control method according to claim 1, characterized in that, The process involves fusing the resonant feature vectors based on the classified sound source groups to generate a preliminary tuning parameter set. The distribution characteristics of this preliminary tuning parameter set are then analyzed using the similarity matrix to generate a diverse candidate list, ultimately yielding the optimal tuning parameter set, including: Based on the classified sound source groups, the resonance feature vectors of each group are extracted to obtain a preliminary tuning parameter set; The preliminary tuning parameter set is standardized, and the distribution characteristics of the preliminary tuning parameter set are analyzed through the similarity matrix to generate a diversified tuning scheme and obtain a diversified scheme list. If the number of solutions in the diversified solution list is less than the preset solution number threshold, the solutions are regenerated; if the threshold is reached, the solutions are output directly to obtain a diversified alternative list. Based on the diverse alternative list, the matching degree with the classified sound source groups is calculated to obtain the optimal tuning parameter set.

6. The AI-based Bluetooth headset voice control method according to claim 1, characterized in that, If the low-frequency gain in the optimal tuning parameter set exceeds the high-frequency gain, smoothing is performed to generate a smoothing parameter set. Simultaneously, multiple candidate tuning schemes are obtained from the candidate list based on this smoothing parameter set. After low-frequency gain checks and parameter weight adjustments, the final dynamic tuning scheme is obtained by prioritizing these schemes, including: If the low-frequency gain value in the optimal tuning parameter set is higher than the high-frequency gain value, then by comparing the difference between the two, it is determined whether smoothing processing is triggered. If it is triggered, the gain difference judgment result is obtained. Based on the gain difference judgment result, the optimal tuning parameter set is smoothed by a preset weight allocation value to obtain a smoothed parameter set; Multiple tuning schemes are obtained from the candidate list. Combined with the smoothing parameter set, the tuning scheme most similar to the smoothing parameter set is selected to obtain a candidate tuning scheme set. The candidate tuning scheme set is subjected to low-frequency gain check to determine whether the low-frequency gain value in each scheme meets the preset gain threshold. If it does, a subset of tuning schemes that meet the conditions is obtained. The subset of tuning schemes is processed a second time to adjust the weight distribution of parameters in each scheme and generate an optimized tuning scheme set. The optimized tuning scheme set is sorted, and the final dynamic tuning scheme is obtained according to the priority rules of the final mode control.

7. The AI-based Bluetooth headset voice control method according to claim 1, characterized in that, The process of preprocessing the speech signal and matching it with a preset command, and activating the dynamic tuning scheme based on the matching result to obtain an activated tuning control mode, includes: Based on the speech signal, an analog-to-digital converter is used to generate a digital speech signal to obtain the original speech data; The original speech data is preprocessed and background noise is removed to obtain clear speech data; Feature extraction is performed on the clear speech data, and instruction parsing is performed to obtain the speech instruction content; If the content of the voice command matches the preset command library, the corresponding tuning scheme is retrieved from the dynamic tuning scheme to obtain the target tuning scheme; Based on the target tuning scheme, it is determined whether the voice signal is a single sound source. If it is determined to be a single sound source, the target tuning scheme is activated to obtain the activated tuning control mode.

8. The AI-based Bluetooth headset voice control method according to claim 1, characterized in that, The step involves real-time equalization adjustment of the real-time audio signal based on the activated tuning control mode, combined with low-frequency dominant scene recognition and smoothing parameter verification, to obtain an optimized audio signal, including: According to the activated tuning control mode, the real-time audio signal is adjusted in real time to obtain the adjusted audio data; Low-frequency features are extracted from the adjusted audio data to determine the scene classification result, and abnormal parameters are removed to obtain optimized audio parameters. Based on the optimized audio parameters and the classified sound source groups, adjustments and optimizations are performed to obtain the optimized audio signal.

9. An AI-based Bluetooth headset voice control system, characterized in that, include: The real-time signal acquisition and frequency domain conversion module is used to acquire real-time audio and speech signals, and perform Fourier transform on the real-time audio signals to obtain frequency distribution characteristics. The resonance feature extraction and single sound source determination module is used to perform multi-layer convolution operation on the frequency distribution features to generate resonance feature vectors of different sound sources, extract peak positions from the resonance feature vectors, detect whether the peak positions match a preset peak threshold, and if they match, calculate the initial matching degree and obtain the set of matched peak positions. The sound source feature clustering and group classification module is used to extract feature subsets of a single sound source based on the matched peak position set, and determine the similarity matrix between feature subsets through cluster analysis to obtain the classified sound source groups; The tuning parameter generation and optimal selection module is used to fuse the resonance feature vectors according to the classified sound source groups to generate a preliminary tuning parameter set, analyze the distribution characteristics of the preliminary tuning parameter set through the similarity matrix, generate a diverse candidate list, and obtain the optimal tuning parameter set. The dynamic tuning scheme optimization module is used to perform smoothing processing if the low-frequency gain in the optimal tuning parameter set exceeds the high-frequency gain, generate a smoothing parameter set, and obtain multiple candidate tuning schemes from the candidate list in combination with the smoothing parameter set. After low-frequency gain check and parameter weight adjustment, the final dynamic tuning scheme is obtained by sorting them by priority. The voice command recognition and tone-tuning mode activation module is used to preprocess the voice signal and match it with a preset command, and activate the dynamic tone-tuning scheme according to the matching result to obtain an activated tone-tuning control mode. The real-time audio equalization adjustment and optimization output module is used to perform real-time equalization adjustment on the real-time audio signal according to the activated tuning control mode, and obtain the optimized audio signal by combining low-frequency dominant scene recognition and smoothing parameter verification.

10. An electronic device, characterized in that, The device includes a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor, wherein the processor, when executing the computer program, implements an AI-based Bluetooth headset voice control method as described in any one of claims 1 to 8.