Wireless earphone tone quality optimization method and system, electronic equipment and storage medium
By performing Mel spectrum processing and feature extraction of wireless headphone sound signals, combining the timing dependence model and ArcFace loss function for abnormal detection and quality evaluation, an adjustment strategy is formulated to optimize sound quality, solving the problem of poor sound quality of wireless headphones and achieving effective improvement of sound quality.
Patent Information
- Application Number
- CN202510229881.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-28
- Publication Date
- 2025-06-03
AI Technical Summary
Existing wireless headphones can cause poor sound quality when output after sound is transmitted.
By obtaining the sound signals in wireless headphones, performing Mel spectrum processing and segmentation, using the CNN network structure to extract the feature vectors of Mel fragments, combining the timing dependence model and the ArcFace loss function for abnormal detection and quality evaluation, and formulating adjustment strategies to optimize sound quality.
Through quality evaluation and abnormal detection, the sound quality of wireless headphones can be effectively improved, ensuring the optimization and adjustment of sound signals, and improving user experience.
Smart Images

Figure CN120091251A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of wireless earphones, and particularly to a method, a system, an electronic device, and a storage medium for optimizing the sound quality of wireless earphones. Background Art
[0002] A wireless earphone is one in which the middle wire is replaced by radio waves. It is connected from the audio output of a computer to a transmitter, and then the transmitter sends the radio waves to the earphone at the receiving end, and the receiving end is equivalent to a radio.
[0003] With the popularization of earphones, most users have increasing requirements for the functions of earphones, not only limited to the functions provided by wired earphones, but also hope that the earphones can be more convenient and adaptable to more application scenarios, such as answering calls while driving and getting rid of the bondage of intertwined earphone wires. Therefore, wireless earphones have emerged and quickly spread, and have been more widely used.
[0004] In the prior art, since wireless earphones do not use earphone wires, the sound quality will deteriorate when the sound is transmitted and output. Summary of the Invention
[0005] Based on this, the purpose of the present invention is to provide a method, a system, an electronic device, and a storage medium for optimizing the sound quality of wireless earphones to solve the deficiencies in the above-mentioned prior art.
[0006] In a first aspect, the present invention provides a method for optimizing the sound quality of wireless earphones, and the method includes:
[0007] Obtain a sound signal in the wireless earphone;
[0008] Process the sound signal to obtain the Mel spectrogram of the sound signal, and segment the Mel spectrogram to obtain a number of Mel segments;
[0009] Obtain the feature vector of the Mel segment based on the CNN network structure, and evaluate the sound signal based on the time series dependence model and the feature vector to obtain the quality evaluation index of the sound signal;
[0010] Perform data augmentation on the Mel spectrogram to obtain an enhanced Mel spectrogram, construct an anomaly detection model, and use the Mel spectrogram as a training set to train the anomaly detection model to obtain a trained anomaly detection model;
[0011] Based on the trained anomaly detection model, and in combination with the ArcFace loss function, detect the sound signal to obtain the anomaly category of the sound signal;
[0012] Form an adjustment strategy based on the abnormal category and the quality evaluation index, and adjust the sound signal of the wireless earphone based on the adjustment strategy.
[0013] Compared with the prior art, the beneficial effects of the present invention are as follows: By using the feature vectors of the Mel segments and evaluating them through the temporal dependence model to obtain the quality evaluation index, the quality judgment of the headphone sound quality can be obtained. Then, by using the trained anomaly detection model and combining with the ArcFace loss function to detect the sound signal, the abnormal category of the sound signal can be obtained, so that the abnormal situation of the headphone sound signal can be obtained. Finally, an adjustment strategy is formulated according to the abnormal category and the quality evaluation index for sound adjustment, so that the sound signal can be adjusted according to the quality evaluation index to improve and optimize the sound quality of the wireless earphone.
[0014] Further, the step of obtaining the sound signal in the wireless earphone includes:
[0015] Obtain the audio information of the wireless earphone, and the audio information includes audio feature information and audio source information.
[0016] Further, the step of processing the sound signal to obtain the Mel spectrogram of the sound signal and segmenting the Mel spectrogram to obtain several Mel segments includes:
[0017] Perform pre-emphasis on the sound signal to enhance the high-frequency components in the sound signal, and window the sound signal to obtain a preprocessed sound signal;
[0018] Extract the Mel spectrogram from the preprocessed sound signal;
[0019] Segment the Mel spectrogram with a preset width to obtain several Mel segments of the same width, and input the several Mel segments of the same width into the CNN network.
[0020] Further, the step of obtaining the feature vector of the Mel segment based on the CNN network structure, evaluating the sound signal based on the temporal dependence model and the feature vector to obtain the quality evaluation index of the sound signal includes:
[0021] Extract the features of the Mel segment based on several residual structures in the CNN network structure to obtain shallow features;
[0022] Extract the deep features of the Mel segment based on the convolutional layer in the CNN network structure, and fuse the shallow features and the deep features to obtain a fused feature map;
[0023] Unfold the fused feature map to obtain a feature vector;
[0024] Based on the Transformer network structure, the self-attention mechanism is adopted to capture the global information in the feature vector and evaluate it to obtain the quality evaluation index of the sound signal.
[0025] Further, the step of performing data augmentation on the Mel spectrogram to obtain the enhanced Mel spectrogram includes:
[0026] Linearly combine the Mel spectrograms of two different sound samples to obtain a new spectrogram;
[0027] And perform time-domain masking and frequency-domain masking on the new spectrogram to perform data augmentation on the Mel spectrogram and obtain the enhanced Mel spectrogram.
[0028] Further, the step of detecting the sound signal based on the trained anomaly detection model and combining the ArcFace loss function to obtain the anomaly category of the sound signal includes:
[0029] Normalize the features of the sound signal based on the Softmax loss function and map them into a hypersphere to obtain the normalized features;
[0030] Based on the ArcFace loss function, calculate the angle between the feature vector of the sound signal and the normalized features, and output the anomaly category of the sound signal based on the Softmax loss function.
[0031] Further, the step of formulating an adjustment strategy based on the anomaly category and the quality evaluation index and adjusting the sound signal of the wireless headset based on the adjustment strategy includes:
[0032] Obtain the abnormal sound signal of the wireless headset based on the anomaly category;
[0033] Based on the abnormal sound signal and according to the quality evaluation index, formulate a sound adjustment strategy to correct the sound signal, and determine whether the corrected sound signal meets the quality evaluation index;
[0034] If not, continue to correct the sound signal until the corrected sound signal meets the quality evaluation index.
[0035] In a second aspect, the present invention further provides a wireless headset sound quality optimization system, and the system includes:
[0036] An acquisition module for acquiring the sound signal in the wireless headset;
[0037] A processing module, configured to process the sound signal to obtain a Mel spectrogram of the sound signal, and segment the Mel spectrogram to obtain a plurality of Mel segments;
[0038] An evaluation module, configured to obtain a feature vector of the Mel segment based on a CNN network structure, and evaluate the sound signal based on a temporal dependence model and the feature vector to obtain a quality evaluation index of the sound signal;
[0039] An enhancement module, configured to perform data enhancement on the Mel spectrogram to obtain an enhanced Mel spectrogram, construct an anomaly detection model, and use the Mel spectrogram as a training set to train the anomaly detection model to obtain a trained anomaly detection model;
[0040] A detection module, configured to detect the sound signal based on the trained anomaly detection model and in combination with an ArcFace loss function to obtain an anomaly category of the sound signal;
[0041] An adjustment module, configured to formulate an adjustment strategy based on the anomaly category and the quality evaluation index, and adjust the sound signal of the wireless earphone based on the adjustment strategy.
[0042] In a third aspect, the present invention further provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor, where when the processor executes the computer program, the above-mentioned wireless earphone sound quality optimization method is implemented.
[0043] In a fourth aspect, the present invention further provides a storage medium, on which a computer program is stored, and when the program is executed by a processor, the above-mentioned wireless earphone sound quality optimization method is implemented. Description of the Drawings
[0044] Figure 1 It is a flowchart of the wireless earphone sound quality optimization method in the first embodiment of the present invention;
[0045] Figure 2 It is a structural block diagram of the wireless earphone sound quality optimization system in the second embodiment of the present invention;
[0046] Figure 3 It is a structural block diagram of the electronic device in the third embodiment of the present invention.
[0047] Main Component Symbol Explanation:
[0048] 10. Acquisition module; 20. Processing module; 30. Evaluation module; 40. Enhancement module; 50. Detection module; 60. Adjustment module;
[0049] 70. Bus; 71. Processor; 72. Memory; 73. Communication interface.
[0050] The following specific embodiments will further illustrate the present invention in conjunction with the above-mentioned drawings. Specific Embodiments
[0051] For ease of understanding of the present invention, the present invention will be described more comprehensively below with reference to the relevant drawings. Several embodiments of the present invention are shown in the drawings. However, the present invention can be implemented in many different forms and is not limited to the embodiments described herein. On the contrary, these embodiments are provided to make the disclosure of the present invention more thorough and comprehensive.
[0052] It should be noted that when an element is referred to as being "fixedly provided on" another element, it can be directly on the other element or there may also be an intermediate element. When an element is considered to be "connected" to another element, it can be directly connected to the other element or there may be an intermediate element at the same time. The terms "vertical", "horizontal", "left", "right" and similar expressions used herein are for illustrative purposes only.
[0053] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the technical field to which the present invention belongs. The terms used in the description of the present invention herein are only for the purpose of describing specific embodiments and are not intended to limit the present invention. The term "and / or" used herein includes any and all combinations of one or more of the related listed items.
[0054] Embodiment 1
[0055] Please refer to Figure 1 , which shows the method for optimizing the sound quality of wireless earphones in the first embodiment of the present invention. The method includes steps S1 to S6:
[0056] S1. Obtain the sound signal in the wireless earphone;
[0057] Specifically, step S1 includes step S11:
[0058] S11. Obtain the audio information of the wireless earphone, where the audio information includes audio feature information and audio source information.
[0059] S2. Process the sound signal to obtain the Mel spectrogram of the sound signal, and segment the Mel spectrogram to obtain several Mel segments;
[0060] Specifically, step S2 includes steps S21 to S23:
[0061] S21, pre-emphasize the sound signal to enhance the high-frequency components in the sound signal, and window the sound signal to obtain a preprocessed sound signal;
[0062] S22, extract the Mel spectrum from the preprocessed sound signal;
[0063] It can be understood that pre-emphasis is a prerequisite for sound signal processing. The main purpose is to enhance the high-frequency components in the sound signal. Usually, some sound signals will have frequency components with high-frequency attenuation. To avoid the loss of high-frequency signals during the transmission of the sound signal, the high-frequency components in the sound signal are enhanced through pre-emphasis filtering, and then the Mel spectrum in the sound signal is extracted.
[0064] S23, segment the Mel spectrum with a preset width to obtain a number of Mel segments with the same width, and input the number of Mel segments with the same width into the CNN network;
[0065] It can be understood that in this embodiment, the maximum Mel frequency is set to 20 kHz to reach the maximum frequency of human ear hearing. 48 Mel filters are used to filter and logarithmically operate on the signal to obtain the Mel spectrum, so as to capture the frequency range that the human auditory system is more sensitive to.
[0066] S3, obtain the feature vector of the Mel segment based on the CNN network structure, and evaluate the sound signal based on the temporal dependence model and the feature vector to obtain the quality evaluation index of the sound signal;
[0067] Specifically, step S3 includes steps S31 to S34:
[0068] S31, extract the features of the Mel segment based on several residual structures in the CNN network structure to obtain shallow features;
[0069] S32, extract the deep features of the Mel segment based on the convolutional layer in the CNN network structure, and fuse the shallow features and the deep features to obtain a fused feature map;
[0070] S33, unfold the fused feature map to obtain a feature vector;
[0071] It can be understood that through the 3×3 small convolutional kernel in the CNN network structure, while reducing the network parameters, the perception ability of the CNN network structure for local features can be improved. Then, the features of the Mel spectrum are extracted through 5 residual structures to obtain shallow features, and the deep features of the Mel segment are extracted through the last convolutional layer in the CNN network structure. Then, the shallow features and the deep features are fused to obtain a fused feature map, and it is flattened into a feature vector with a dimension of 960.
[0072] S34, based on the Transformer network structure, and adopting a self-attention mechanism to capture the global information in the feature vector and evaluate it to obtain the quality evaluation index of the sound signal;
[0073] It can be understood that the Transformer network avoids using recurrent neural networks and convolutional layer neural networks. By using the self-attention mechanism to automatically extract features, it can effectively capture the global information between time steps at different distances in the input data, and can be trained efficiently in a parallelized manner. By evaluating the sound signal through the Transformer network structure and the feature vector, the quality evaluation index of the sound signal emitted by the wireless headset can be obtained.
[0074] S4, perform data augmentation on the Mel spectrogram to obtain an enhanced Mel spectrogram, construct an anomaly detection model, and use the Mel spectrogram as a training set to train the anomaly detection model to obtain a trained anomaly detection model;
[0075] Specifically, the step S4 includes steps S41 to S42:
[0076] S41, linearly combine the Mel spectrograms of two different sound samples to obtain a new spectrogram;
[0077] S42, and perform time-domain masking and frequency-domain masking on the new spectrogram to perform data augmentation on the Mel spectrogram to obtain an enhanced Mel spectrogram;
[0078] It can be understood that by linearly combining the spectrograms of two different sound samples, a new spectrogram is obtained. At the same time, their labels are also linearly combined to obtain the label of the new sound sample, so that a new spectrogram can be obtained. Then, by performing time-domain masking and frequency-domain masking on the new spectrogram in sequence, data augmentation can be performed on the Mel spectrogram, and then an enhanced Mel spectrogram can be obtained. Among them, time-domain masking is to randomly select several time periods in the audio and cover them to generate some new samples different from the original audio. Frequency-domain masking is to randomly select several frequency ranges and set their coefficients to 0, or randomly crop or scale the coefficients in certain frequency ranges.
[0079] It should be noted that in this embodiment, a model combining the MobileFaceNet network and the SqueezeNet is used as the anomaly detection model and trained with the Mel spectrogram as the training set.
[0080] S5. Based on the trained anomaly detection model and combined with the ArcFace loss function, detect the voice signal to obtain the anomaly category of the voice signal;
[0081] Specifically, in this embodiment, step S5 includes steps S51 to S52:
[0082] S51. Normalize the features of the voice signal based on the Softmax loss function and map them into a hypersphere to obtain normalized features;
[0083] S52. Based on the ArcFace loss function, calculate the angle between the feature vector of the voice signal and the normalized feature, and output the anomaly category of the voice signal based on the Softmax loss function;
[0084] It can be understood that the anomaly category of the voice signal includes whether there is noise, echo, and other noises in the voice.
[0085] S6. Based on the anomaly category and the quality evaluation index, formulate an adjustment strategy, and adjust the voice signal of the wireless earphone based on the adjustment strategy;
[0086] Specifically, step S6 includes steps S61 to S63:
[0087] S61. Obtain the abnormal voice signal of the wireless earphone based on the anomaly category;
[0088] S62. Based on the abnormal voice signal and according to the quality evaluation index, formulate a voice adjustment strategy to correct the voice signal, and determine whether the corrected voice signal meets the quality evaluation index;
[0089] S63. If not, continue to correct the voice signal until the corrected voice signal meets the quality evaluation index;
[0090] It can be understood that the abnormal voice signal includes whether there is echo noise in the voice signal, and the voice signal is corrected through the quality evaluation index so that the abnormal voice in the voice signal returns to the range that does not affect listening, thereby effectively improving the sound quality of the wireless earphone.
[0091] In summary, in the above embodiments of the present invention, the method for optimizing the sound quality of wireless earphones obtains a quality evaluation index by means of the feature vectors of Mel segments and evaluates them through a temporal dependence model, so as to obtain the quality judgment of the earphone sound quality. Then, by using the trained anomaly detection model and combining the ArcFace loss function to detect the sound signal, the anomaly category of the sound signal can be obtained, so as to obtain the anomaly situation of the earphone sound signal. Finally, an adjustment strategy is formulated based on the anomaly category and the quality evaluation index to adjust the sound, so that the sound signal can be adjusted according to the quality evaluation index to improve and optimize the sound quality of the wireless earphone.
[0092] Embodiment 2
[0093] Please refer to Figure 2 , which shows the wireless earphone sound quality optimization system in the second embodiment of the present invention. The system includes:
[0094] An acquisition module 10, configured to acquire a sound signal in the wireless earphone;
[0095] A processing module 20, configured to process the sound signal to obtain the Mel spectrogram of the sound signal, and segment the Mel spectrogram to obtain a plurality of Mel segments;
[0096] An evaluation module 30, configured to obtain the feature vectors of the Mel segments based on a CNN network structure, and evaluate the sound signal based on a temporal dependence model and the feature vectors to obtain a quality evaluation index of the sound signal;
[0097] An enhancement module 40, configured to perform data enhancement on the Mel spectrogram to obtain an enhanced Mel spectrogram, construct an anomaly detection model, and use the Mel spectrogram as a training set to train the anomaly detection model to obtain a trained anomaly detection model;
[0098] A detection module 50, configured to detect the sound signal based on the trained anomaly detection model and in combination with an ArcFace loss function to obtain the anomaly category of the sound signal;
[0099] An adjustment module 60, configured to formulate an adjustment strategy based on the anomaly category and the quality evaluation index, and adjust the sound signal of the wireless earphone based on the adjustment strategy.
[0100] In some alternative embodiments, the acquisition module 10 includes:
[0101] A first acquisition unit, configured to acquire the audio information of the wireless earphone, where the audio information includes audio feature information and audio source information.
[0102] In some alternative embodiments, the processing module 20 includes:
[0103] A first enhancement unit for pre-emphasizing the sound signal to enhance the high-frequency components in the sound signal and windowing the sound signal to obtain a preprocessed sound signal;
[0104] A first extraction unit for extracting the Mel spectrogram from the preprocessed sound signal;
[0105] A segmentation unit for segmenting the Mel spectrogram with a preset width to obtain a number of Mel segments of the same width, and inputting the number of Mel segments of the same width into a CNN network.
[0106] In some alternative embodiments, the evaluation module 30 includes:
[0107] A second extraction unit for extracting features of the Mel segment based on a number of residual structures in the CNN network structure to obtain shallow features;
[0108] A third extraction unit for extracting deep features of the Mel segment based on the convolutional layer in the CNN network structure and fusing the shallow features with the deep features to obtain a fused feature map;
[0109] An unfolding unit for unfolding the fused feature map to obtain a feature vector;
[0110] An evaluation unit for capturing global information in the feature vector based on the Transformer network structure and using the self-attention mechanism for evaluation to obtain a quality evaluation index of the sound signal.
[0111] In some alternative embodiments, the enhancement module 40 includes:
[0112] A combination unit for linearly combining the Mel spectrograms of two different sound samples to obtain a new spectrogram;
[0113] A second enhancement unit for performing time-domain masking and frequency-domain masking on the new spectrogram to perform data enhancement on the Mel spectrogram and obtain an enhanced Mel spectrogram.
[0114] In some alternative embodiments, the detection module 50 includes:
[0115] A normalization unit for normalizing the features of the sound signal based on the Softmax loss function and mapping them into a hypersphere to obtain normalized features;
[0116] An output unit, configured to calculate the angle between the feature vector of the sound signal and the normalized feature number based on the ArcFace loss function, and output the abnormal category of the sound signal based on the Softmax loss function.
[0117] In some alternative embodiments, the adjustment module 60 includes:
[0118] A second acquisition unit, configured to acquire the abnormal sound signal of the wireless earphone based on the abnormal category;
[0119] A formulation unit, configured to formulate a sound adjustment strategy based on the abnormal sound signal and according to the quality evaluation index to correct the sound signal, and determine whether the corrected sound signal meets the quality evaluation index;
[0120] A judgment unit, configured to continue to correct the sound signal if the sound signal does not meet the quality evaluation index until the corrected sound signal meets the quality evaluation index.
[0121] The functions or operation steps implemented when the above-mentioned modules and units are executed are substantially the same as those in the above method embodiments, and will not be elaborated herein.
[0122] The wireless earphone sound quality optimization system provided by the embodiments of the present invention has the same implementation principle and technical effects as those in the foregoing method embodiments. For the sake of brief description, for the parts not mentioned in the system embodiments, reference may be made to the corresponding content in the foregoing method embodiments.
[0123] Embodiment III
[0124] Please refer to Figure 3 , which shows the electronic device in the third embodiment of the present invention. The electronic device may include a processor 71 and a memory 72 storing computer program instructions.
[0125] Specifically, the above-mentioned processor 71 may include a central processing unit (CPU), or an application specific integrated circuit (ASIC), or one or more integrated circuits configured to implement the present application.
[0126] Among them, the memory 72 may include a mass storage for data or instructions. By way of example and not limitation, the memory 72 may include a hard disk drive (HDD), a floppy disk drive, a solid state drive (SSD), a flash memory, an optical disc, a magneto-optical disc, a magnetic tape, or a universal serial bus (USB) drive, or a combination of two or more of these. Where appropriate, the memory 72 may include removable or non-removable (or fixed) media. Where appropriate, the memory 72 may be internal or external to the data processing device. In a particular embodiment, the memory 72 is non-volatile memory. In a particular embodiment, the memory 72 includes a read-only memory (ROM) and a random access memory (RAM). Where appropriate, the ROM may be a mask-programmed ROM, a programmable ROM (PROM), an erasable PROM (EPROM), an electrically erasable PROM (EEPROM), an electrically alterable ROM (EAROM), or a flash memory, or a combination of two or more of these. Where appropriate, the RAM may be a static random access memory (SRAM) or a dynamic random access memory (DRAM), where the DRAM may be a fast page mode dynamic random access memory (FPMDRAM), an extended date out dynamic random access memory (EDODRAM), a synchronous dynamic random access memory (SDRAM), etc.
[0127] The memory 72 can be used to store or cache various data files required for processing and / or communication, as well as possible computer program instructions executed by the processor 71.
[0128] The processor 71 reads and executes the computer program instructions stored in the memory 72 to implement the wireless headphone sound quality optimization method of the first embodiment above.
[0129] In some of the embodiments, the electronic device may further include a communication interface 73 and a bus 70. Among them, as Figure 3 shown, the processor 71, the memory 72, and the communication interface 73 are connected through the bus 70 and complete communication with each other.
[0130] The communication interface 73 is used to implement communication between various modules, devices, units, and / or devices in the present application. The communication interface 73 can also implement data communication with other components such as external devices, image / data acquisition devices, databases, external storage, and image / data processing workstations.
[0131] Bus 70 includes hardware, software, or both, and couples components of a device to each other. Bus 70 includes, but is not limited to, at least one of the following: Data Bus, Address Bus, Control Bus, Expansion Bus, Local Bus. By way of example and not limitation, Bus 70 may include an Accelerated Graphics Port (AGP) or other graphics bus, an Extended Industry Standard Architecture (EISA) bus, a Front Side Bus (FSB), a Hyper Transport (HT) interconnect, an Industry Standard Architecture (ISA) bus, an InfiniBand interconnect, a Low Pin Count (LPC) bus, a memory bus, a Micro Channel Architecture (MCA) bus, a Peripheral Component Interconnect (PCI) bus, a PCI-Express (PCI-X) bus, a Serial Advanced Technology Attachment (SATA) bus, a Video Electronics Standards Association Local Bus (VLB) bus, or other suitable bus or a combination of two or more of these. In a suitable case, Bus 70 may include one or more buses. Although this application describes and illustrates specific buses, this application contemplates any suitable bus or interconnect.
[0132] The electronic device can obtain a wireless headphone sound quality optimization system and execute the wireless headphone sound quality optimization method of Embodiment 1.
[0133] In addition, in combination with the wireless headphone sound quality optimization method in Embodiment 1 above, this application can be implemented by providing a storage medium. Computer program instructions are stored on the storage medium; when the computer program instructions are executed by a processor, the wireless headphone sound quality optimization method of Embodiment 1 above is implemented.
[0134] In the description of this specification, the description referring to terms such as "one embodiment", "some embodiments", "examples", "specific examples", or "some examples", etc. means that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in a suitable manner in any one or more embodiments or examples.
[0135] The above-described embodiments merely represent several implementation manners of the present invention. The description thereof is relatively specific and detailed, but it should not be construed as a limitation to the scope of the patent of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present invention, several modifications and improvements can still be made, and these all belong to the protection scope of the present invention. Therefore, the protection scope of the patent of the present invention shall be subject to the appended claims.
Claims
1. A method for optimizing the sound quality of a wireless headset, characterized in that: The method comprises: Acquire sound signals from wireless headphones; Processing the sound signal to obtain a Mel spectrum of the sound signal, and dividing the Mel spectrum to obtain a plurality of Mel segments; Acquire a feature vector of the Mel segment based on a CNN network structure, and evaluate the sound signal based on a temporal dependency model and the feature vector to obtain a quality evaluation index of the sound signal; Performing data enhancement on the Mel spectrum to obtain an enhanced Mel spectrum, building an anomaly detection model, and training the anomaly detection model using the Mel spectrum as a training set to obtain a trained anomaly detection model; Based on the trained anomaly detection model, the sound signal is detected in combination with an ArcFace loss function to obtain an abnormal category of the sound signal; An adjustment strategy is formulated based on the abnormal category and the quality evaluation index, and the sound signal of the wireless headset is adjusted based on the adjustment strategy.
2. The wireless headset sound quality optimization method according to claim 1, characterized in that: The step of obtaining the sound signal in the wireless headset comprises: The audio information of the wireless headset is obtained, where the audio information includes audio feature information and audio source information.
3. The wireless headset sound quality optimization method according to claim 1, characterized in that: The step of processing the sound signal to obtain a Mel spectrum of the sound signal, and dividing the Mel spectrum to obtain a plurality of Mel segments comprises: Pre-emphasize the sound signal to enhance the high frequency component of the sound signal, and window the sound signal to obtain a pre-processed sound signal; Extracting the Mel spectrum from the preprocessed sound signal; The Mel spectrum is divided into a preset width to obtain a plurality of Mel segments of the same width, and the plurality of Mel segments of the same width are input into a CNN network.
4. The wireless headset sound quality optimization method according to claim 1, characterized in that: The step of obtaining the feature vector of the Mel segment based on the CNN network structure, and evaluating the sound signal based on the temporal dependency model and the feature vector to obtain the quality evaluation index of the sound signal includes: Extracting the features of the Mel segment based on a plurality of residual structures in the CNN network structure to obtain shallow features; Extracting deep features of the Mel segment based on the convolutional layer in the CNN network structure, and fusing the shallow features with the deep features to obtain a fused feature map; Expanding the fused feature map to obtain a feature vector; Based on the Transformer network structure, a self-attention mechanism is used to capture the global information in the feature vector and perform evaluation to obtain the quality evaluation index of the sound signal.
5. The wireless headset sound quality optimization method according to claim 1, characterized in that: The step of performing data enhancement on the Mel spectrum to obtain an enhanced Mel spectrum includes: Linearly combine the Mel spectrograms of two different sound samples to obtain a new spectrogram; The new spectrum graph is masked in the time domain and the frequency domain to perform data enhancement on the Mel spectrum to obtain an enhanced Mel spectrum.
6. The wireless headset sound quality optimization method according to claim 1, characterized in that: The step of detecting the sound signal based on the trained anomaly detection model and combining the ArcFace loss function to obtain the anomaly category of the sound signal includes: Normalizing the features of the sound signal based on a Softmax loss function and mapping the features into a hypersphere to obtain normalized features; The feature vector of the sound signal and the angle between the feature vectors of the normalized feature vector are calculated based on the ArcFace loss function, and the abnormal category of the sound signal is output based on the Softmax loss function.
7. The wireless headset sound quality optimization method according to claim 1, characterized in that: The step of formulating an adjustment strategy based on the abnormality category and the quality evaluation index, and adjusting the sound signal of the wireless headset based on the adjustment strategy includes: Acquire an abnormal sound signal of the wireless headset based on the abnormal category; Formulate a sound adjustment strategy based on the abnormal sound signal and according to the quality evaluation index to correct the sound signal, and determine whether the corrected sound signal meets the quality evaluation index; If not, continue to correct the sound signal until the corrected sound signal meets the quality evaluation index.
8. A wireless headset sound quality optimization system, characterized in that: The system comprises: An acquisition module, used to acquire sound signals from the wireless headset; A processing module, configured to process the sound signal to obtain a Mel spectrum of the sound signal, and to segment the Mel spectrum to obtain a plurality of Mel segments; An evaluation module, used for obtaining a feature vector of the Mel segment based on a CNN network structure, and evaluating the sound signal based on a temporal dependency model and the feature vector to obtain a quality evaluation index of the sound signal; An enhancement module is used to perform data enhancement on the Mel spectrum to obtain an enhanced Mel spectrum, construct an anomaly detection model, and train the anomaly detection model using the Mel spectrum as a training set to obtain a trained anomaly detection model; A detection module, configured to detect the sound signal based on the trained anomaly detection model and in combination with an ArcFace loss function to obtain an abnormality category of the sound signal; The adjustment module is used to formulate an adjustment strategy based on the abnormal category and the quality evaluation index, and adjust the sound signal of the wireless headset based on the adjustment strategy.
9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the method for optimizing the sound quality of a wireless headset according to any one of claims 1 to 7 is implemented.
10. A storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the wireless headset sound quality optimization method according to any one of claims 1 to 7 is implemented.