Audio device mode language control method and system
By performing energy enhancement, power spectrum analysis, and feature recognition on the voice commands of audio devices, the problems of cumbersome and inaccurate control of audio devices are solved, enabling efficient and accurate control in complex environments and improving user experience.
Patent Information
- Application Number
- CN202411302834.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-09-18
- Publication Date
- 2025-11-18
- Estimated Expiration
- 2044-09-18
AI Technical Summary
Existing audio equipment control methods are cumbersome and difficult to achieve precise control in complex environments, especially in the case of noise interference or unclear speech, resulting in inconvenient and inefficient control.
By acquiring voice commands, the system performs energy enhancement, power spectrum analysis, signal filtering, noise suppression, time-frequency adjustment, and codebook matching; identifies audio features; parses command text; and performs mode switching and parameter adjustment to ensure that audio devices accurately respond to user needs.
It improves the accuracy and convenience of audio device control, ensuring precise response to user commands even in complex environments, thus enhancing the user experience.
Smart Images

Figure CN119169996B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of voice control, and in particular to an audio device mode language control method and system. BACKGROUND
[0002] Audio devices play an important role in people's daily life and work. Building a mode language control platform for audio devices can allow users to easily switch between different modes of audio devices through voice commands, thereby improving the convenience and efficiency of using audio devices. At the same time, users can quickly adjust the parameter settings of audio devices according to different scene requirements to obtain a more personalized audio experience, and facilitate interaction and collaboration between different users, audio devices and other smart devices, thereby reducing the distress caused by complex operations and improving user satisfaction.
[0003] Currently, the control of audio devices usually adopts traditional control methods such as manual buttons, touch screen operations or remote controllers. However, through such traditional control methods, users need to perform tedious operations, and in some cases (such as environmental noise interference, unclear or unclear voice, etc.), it is difficult to achieve precise control. Moreover, due to the diverse use scenarios of audio devices, there are a large number of different mode and parameter setting requirements. The traditional control method is not efficient and accurate in data processing when facing complex operations and a large number of instructions, thereby making the control of audio devices not convenient. SUMMARY
[0004] The present application provides an audio device mode language control method and system, which mainly aims to improve the accuracy of audio device control.
[0005] To achieve the above-mentioned purpose, the present application provides an audio device mode language control method, which comprises:
[0006] Obtaining a voice command of a user corresponding to an audio device, extracting a voice audio of the voice command, performing energy enhancement on the voice audio to obtain an enhanced audio;
[0007] Analyzing the power spectrum of the enhanced audio, performing signal filtering on the enhanced audio based on the power spectrum to obtain a filtered audio, calculating the long-time frame power of the filtered audio, performing noise suppression on the filtered audio based on the long-time frame power to obtain a suppressed audio;
[0008] Querying the multi-channel time-frequency and power of the suppressed audio, performing time-frequency scale adjustment on the channel time-frequency and the power to obtain adjusted time-frequency and adjusted power, determining a time-frequency adjusted audio of the suppressed audio based on the adjusted time-frequency and adjusted power, querying the best codebook of the time-frequency adjusted audio, and identifying the audio feature of the suppressed audio based on the best codebook.
[0009] resolve an instruction text corresponding to the voice instruction based on the audio feature, perform mode switching on the audio device based on the instruction text, and obtain a target audio device;
[0010] analyze an audio effect of the target audio device;
[0011] if the audio effect does not meet a preset user demand effect, query a user feedback instruction of the target audio device, perform parameter adjustment on the target audio device based on the user feedback instruction, and return to the step of analyzing the audio effect of the target audio device;
[0012] if the audio effect meets the preset user demand effect, take the target audio device as an audio control device of the voice instruction.
[0013] Optionally, the energy enhancement on the voice audio to obtain enhanced audio includes:
[0014] perform frequency domain conversion on the voice audio to obtain converted audio;
[0015] calculate energy levels of different frequency intervals in the converted audio to obtain regional audio energy levels;
[0016] determine an energy enhancement interval of the converted audio based on the regional audio energy levels;
[0017] perform energy enhancement on the converted audio based on the energy enhancement interval to obtain enhanced audio.
[0018] Optionally, the analysis of the power spectrum of the enhanced audio includes:
[0019] extract an initial signal of the enhanced audio;
[0020] calculate a wavelet coefficient spectrum of the initial signal;
[0021] calculate the power spectrum of the enhanced audio based on the wavelet coefficient spectrum.
[0022] Optionally, the signal filtering on the enhanced audio based on the power spectrum to obtain filtered audio includes:
[0023] determine a filtering range of the enhanced audio based on the power spectrum;
[0024] calculate a signal pass degree of the enhanced audio based on the filtering range;
[0025] perform signal transformation on a signal of the enhanced audio based on the signal pass degree to obtain filtered audio.
[0026] Optionally, the calculating the long-time frame power of the filtered audio comprises:
[0027] performing audio segmentation on the filtered audio to obtain segmented audio;
[0028] calculating signal power of the segmented audio;
[0029] calculating the long-time frame power of the filtered audio based on the frame power.
[0030] Optionally, the querying the optimal codebook of the time-frequency adjusted audio comprises:
[0031] converting the time-frequency adjusted audio into an audio vector;
[0032] performing vector clustering on the audio vector to obtain a clustered vector;
[0033] querying a cluster center of the clustered vector;
[0034] performing cluster optimization on the clustered vector based on the cluster center to obtain an optimized clustered vector;
[0035] performing screening on the optimized clustered vector to obtain a screened clustered vector;
[0036] determining the optimal codebook of the time-frequency adjusted audio based on the screened clustered vector.
[0037] Optionally, the converting the time-frequency adjusted audio into an audio vector comprises:
[0038] performing frame processing on the time-frequency adjusted audio to obtain a plurality of frames of audio;
[0039] performing frequency scale conversion on the plurality of frames of audio to obtain converted audio;
[0040] performing discrete cosine transform on the converted audio to obtain an audio vector.
[0041] Optionally, the identifying the audio feature of the suppression audio based on the optimal codebook comprises:
[0042] querying a feature parameter of the suppression audio;
[0043] performing vectorization processing on the feature parameter to obtain a parameter vector;
[0044] matching the parameter vector with the optimal codebook to obtain a matching result;
[0045] based on the matching result, performing time-frequency domain feature labeling on the suppression audio to determine the audio feature of the suppression audio.
[0046] Optionally, the analyzing the audio effect of the target audio device comprises:
[0047] inquiring a control parameter corresponding to the target audio device;
[0048] performing an audio effect test on the target audio device according to the control parameter;
[0049] after the audio effect test ends, performing intelligent voice interaction with a user corresponding to the target audio device;
[0050] determining the audio effect of the target audio device based on an interaction result of the intelligent voice interaction.
[0051] To solve the above problems, the present application also provides an audio device mode language control system, which comprises:
[0052] an audio enhancement module configured to acquire a voice instruction of a user corresponding to an audio device, extract voice audio of the voice instruction, perform energy enhancement on the voice audio, and obtain enhanced audio;
[0053] an audio suppression module configured to analyze a power spectrum of the enhanced audio, perform signal filtering on the enhanced audio based on the power spectrum, obtain filtered audio, calculate long-time frame power of the filtered audio, perform noise suppression on the filtered audio based on the long-time frame power, and obtain suppressed audio;
[0054] an audio feature recognition module configured to inquire multi-channel time-frequency and power of the suppressed audio, perform time-frequency scale adjustment on the channel time-frequency and the power, obtain adjusted time-frequency and adjusted power, determine time-frequency adjusted audio of the suppressed audio based on the adjusted time-frequency and the adjusted power, inquire an optimal codebook of the time-frequency adjusted audio, and recognize an audio feature of the suppressed audio based on the optimal codebook;
[0055] a device mode switching module configured to parse an instruction text corresponding to the voice instruction based on the audio feature, perform mode switching on the audio device based on the instruction text, and obtain a target audio device;
[0056] an audio effect analysis module configured to analyze an audio effect of the target audio device;
[0057] a device adjustment module configured to, if the audio effect does not meet a preset user demand effect, inquire a user feedback instruction of the target audio device, perform parameter adjustment on the target audio device based on the user feedback instruction, and return to perform the step of analyzing the audio effect of the target audio device;
[0058] The target device acquisition module is configured to acquire the target audio device as an audio control device of the voice instruction if the audio effect meets a preset user demand effect.
[0059] The application can make the voice signal more prominent and improve the accuracy and reliability of subsequent processing by enhancing the energy of the voice audio to obtain enhanced audio. The application can understand the energy distribution of the audio signal at different frequencies by analyzing the power spectrum of the enhanced audio, so as to better remove audio noise. The application can more accurately capture the change characteristics of the audio signal of the suppression audio in time and frequency by querying the multi-channel time-frequency and power of the suppression audio. The application can determine the instruction issued by the user by analyzing the instruction text corresponding to the voice instruction based on the audio features, so as to accurately control the audio device. Therefore, the application can improve the accuracy of audio device control. BRIEF DESCRIPTION OF DRAWINGS
[0060] Figure 1 A flowchart of an audio device mode language control method provided by an embodiment of the application is shown in the figure.
[0061] Figure 2 A function module diagram of an audio device mode language control system provided by an embodiment of the application is shown in the figure.
[0062] The implementation, functional features and advantages of the application will be further described with reference to the embodiments and the accompanying drawings. DETAILED DESCRIPTION
[0063] It should be understood that the specific embodiments described herein are only used to explain the application and not to limit the application.
[0064] An audio device mode language control method is provided in the embodiments of the application. The execution subject of the audio device mode language control method includes but is not limited to at least one of electronic devices such as a server and a terminal, which can be configured to execute the method provided in the embodiments of the application. In other words, the audio device mode language control method can be executed by software or hardware installed in a terminal device or a server device, and the software can be a blockchain platform. The server includes but is not limited to a single server, a server cluster, a cloud server or a cloud server cluster, etc. The server can be a stand-alone server, or a cloud server providing cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content distribution networks (CDN), and basic cloud computing services such as big data and artificial intelligence platforms.
[0065] REFERENCE Figure 1As shown, a flowchart of an audio device mode language control method provided by an embodiment of the present application is shown. In this embodiment, the audio device mode language control method includes:
[0066] S1, obtaining a voice instruction of a user corresponding to an audio device, extracting voice audio of the voice instruction, performing energy enhancement on the voice audio to obtain enhanced audio.
[0067] In this embodiment, the audio device is a hardware device capable of receiving, processing and outputting audio signals, and the voice instruction is a specific command sentence expressed by voice. The audio device can execute corresponding operations by identifying these instructions. The voice instruction can be obtained by receiving a voice signal of a user through an audio input device (such as a microphone) in the audio device.
[0068] Further, in this embodiment, the voice audio can be obtained by focusing on the voice signal of the user and excluding other irrelevant background noise interference.
[0069] The voice audio is a signal of voice, which can be obtained by filtering the received audio signal to separate the voice part of the user.
[0070] Further, in this embodiment, the enhanced audio can be obtained by performing energy enhancement on the voice audio, which can make the voice signal more prominent and improve the accuracy and reliability of subsequent processing.
[0071] As an embodiment of the present application, the energy enhancement on the voice audio to obtain the enhanced audio includes: performing frequency domain conversion on the voice audio to obtain converted audio, calculating energy levels of different frequency intervals in the converted audio to obtain regional audio energy levels, determining an energy enhancement interval of the converted audio based on the regional audio energy levels, and performing energy enhancement on the converted audio based on the energy enhancement interval to obtain the enhanced audio.
[0072] Optionally, the converted audio can be obtained by performing frequency domain conversion on the speech audio through fast Fourier transform, the energy level of the different frequency intervals can be determined by calculating the square sum of the audio signal corresponding to the speech audio, the energy enhancement interval can be obtained by comparing the energy of different intervals with a preset energy threshold, and the energy of the different intervals needs to be enhanced if the energy of the different intervals is less than the energy threshold, the preset energy threshold can be set to 0.5J, or a suitable value can be set according to an actual application scenario, the energy of the converted audio is enhanced to obtain enhanced audio through a spectrum enhancement method, that is, the energy of a specific frequency interval is enhanced by processing the audio signal in the frequency domain.
[0073] S2, analyze the power spectrum of the enhanced audio, perform signal filtering on the enhanced audio based on the power spectrum to obtain filtered audio, calculate the long-time frame power of the filtered audio, and perform noise suppression on the filtered audio based on the long-time frame power to obtain suppressed audio.
[0074] The embodiment of the present application can understand the energy distribution of the audio signal at different frequencies by analyzing the power spectrum of the enhanced audio, so as to better remove audio noise. The power spectrum refers to a representation method for describing the distribution of signal power at different frequencies.
[0075] As an embodiment of the present application, the analysis of the power spectrum of the enhanced audio includes: extracting an initial signal of the enhanced audio, calculating the wavelet coefficient spectrum of the initial signal, and calculating the power spectrum of the enhanced audio based on the wavelet coefficient spectrum using the following formula:
[0076]
[0077] In the formula, P(ω) represents the power spectrum, ω represents the wavelet coefficient spectrum, E represents the time-frequency conversion value, and T represents the total number of frames of the enhanced audio.
[0078] Optionally, the wavelet coefficient spectrum of the initial signal can be calculated through a wavelet function.
[0079] Further, the embodiment of the present application can remove unnecessary frequency components and reduce the amount of audio data by performing signal filtering on the enhanced audio based on the power spectrum to obtain filtered audio, thereby reducing the computational complexity and storage requirements of subsequent processing.
[0080] As an embodiment of the present application, the signal filtering on the enhanced audio based on the power spectrum to obtain filtered audio includes: determining a filtering range of the enhanced audio based on the power spectrum, and calculating the signal pass of the enhanced audio based on the filtering range using the following formula:
[0081]
[0082] wherein H(f) represents a signal pass, f represents one frequency in the enhanced audio, f l represents the minimum value of the filtering range, f h represents the maximum value of the filtering range;
[0083] performing signal transformation on the signal of the enhanced audio based on the signal pass to obtain filtered audio.
[0084] wherein the signal pass refers to the pass degree of different frequency components of an input signal in a filter, when H(f) = 1, it means that the signal at frequency f is completely passed through the filter; when H(f) = 0, it means that the signal at the frequency is completely blocked.
[0085] Optionally, the filtering range of the enhanced audio can be determined by analyzing the peak and valley of the power spectrum, for example, when the power spectrum has an obvious peak at a certain frequency, it may mean that the frequency corresponds to the fundamental frequency or an important harmonic component of the signal, for a periodic signal, the peak in the power spectrum will appear at the integer multiple of the period frequency, the period characteristics of the signal can be determined according to these peaks, and the relevant frequency components are retained, and the signal transformation on the signal of the enhanced audio based on the signal pass to obtain filtered audio can be performed by performing Fourier transform on the signal of the enhanced audio corresponding to the signal pass to obtain a filtered transform signal, and then performing inverse Fourier transform on the filtered transform signal to obtain filtered audio.
[0086] The long-time frame power of the filtered audio calculated by the embodiment of the application can help the user to analyze the noise characteristics in the audio, and by observing the changes of the long-time frame power in different time periods, it can be judged whether the intensity of the noise is stable and the frequency distribution of the noise.
[0087] wherein the long-time frame power refers to a comprehensive index obtained by statistically processing the power of a plurality of continuous frame audio signals.
[0088] As an embodiment of the application, the calculation of the long-time frame power of the filtered audio comprises: performing audio segmentation on the filtered audio to obtain segmented audio, and calculating the signal power of the segmented audio by using the following formula:
[0089]
[0090] wherein P framerepresents signal power, N represents frame length of segmented audio, w(n) represents weighting function, x(n) represents value of audio signal corresponding to segmented audio at n-th sampling point;
[0091] Based on the frame power, the long-time frame power of the filtered audio is calculated by using the following formula:
[0092]
[0093] wherein, P long-term represents long-time frame power, M represents total frame number of filtered audio, P frame (i) represents i-th frame of filtered audio.
[0094] Further, the embodiment of the present application can improve signal processing efficiency and optimize performance of the audio processing system by performing noise suppression on the filtered audio based on the long-time frame power to obtain suppressed audio.
[0095] Optionally, the noise suppression on the suppressed audio can be performed by using the long-time frame power in combination with an asymmetric filter function.
[0096] S3, a multi-channel time-frequency and power of the suppressed audio are queried, time-frequency scale adjustment is performed on the channel time-frequency and the power to obtain adjusted time-frequency and adjusted power, time-frequency adjusted audio of the suppressed audio is determined based on the adjusted time-frequency and the adjusted power, an optimal codebook of the time-frequency adjusted audio is queried, and an audio feature of the suppressed audio is identified based on the optimal codebook.
[0097] The embodiment of the present application can more accurately capture the change characteristics of the audio signal of the suppressed audio in time and frequency by querying the multi-channel time-frequency and power of the suppressed audio.
[0098] The multi-channel time-frequency refers to the change of the audio signal in multiple channels over time and frequency, and the power refers to the energy intensity of the audio signal in a specific time period.
[0099] Further, the embodiment of the present application can make the features of the audio signal more obvious and facilitate feature extraction and analysis by performing time-frequency scale adjustment on the channel time-frequency and the power to obtain adjusted time-frequency and adjusted power.
[0100] Optionally, the time-frequency scale adjustment on the channel time-frequency and the power to obtain adjusted time-frequency and adjusted power can be performed by resampling method, and the resampling parameters are determined according to the required time-frequency scale adjustment. The resampling parameters include change ratio of sampling rate, filter type, etc.
[0101] Further, the embodiment of the present application can quickly find the code word that best matches the data by querying the optimal codebook of the time-frequency adjusted audio, thereby realizing effective processing and division of the data.
[0102] The codebook refers to a set used for encoding or representing data. In audio processing, the codebook is usually composed of specific code words that can be used to represent the characteristics or patterns of the audio.
[0103] As an embodiment of the present application, the querying of the optimal codebook of the time-frequency adjusted audio comprises: converting the time-frequency adjusted audio into an audio vector, performing vector clustering on the audio vector to obtain a clustered vector, querying a cluster center of the clustered vector, performing cluster optimization on the clustered vector based on the cluster center to obtain an optimized clustered vector, performing screening on the optimized clustered vector to obtain a screened clustered vector, and determining the optimal codebook of the time-frequency adjusted audio based on the screened clustered vector.
[0104] Optionally, the vector clustering of the audio vector can be realized by the K-Means method, for example, K initial cluster centers are randomly selected, the distance of each audio vector to each cluster center is calculated, the audio vector is assigned to the cluster where the nearest cluster center is located, then the center of each cluster is recalculated, and the process is repeated until the cluster center no longer changes or a certain number of iterations is reached. It should be noted that after the vector clustering is completed, each cluster has a center, which can be obtained by calculating the mean of all audio vectors in the cluster. For example, using K-Means clustering, the center of each cluster is the K points determined in the last iteration. The cluster optimization of the clustered vector based on the cluster center to obtain the optimized clustered vector can be realized by further refining the initial clustering result by fuzzy C-means clustering (FCM). The screening of the optimized clustered vector to obtain the screened clustered vector can be realized by calculating the distance of each audio vector to the center of the cluster to which it belongs. If the distance exceeds a preset threshold, the audio vector is considered to be an outlier and is removed. It should be noted that the preset threshold needs to be set in combination with actual application operation. The optimal codebook can be obtained by taking the center of the screened clustered vector as the code word of the codebook. These code words constitute the optimal codebook.
[0105] Further, the embodiment of the present application can accurately classify different types of suppression audio by identifying the audio features of the suppression audio based on the optimal codebook. For example, the audio can be classified into different categories such as speech, music, environmental sound, etc. For a speech recognition system, the identity of the speaker and the language content can be identified. For example, in an intelligent voice assistant application, the user's voice instruction can be accurately identified to better serve the user.
[0106] As an optional embodiment of the present application, the converting the time-frequency adjusted audio into an audio vector comprises: performing frame processing on the time-frequency adjusted audio to obtain a plurality of frames of audio, performing frequency scale conversion on the plurality of frames of audio to obtain converted audio, and performing discrete cosine transformation on the converted audio to obtain an audio vector.
[0107] Optionally, the frame processing on the time-frequency adjusted audio to obtain a plurality of frames of audio can be implemented by a window function, and the frequency scale conversion on the plurality of frames of audio to obtain converted audio can be implemented by a mel frequency scale conversion.
[0108] As an embodiment of the present application, the identifying the audio feature of the suppression audio based on the optimal codebook comprises: querying a feature parameter of the suppression audio, performing vectorization processing on the feature parameter to obtain a parameter vector, matching the parameter vector with the optimal codebook to obtain a matching result, and based on the matching result, performing time-frequency domain feature labeling on the suppression audio to determine the audio feature of the suppression audio. The feature parameter refers to a numerical value used to describe a specific attribute of an audio signal. These parameters can reflect the characteristics of the audio from different angles and provide a basis for audio analysis, identification, classification and other tasks.
[0109] Optionally, the feature parameter can be obtained by a mel frequency cepstrum coefficient, the vectorization processing on the feature parameter to obtain a parameter vector can be implemented by a vector matrix, the matching of the parameter vector with the optimal codebook to obtain a matching result can be implemented by using a distance measurement method, such as a Euclidean distance, a cosine distance and the like, and based on the matching result, the time-frequency domain feature labeling on the suppression audio to determine the audio feature of the suppression audio can be implemented according to the code word obtained by the matching to determine the time-frequency domain feature of the suppression audio. It needs to be explained that each code word represents a specific audio feature, and a specific time-frequency domain feature label can be assigned to each code word in the codebook in advance. For example, a code word in the codebook represents a high-frequency speech feature, and when the parameter vector matches this code word, the suppression audio can be labeled with a high-frequency speech feature label. By matching and labeling different parts of the suppression audio, the overall audio feature of the suppression audio can be determined.
[0110] S4, based on the audio feature, analyzing instruction text corresponding to the voice instruction, and based on the instruction text, performing mode switching on the audio device to obtain a target audio device.
[0111] By analyzing the instruction text corresponding to the voice instruction based on the audio feature, the embodiment of the present application can clearly determine the instruction issued by the user to accurately control the audio device.
[0112] Optionally, the instruction text can be translated by using a convolutional neural network to perform feature recognition on the audio features.
[0113] Further, the embodiment of the present application can improve the accuracy and reliability of instruction execution by switching the mode of the audio device based on the instruction text to obtain a target audio device, and ensure that the audio device switches the mode according to the real intention of the user.
[0114] As an embodiment of the present application, the mode of the audio device is switched based on the instruction text to obtain a target audio device, which includes: performing natural language processing on the instruction text to obtain a processed text, performing intention recognition on the instruction text based on the processed text, and calculating the intention score of the intention recognition using the following formula:
[0115]
[0116] wherein Score M (T) represents the intention score, w represents the keyword of the instruction text, T represents the instruction text, M represents a specific mode of the audio device, K M represents the keyword set of mode M, and Weight(w) represents the weight of w.
[0117] The mode of the audio device is switched based on the intention score to obtain a target audio device. Optionally, the mode of the audio device is switched based on the specific score value of the intention score to obtain a target audio device. For example, there is a smart audio device that supports Bluetooth mode, AUX mode and Wi-Fi music playing mode. For the Bluetooth mode, the keyword set K 蓝牙 = {"Bluetooth", "Connect Bluetooth", "Pair Bluetooth", "Switch to Bluetooth"}, and the weights can be set as follows: the weight of "Bluetooth" is 0.5, the weight of "Connect Bluetooth" is 0.8, the weight of "Pair Bluetooth" is 0.7, the weight of "Switch to Bluetooth" is 0.9, the weight of "AUX" is 0.6, the weight of "Connect AUX" is 0.8, and the weight of "Switch to AUX" is 0.85. For the Wi-Fi music playing mode, the keyword set K Wi-Fi = {"W i-Fi"play music", "connect Wi-Fi", "switch to Wi-Fi music", and the weights are set as: the weight of "Wi-Fi" is 0.4, the weight of "play music" is 0.6, the weight of "connect Wi-Fi" is 0.7, and the weight of "switch to Wi-Fi music" is 0.8, assuming that the user issues the instruction "switch to Bluetooth mode", the instruction text T1 is "switch to Bluetooth mode". One of the keyword set, and the weight is 0.9), the intent score of the AUX mode: Score AUX (T1) = 0 (the instruction text does not contain the keyword of the AUX mode), the intent score of the Wi-Fi music playing mode: Score Wi-Fi (T1) = 0 (the instruction text does not contain the keyword of the Wi-Fi music playing mode), assuming that the user issues the instruction "I want to listen to music, and play through Wi-Fi connection", the instruction text T2 is "I want to listen to music, and play through Wi-Fi connection", the intent score of the Bluetooth mode: Score 蓝牙 (T2) = 0, the intent score of the AUX mode: Score AUX (T2) = 0, the intent score of the Wi-Fi music playing mode: Score Wi-Fi (T2) = 0.4 + 0.6 + 0.7 + 0.8 = 2.5 (the instruction text contains the keywords of "Wi-Fi", "play music", "connect Wi-Fi", and "switch to Wi-Fi music", and the weights are added respectively), according to the calculated intent score, the smart speaker makes mode switching decision,
[0118] For instruction T1, since the intent score of the Bluetooth mode is the highest (0.9), and the intent scores of the AUX mode and the Wi-Fi music playing mode are both 0, the smart speaker switches to the Bluetooth mode, for instruction T2, the intent score of the Wi-Fi music playing mode is the highest (2.5), and the intent scores of the Bluetooth mode and the AUX mode are both 0, so the smart speaker switches to the Wi-Fi music playing mode and starts playing music through the Wi-Fi connection.
[0119] S5, analyzing the audio effect of the target audio device.
[0120] The embodiment of the present application can understand the accuracy of the execution of the audio device for the user instruction by analyzing the audio effect of the target audio device.
[0121] As an embodiment of the present application, the analyzing the audio effect of the target audio device comprises: querying the control parameters corresponding to the target audio device, performing an audio effect test on the target audio device according to the control parameters, performing intelligent voice interaction with the user corresponding to the target audio device after the audio effect test ends, and determining the audio effect of the target audio device based on the interaction result of the intelligent voice interaction.
[0122] The control parameters refer to a series of specific numerical values or setting options for adjusting and controlling the performance and audio output effect of the target audio device, such as volume-related parameters, sound effect mode parameters, and audio input and output parameters.
[0123] Optionally, the control parameters can be obtained by querying the user instructions in the instruction text, the audio effect test on the target audio device according to the control parameters can be performed by adjusting the corresponding settings of the target audio device according to the control parameters, the intelligent voice interaction with the user corresponding to the target audio device can be performed through the voice interaction system of the audio device, and the audio effect of the target audio device can be determined based on the interaction result of the intelligent voice interaction by recognizing the feedback information of the user, such as the user indicating "ok" or "can" indicating that the audio effect is fine, and the user indicating "small volume" or "change sound channel mode" indicating that the audio effect has a problem.
[0124] S6, if the audio effect does not meet the preset user demand effect, querying the user feedback instructions of the target audio device, and after adjusting the parameters of the target audio device based on the user feedback instructions, returning to the step of analyzing the audio effect of the target audio device.
[0125] The embodiment of the present application can adjust the settings of the audio device to meet the user's demand if the audio effect does not meet the preset user demand effect, so as to achieve the user's satisfaction. The user feedback instructions can be obtained by interacting with the user through the voice interaction system of the audio device.
[0126] It should be noted that if the audio effect does not meet the preset user demand effect, it means that the audio effect of the audio device does not meet the user's requirements, such as the volume, sound channel, and playback mode settings not meeting the user's demand.
[0127] Further, the embodiment of the present application returns to execute the step of analyzing the audio effect of the target audio device after adjusting the parameters of the target audio device based on the user feedback instruction, which can gradually adjust the control parameters of the audio device until the control requirements of the user for the audio device are met.
[0128] S7, if the audio effect meets the preset user demand effect, the target audio device is taken as the audio control device of the voice instruction.
[0129] It should be noted that if the audio effect meets the preset user demand effect, it means that the audio effect of the audio device meets the requirements of the user, such as volume, channel, and playing mode, which meet the user's demand, and the audio control device can be used to execute the user's audio control task.
[0130] As shown in Figure 2 is a functional module diagram of an audio device mode language control system provided by an embodiment of the present application.
[0131] The audio device mode language control system 200 can be installed in an electronic device. According to the functions to be implemented, the audio device mode language control system 200 can include a data acquisition module 201, a data class center identification module 202, a data filtering module 203, and a data storage module 204. The modules of the present application can also be referred to as units, which refer to a series of computer program segments that can be executed by an electronic device processor and can complete a fixed function, which are stored in the memory of the electronic device.
[0132] In the present embodiment, the functions of each module / unit are as follows:
[0133] The audio enhancement module 201 is configured to obtain a voice instruction of an audio device corresponding to a user, extract a voice audio of the voice instruction, perform energy enhancement on the voice audio, and obtain enhanced audio;
[0134] The audio suppression module 202 is configured to analyze a power spectrum of the enhanced audio, perform signal filtering on the enhanced audio based on the power spectrum, obtain filtered audio, calculate a long-time frame power of the filtered audio, perform noise suppression on the filtered audio based on the long-time frame power, and obtain suppressed audio;
[0135] The audio feature recognition module 203 is configured to query a multi-channel time-frequency and power of the suppressed audio, perform time-frequency scale adjustment on the channel time-frequency and the power, obtain adjusted time-frequency and adjusted power, determine a time-frequency adjusted audio of the suppressed audio based on the adjusted time-frequency and the adjusted power, query an optimal codebook of the time-frequency adjusted audio, and identify an audio feature of the suppressed audio based on the optimal codebook.
[0136] The device mode switching module 204 is configured to parse instruction text corresponding to the voice instruction based on the audio feature, and perform mode switching on the audio device based on the instruction text to obtain a target audio device.
[0137] The audio effect analysis module 205 is configured to analyze an audio effect of the target audio device.
[0138] The device adjustment module 206 is configured to, if the audio effect does not meet a preset user demand effect, query a user feedback instruction of the target audio device, and after performing parameter adjustment on the target audio device based on the user feedback instruction, return to the step of analyzing the audio effect of the target audio device.
[0139] The target device acquisition module 207 is configured to, if the audio effect meets the preset user demand effect, take the target audio device as an audio control device of the voice instruction.
[0140] In detail, each module in the audio device mode language control system 200 in the embodiment of the present application adopts the same technical means as the audio device mode language control method in the accompanying drawings when used, and can produce the same technical effects, which will not be described here.
[0141] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application and not to limit it. Although the present application has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solutions of the present application can be modified or replaced by equivalents without departing from the spirit and scope of the technical solutions of the present application.
Claims
1. A method for controlling audio device mode language, characterized in that, The methods include: S1. Obtain the voice command corresponding to the user of the audio device, extract the voice audio of the voice command, and perform energy enhancement on the voice audio to obtain enhanced audio; S2. Analyze the power spectrum of the enhanced audio. Based on the power spectrum, perform signal filtering on the enhanced audio to obtain the filtered audio. Calculate the long-time frame power of the filtered audio. Based on the long-time frame power, perform noise suppression on the filtered audio to obtain the suppressed audio. S3. Query the multi-channel time-frequency and power of the suppressed audio, perform time-frequency scaling on the channel time-frequency and power to obtain the adjusted time-frequency and adjusted power, determine the time-frequency adjusted audio of the suppressed audio based on the adjusted time-frequency and adjusted power, query the optimal codebook for the time-frequency adjusted audio, and identify the audio features of the suppressed audio based on the optimal codebook; wherein, querying the optimal codebook for the time-frequency adjusted audio includes: The time-frequency adjusted audio is converted into audio vectors; vector clustering is then performed on the audio vectors to obtain cluster vectors; vector clustering of audio vectors is implemented using the K-Means method: first, K initial cluster centers are randomly selected; then, the distance from each audio vector to each cluster center is calculated, and the audio vector is assigned to the cluster containing the nearest cluster center. Next, the center of each cluster is recalculated, and this process is repeated until the cluster centers no longer change or a certain number of iterations are reached. After vector clustering is completed, the centers of each cluster are the K points determined in the last iteration. Based on the cluster centers, cluster optimization is performed on the cluster vectors using fuzzy C-Means. Mean clustering yields optimized cluster vectors from the initial clustering results. These optimized vectors are then filtered to obtain filtered cluster vectors. This is achieved by calculating the distance from each audio vector to its cluster center; if the distance exceeds a preset threshold, the audio vector is considered an outlier and removed. The cluster centers of the selected audio vectors are then queried. Based on these centers, the cluster vectors are further optimized to obtain optimized cluster vectors. These optimized vectors are then filtered to obtain filtered cluster vectors. Finally, based on these filtered cluster vectors, the optimal codebook for time-frequency adjusted audio is determined. Based on the best codebook, identify the audio features of suppressed audio, including: Query the characteristic parameters of suppressed audio; The feature parameters are vectorized to obtain a parameter vector; The parameter vector is matched with the best codebook to obtain the matching result; Based on the matching results, the suppressed audio is labeled with time-frequency domain features to determine the audio features of the suppressed audio; S4. Based on audio features, parse the instruction text corresponding to the voice command; based on the instruction text, switch the mode of the audio device to obtain the target audio device; perform natural language processing on the instruction text to obtain processed text; based on the processed text, perform intent recognition on the instruction text; and based on the intent score, switch the mode of the audio device to obtain the target audio device. S5. Analyze the audio effects of the target audio device; S6. If the audio effect does not meet the preset user requirements, query the user feedback instructions of the target audio device, adjust the parameters of the target audio device based on the user feedback instructions, and then return to the step of analyzing the audio effect of the target audio device. S7. If the audio effect meets the preset user requirements, use the target audio device as the audio control device for voice commands.
2. The audio device pattern language control method as described in claim 1, characterized in that, Energy enhancement is performed on the speech audio to obtain enhanced audio, including: The speech audio is converted to the frequency domain to obtain the converted audio. Calculate the energy levels of different frequency ranges in the converted audio to obtain the regional audio energy levels; Based on the regional audio energy level, determine the energy enhancement range of the converted audio; The converted audio is enhanced by applying energy to the energy enhancement range to obtain enhanced audio.
3. The audio device pattern language control method as described in claim 1, characterized in that, Analyze the power spectrum of the enhanced audio, including: Extract the initial signal for enhanced audio; Calculate the wavelet coefficient spectrum of the initial signal; The power spectrum of the enhanced audio is calculated based on the wavelet coefficient spectrum.
4. The audio device pattern language control method as described in claim 1, characterized in that, Based on the power spectrum, the enhanced audio is filtered to obtain the filtered audio, including: Based on the power spectrum, determine the filtering range for enhanced audio; Calculate the signal throughput of the enhanced audio based on the filtering range; Based on signal throughput, the enhanced audio signal is transformed to obtain filtered audio.
5. The audio device pattern language control method as described in claim 1, characterized in that, Calculating the long-time frame power of the filtered audio includes: The filtered audio is segmented to obtain the segmented audio; Calculate the signal power of the segmented audio; The long-time frame power of the filtered audio is calculated based on the frame power.
6. The audio device pattern language control method as described in claim 1, characterized in that, Converting time-frequency adjusted audio into audio vectors includes: The time-frequency adjusted audio is processed into frames to obtain multiple audio frames. Frequency scaling is performed on multiple audio frames to obtain the converted audio; Perform a discrete cosine transform on the converted audio to obtain the audio vector.
7. The audio device pattern language control method as described in claim 1, characterized in that, Analyze the audio effects of the target audio device, including: Query the control parameters corresponding to the target audio device; Based on the control parameters, the audio effect of the target audio device is tested; After the audio effect test is completed, intelligent voice interaction is performed with the user corresponding to the target audio device. Based on the interaction results of intelligent voice interaction, determine the audio effect of the target audio device.
8. An audio device pattern language control system, characterized in that, The system is used to execute an audio device mode language control method as described in any one of claims 1-7, and includes: The audio enhancement module is used to acquire the voice commands of the user corresponding to the audio device, extract the voice audio of the voice commands, and enhance the energy of the voice audio to obtain enhanced audio; The audio suppression module is used to analyze the power spectrum of the enhanced audio, filter the enhanced audio based on the power spectrum to obtain the filtered audio, calculate the long-time frame power of the filtered audio, and suppress noise based on the long-time frame power to obtain the suppressed audio. The audio feature recognition module is used to query the multi-channel time-frequency and power of suppressed audio, adjust the time-frequency and power of the channels to obtain the adjusted time-frequency and adjusted power, determine the time-frequency adjusted audio of suppressed audio based on the adjusted time-frequency and adjusted power, query the best codebook of the time-frequency adjusted audio, and identify the audio features of suppressed audio based on the best codebook. The device mode switching module is used to parse the command text corresponding to the voice command based on audio features, and switch the audio device mode based on the command text to obtain the target audio device. The audio effect analysis module is used to analyze the audio effect of the target audio device; The device adjustment module is used to query the user feedback instructions of the target audio device if the audio effect does not meet the preset user needs. Based on the user feedback instructions, the parameters of the target audio device are adjusted, and then the process returns to analyze the audio effect of the target audio device. The target device acquisition module is used to select the target audio device as the audio control device for voice commands if the audio effect meets the preset user requirements.
Citation Information
Patent Citations
Intelligent voice recognition control system and method based on Internet of Things
CN118449800A