Bluetooth speaker control method and system based on AI voice analysis
Through AI voice analysis, the wake-up audio and control audio attributes of Bluetooth speakers are identified and the volume is automatically adjusted, which solves the problem that the volume adjustment of Bluetooth speakers is not suitable for user needs, and improves the interaction quality and user experience.
Patent Information
- Application Number
- CN202510293299.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-13
- Publication Date
- 2025-09-02
- Estimated Expiration
- 2045-03-13
AI Technical Summary
The volume adjustment of existing Bluetooth speakers is difficult to meet user needs, resulting in too high or too low volume at different distances, affecting the interactive effect.
Through AI voice analysis, the audio attributes of wake-up audio and control audio are identified, including direction and distance, automatically adjust the feedback volume, and generate control instructions based on the control audio to achieve real-time adjustment of volume.
Improve the interaction quality between Bluetooth speakers and users, ensure that the volume adapts to user needs and improves user experience, especially in different dialect areas and mobile scenarios.
Smart Images

Figure CN119811369B_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the field of Bluetooth speaker control and involves AI voice recognition and analysis technology, specifically a Bluetooth speaker control method and system based on AI voice analysis. Background Art
[0002] A Bluetooth speaker is an audio device that uses a built-in Bluetooth chip, replacing traditional wired connections with Bluetooth connectivity. Bluetooth speakers can provide audio playback, Q&A, and local device control through control commands or voice interaction. Their compact size and rich functionality are increasingly gaining popularity among consumers.
[0003] When used indoors, Bluetooth speakers are typically placed in a fixed location. Their primary function is to interact with the user and help them control smart home devices. For both interaction and control, Bluetooth speakers should provide users with timely feedback. However, current Bluetooth speaker volume adjustment is primarily manual or smartphone-based. Furthermore, due to the relatively fixed location of Bluetooth speakers, it's difficult to adapt the playback volume to user needs, severely impacting the interactive nature of Bluetooth speakers.
[0004] This application provides a Bluetooth speaker control method and system based on AI voice analysis to solve the above technical problems. Summary of the Invention
[0005] The present application aims to solve at least one of the technical problems existing in the prior art; to this end, the present application proposes a Bluetooth speaker control method and system based on AI voice analysis, which is used to solve the technical problem that the volume of the Bluetooth speaker is difficult to adapt to user needs by manually adjusting the volume of the prior art.
[0006] To achieve the above objectives, the first aspect of the present application provides a Bluetooth speaker control method based on AI voice analysis, comprising:
[0007] Receive the wake-up audio, recognize the wake-up audio and call the audio recognition library;
[0008] receiving control audio, identifying audio attributes of the control audio, and automatically adjusting the feedback volume of the Bluetooth speaker according to the audio attributes, wherein the audio attributes include direction and distance; and
[0009] The control audio is recognized through the audio recognition library, and control instructions are generated based on the recognition results; the operation is completed according to the control instructions, and feedback information is issued based on the feedback volume.
[0010] Preferably, before receiving the wake-up audio, the wake-up word of the Bluetooth speaker is set, including:
[0011] Determine a basic word and several alternative words; and sequentially combine the basic word and the several alternative words to form an alternative phrase; wherein the basic word is a required word set by the Bluetooth speaker manufacturer;
[0012] Perform audio tests on several candidate phrases and select at least one candidate phrase as the wake-up word based on the test results; the audio test includes a wake-up rate test, a stability test, and a discrimination rate test.
[0013] Preferably, audio testing is performed on several candidate phrases, including:
[0014] Test the arousal rate and stability of several alternative phrases;
[0015] The test obtains audio signals of different dialects corresponding to each candidate phrase, and calculates the discrimination rate of the corresponding candidate phrase based on the similarity between the corresponding audio signals of different dialects;
[0016] Based on the awakening rate, stability and discrimination rate, the alternative phrase with the best comprehensive performance is selected as the wake-up word.
[0017] Preferably, identifying audio attributes of the control audio includes:
[0018] Calculate the time difference between the control audio and the wake-up audio;
[0019] When the time difference is greater than the set time threshold, the device is prompted to wake up again; otherwise, the relative position of the control audio and the Bluetooth speaker is identified;
[0020] Gets the audio properties that control the audio based on the relative position.
[0021] Preferably, automatically adjusting the feedback volume of the Bluetooth speaker according to the audio attributes includes:
[0022] Retrieving a property-volume curve; the property-volume curve represents the correspondence between audio properties and optimal volume, and is constructed through simulation testing;
[0023] The audio attributes are input into the attribute-volume curve, and the optimal volume corresponding to the audio attributes is calculated. The optimal volume is used as the feedback volume.
[0024] Preferably, before calculating the time difference between the control audio and the wake-up audio, determining the dialect similarity between the two includes:
[0025] Determine the dialect similarity between the control audio and the wake-up audio;
[0026] When the dialect similarity is greater than the similarity threshold, the control audio is recognized by the audio recognition library; otherwise, the control audio is analyzed to re-match the audio recognition library.
[0027] Preferably, when several control audios are received, the timbre characteristics of each control audio are identified; the timbre characteristics are compared with the timbre characteristics of the preset user, the control audio with the consistent comparison is used as the target audio, and the target audio is analyzed to generate a control instruction; wherein the preset user has priority.
[0028] Preferably, the feedback volume is determined according to several audio properties of the control audio, including:
[0029] Identify several audio attributes that control audio, select the audio attribute closest to the Bluetooth speaker from the several audio attributes as target attribute one, and use the audio attribute of the target audio as target attribute two;
[0030] Volume 1 is determined according to target attribute 1 and the attribute-volume curve, and volume 2 is determined according to target attribute 2 and the attribute-volume curve; feedback volume is determined based on volume 1 and volume 2; volume 1 is the reminder volume.
[0031] Preferably, determining the feedback volume based on the volume one and the volume two includes:
[0032] Set volume 1 as the reminder volume and volume 2 as the target volume;
[0033] After setting the transition between the alert volume and the target volume, generate the feedback volume.
[0034] The second aspect of the present application provides a Bluetooth speaker control system based on AI voice analysis, which is applied to a Bluetooth speaker and includes:
[0035] Data acquisition module: used to receive wake-up audio and control audio;
[0036] Speech analysis module: used to analyze the wake-up audio and call the audio recognition library, determine the feedback volume based on the audio properties of the control audio, and issue feedback information based on the feedback volume.
[0037] Compared with the prior art, the present invention has the following advantages:
[0038] 1. When controlling a Bluetooth speaker, the user first generates a wake-up audio using a wake-up word. The Bluetooth speaker then identifies the corresponding dialect based on the wake-up audio and accesses the audio recognition library corresponding to the dialect. The audio recognition library then identifies the control audio received subsequently, determines the feedback volume based on the audio properties of the control audio, generates a control instruction based on the audio content of the control audio, and finally sends feedback information at the feedback volume. The feedback volume of the present application is automatically adjusted based on the control audio received in real time, which can effectively prevent the volume of the Bluetooth speaker from being too loud or too soft, improve the quality of interaction between the Bluetooth speaker and the user, and enhance the user experience.
[0039] 2. This application obtains several alternative phrases based on the combination of basic words and alternative words, conducts audio tests on several alternative phrases in different dialects, calculates the comprehensive score of each alternative phrase based on the wake-up rate, stability and discrimination rate, and selects one or more with the highest comprehensive score as the wake-up word. Generating multiple alternative phrases based on basic words and alternative words can meet the user's customization needs and the manufacturer's advertising needs; testing the wake-up rate, stability and discrimination rate of alternative phrases in different dialects, selecting the best performing alternative phrase as the wake-up word, calling the audio recognition library based on the wake-up word to identify subsequent control audio, can improve the voice interaction accuracy of Bluetooth speakers in different dialect areas.
[0040] 3. After verifying the interval time and dialect consistency between the control audio and the wake-up audio, this application identifies the audio attributes of the control audio, calls the attribute-volume curve based on the scene in which the Bluetooth speaker is located, obtains the optimal volume corresponding to the audio attributes, and uses this optimal volume as the feedback volume. This application determines the optimal volume by controlling the audio, which is equivalent to determining the feedback volume based on the user's latest location. It can provide the most appropriate feedback volume for mobile users. Moreover, when using the two parameters of direction and distance to obtain the feedback volume, it can improve the scenario applicability and ensure that the feedback volume can meet the user's requirements without the need for frequent adjustments by the user. BRIEF DESCRIPTION OF THE DRAWINGS
[0041] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0042] Figure 1 This is a flowchart of a method for controlling a Bluetooth speaker in Embodiment 1 of the present application;
[0043] Figure 2 This is a schematic diagram of the process of determining the wake-up words for identifying various dialects in Example 2 of the present application;
[0044] Figure 3 This is a flowchart of determining the feedback volume according to the audio attributes of the controlled audio in the third embodiment of the present application;
[0045] Figure 4 This is a flowchart of determining the feedback volume when the Bluetooth speaker receives a plurality of control audios in the fourth embodiment of the present application;
[0046] Figure 5 This is a schematic diagram of the system principle of the Bluetooth speaker control system in this application. DETAILED DESCRIPTION
[0047] The following will clearly and completely describe the technical solutions of this application in conjunction with the embodiments. Obviously, the embodiments described are only a part of the embodiments of this application, not all of them. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.
[0048] Bluetooth speakers are increasingly used in our daily lives, for example, being placed indoors to play music, engage in fun conversations, and even function as control hubs for smart home devices. Whether playing music, engaging conversations, or controlling devices, Bluetooth speakers provide corresponding outputs. However, when used indoors, Bluetooth speakers are relatively fixed in position, and their volume control methods are also relatively fixed: either manually controlled via buttons on the speaker or controlled via a connected smart device (such as a smartphone or tablet). Once the volume is set, it remains fixed and does not automatically adjust unless the volume is readjusted. This results in the Bluetooth speaker outputting a constant volume regardless of the distance from the user. Obviously, when the distance between the two is far, the user will have difficulty receiving feedback from the Bluetooth speaker. When the distance is close, the user may be startled by the feedback, which can affect the interaction and reduce the user experience.
[0049] Example 1
[0050] The Bluetooth speaker control method based on AI voice analysis provided in this embodiment controls the feedback volume by identifying the direction and distance of the audio signal, thereby ensuring that the feedback volume can clearly convey the feedback information to the user while avoiding the feedback volume being too loud and affecting the user experience.
[0051] See also Figure 1 The first embodiment of the present application provides a Bluetooth speaker control method based on AI voice analysis, including:
[0052] Receive the wake-up audio, recognize the wake-up audio and call the audio recognition library;
[0053] receiving control audio, identifying audio attributes of the control audio; automatically adjusting the feedback volume of the Bluetooth speaker according to the audio attributes, wherein the audio attributes include direction and distance; and,
[0054] The control audio is recognized through the audio recognition library, and control instructions are generated based on the recognition results; the operation is completed according to the control instructions, and feedback information is issued based on the feedback volume.
[0055] The Bluetooth speaker control method provided in this embodiment involves two audio signals: a wake-up audio signal and a control audio signal. The wake-up audio signal is used to remind the Bluetooth speaker to accurately receive the subsequent control audio signal, while the control audio signal contains the control content required by the user. Both the wake-up audio signal and the control audio signal originate from the user. It should be noted that the wake-up audio signal generally precedes the control audio signal, and originating from the user can mean directly or indirectly from the user.
[0056] During daily use, Bluetooth speakers may be in various dialect areas, that is, they need to deal with various dialect scenarios. In order to ensure that the dialect can be accurately identified so as to accurately identify the subsequent control audio, this embodiment determines the corresponding dialect type through a specific wake-up audio, calls the corresponding audio recognition library according to the dialect type, and uses the audio recognition library to identify the subsequent control audio. There are multiple audio recognition libraries built into the Bluetooth speaker or stored in the cloud connected to it, and each audio recognition library is used to identify a dialect. After determining the dialect type used by the user, calling the corresponding audio recognition library to identify the control audio can improve the recognition accuracy of the control audio, and calling the audio recognition library in advance can improve the recognition efficiency of the control audio, avoiding the situation where the control audio has been received but the audio recognition library call has not been completed. It should be noted that the audio recognition library is pre-trained with various dialect training data sets, which can be an existing commercial model or an artificial intelligence model trained with self-collected data. It should be noted that the speech recognition of this application is based on existing speech recognition methods and AI large model technology.
[0057] After receiving the control audio, the audio properties of the control audio are first identified. These audio properties specifically include the relative distance and relative direction between the control audio source and the Bluetooth speaker. Bluetooth Direction Finding technology can be used to identify the direction of the user's voice. This technology includes two methods: Angle of Arrival (AoA) and Angle of Departure (AoD), which determine the direction of the signal by measuring the phase difference of the signal. For example, in the AoA method, the receiving device (speaker) uses an antenna array to measure the phase difference of a special signal sent by a single transmitting antenna to estimate the relative direction of the signal. Furthermore, Bluetooth® Channel Sounding technology uses phase-based ranging (PBR) to achieve high-accuracy distance measurement (HADM).
[0058] In some other preferred embodiments, the relative distance may be measured by a built-in distance measuring sensor based on the relative direction, such as by using an infrared sensor to identify the relative distance to the user along the relative direction.
[0059] The audio attributes identified by the above method can determine the relative position (mainly including relative direction and relative distance) between the user and the Bluetooth speaker. The feedback volume of the Bluetooth speaker is then set based on this relative position. Each time control audio is received, the feedback volume is adjusted based on the audio attributes of the control audio (frequent adjustment is not required when the audio attributes have not changed or have changed slightly). This allows for automatic volume setting of the Bluetooth speaker. This solves the problem of fixed volume settings of Bluetooth speakers using existing solutions, where the user may not be able to clearly hear feedback information when the relative distance is far, or the volume may be too loud when the relative distance is close, affecting the user experience.
[0060] While setting the feedback volume according to the audio properties, the audio recognition library is used to identify the content of the audio control, and a control instruction is generated based on the identified content. The relevant operation is performed according to the control instruction. If the operation is music playback, smart question and answer, etc. that requires audio playback, the aforementioned feedback volume is the playback volume; if the operation is to control other devices, feedback information will be issued according to the feedback volume after the control is completed. Of course, if the user has set that no feedback information is required, no feedback information will be issued.
[0061] In this embodiment, when controlling a Bluetooth speaker, the user first generates a wake-up audio signal using a wake-up word. The Bluetooth speaker then identifies the corresponding dialect based on the wake-up audio and accesses the audio recognition library corresponding to that dialect. The audio recognition library then identifies the control audio received, determines the feedback volume based on the audio properties of the control audio, generates a control instruction based on the audio content of the control audio, and finally sends feedback information at the feedback volume. The feedback volume in this embodiment is automatically adjusted based on the control audio received in real time, effectively preventing the Bluetooth speaker from emitting excessive or insufficient volume, improving the quality of interaction between the Bluetooth speaker and the user and enhancing the user experience.
[0062] Example 2
[0063] Based on the first embodiment, this embodiment provides a method for determining a wake-up word to improve the recognizability of the wake-up audio, improve the matching degree of the audio recognition library, and thus improve the recognition accuracy of the control audio. Figure 2 , as follows:
[0064] Before receiving the wake-up audio, set the wake-up word for the Bluetooth speaker, including:
[0065] Determine a basic word and several alternative words; combine the basic word and several alternative words in sequence to form alternative phrases; perform audio tests on the several alternative phrases, and select at least one alternative phrase as the wake-up word based on the test results.
[0066] The wake-up word can be selected from a set of candidate words and a base word entered by the user after the Bluetooth speaker is activated. Alternatively, at least one candidate word that meets the requirements can be pre-determined when the Bluetooth speaker leaves the factory, and the user selects one from the built-in candidate words as the wake-up word. In other words, the wake-up word testing process can be performed after the Bluetooth speaker is activated or before it leaves the factory.
[0067] The aforementioned base words are fixed words, typically code names or nicknames for Bluetooth speakers. Alternative words are phrases set by the manufacturer or user. When combined with the base word, the alternative words should be semantically coherent and ideally have positive meaning. There's no specific order for the alternative words and base words within an alternative phrase; adjusting the order of an alternative word and a base word can result in two alternative phrases.
[0068] After the base words and alternative words are combined into several alternative phrases, audio testing is required for each alternative phrase to ensure that the alternative phrases with the required wake-up rate and stability and that can effectively distinguish dialects are selected as wake-up words. Audio testing of several alternative phrases includes the following steps:
[0069] Test the wake-up rate and stability of several alternative phrases; obtain audio signals corresponding to different dialects for each alternative phrase, and calculate the discrimination rate of the corresponding alternative phrase based on the similarity between the corresponding audio signals of different dialects; select the alternative phrase with the best overall performance as the wake-up word based on the wake-up rate, stability and discrimination rate.
[0070] The wake-up rate is the probability of correctly recognizing a candidate phrase. Specifically, it measures the proportion of utterances that successfully wake up the Bluetooth speaker among all utterances that are actually candidate phrases. Stability primarily refers to the stability of recognition in both the awake and non-awakened states, and can be used to assess the Bluetooth speaker's ability to recognize candidate phrases over extended periods of operation. In addition to the improved wake-up rate and stability in this embodiment, other metrics may include false wake-up rate, response time, and false rejection rate.
[0071] The above discrimination rate is used to represent the ability of several candidate phrases to distinguish various dialects. It can be calculated by the similarity between audio signals. Please refer to the following steps:
[0072] Extracting multiple audio signals in different dialects for each candidate phrase and extracting MFCC (Mel-Frequency Cepstral Coefficients) coefficients of the multiple audio signals; the calculation method of the MFCC coefficients is widely documented in existing solutions, and the calculation process will not be repeated here;
[0073] Calculate the similarity of any two MFCC coefficients. The higher the similarity, the less obvious the discrimination between the two dialects is for the corresponding candidate phrase (which can be distinguished by the set similarity threshold). The proportion of dialects that are clearly distinguished in the same candidate phrase can be statistically calculated as the discrimination rate. The higher the discrimination rate, the better the discrimination ability of the corresponding candidate phrase is for different dialects.
[0074] It should be noted that the similarity calculation methods mentioned above include Euclidean distance and cosine similarity. Euclidean distance: Calculates the Euclidean distance between two sets of MFCC coefficients. The smaller the distance, the higher the similarity. Cosine similarity: Calculates the cosine similarity between two sets of MFCC coefficients. The closer the cosine value is to 1, the higher the similarity.
[0075] In some other preferred embodiments, the audio similarity can also be compared using Mel-spectrum, for example, first extracting the Mel-spectrum of the audio signal, and then calculating the sum of the square differences between the Mel-spectra of two audio signals to evaluate the similarity between the two.
[0076] Finally, you need to select the best candidate phrase based on wake-up rate, stability, and discrimination rate as the wake-up word. You can refer to the following steps:
[0077] 1. All evaluation indicators are processed positively, that is, all indicators are ensured to be as high as possible, and then the data is standardized to eliminate the influence of different indicator dimensions. The standardization formula is:
[0078] ;in, Indicates the The alternative phrases are in The original data on the evaluation indicators, The data are standardized.
[0079] 2. Calculate positive and negative ideal solutions: For each evaluation indicator, calculate the maximum and minimum values of all candidate phrases and use them as positive ideal solutions. and negative ideal solutions .
[0080] 3. Calculate the Euclidean distance between the candidate phrase and the positive ideal solution and the negative ideal solution respectively:
[0081]
[0082]
[0083] in, It is The distance between the candidate phrases and the ideal solution, It is The distance between the candidate phrases and the negative ideal solution.
[0084] 4. Calculate the comprehensive score of each candidate phrase , Comprehensive score The higher it is, the closer the corresponding alternative phrase is to the ideal solution, and the better the effect.
[0085] It is understandable that the weight of each evaluation indicator in the above steps needs to be assigned according to its importance. The weight can be determined by expert scoring method, entropy method, hierarchical analysis method, etc. The weighted standardized evaluation indicator data is used to calculate the comprehensive score.
[0086] After calculating the combined scores of each candidate phrase, the phrase with the highest combined score can be selected as the wake-up word. Taking into account different user preferences, several candidate phrases with the highest combined scores can be selected, and the user can choose one from them as the wake-up word.
[0087] This embodiment obtains several alternative phrases based on the combination of basic words and alternative words, conducts audio tests on these alternative phrases in different dialects, calculates the comprehensive score of each alternative phrase based on the wake-up rate, stability, and discrimination rate, and selects one or more with the highest comprehensive score as the wake-up word. Generating multiple alternative phrases based on basic words and alternative words can meet the user's customization needs and the manufacturer's advertising needs; testing the wake-up rate, stability, and discrimination rate of the alternative phrases in different dialects, selecting the best-performing alternative phrase as the wake-up word, and calling the audio recognition library based on the wake-up word to identify subsequent control audio, can improve the voice interaction accuracy of Bluetooth speakers in different dialect areas.
[0088] Example 3
[0089] Based on the first or second embodiment, this embodiment provides a method for determining the feedback volume. This method not only considers the relative position between the user who sends the wake-up audio and the Bluetooth speaker, but also considers the regional characteristics of the area where the Bluetooth speaker is placed. After comprehensive analysis, it ensures the optimal feedback volume and avoids the impact of ambient noise on the user's reception of feedback information. Figure 3 , please refer to the following steps for details:
[0090] First, identify the audio properties that control the audio, including:
[0091] Calculate the time difference between the control audio and the wake-up audio; when the time difference is greater than the set time threshold, prompt to re-wake the device; otherwise, identify the relative position of the control audio and the Bluetooth speaker; and obtain the audio properties of the control audio based on the relative position.
[0092] Before determining the feedback volume, it is also necessary to prioritize whether feedback is required for the control audio, that is, whether the control audio is connected to the wake-up audio. Therefore, the time difference between the control audio and the wake-up audio is calculated to determine whether the two are connected. If the time difference is greater than the set time threshold, the user is prompted to re-wake the device. If the time difference is less than the set time threshold, the audio attributes of the control audio are identified. It should be noted that when prompting the user to re-wake the device, the volume of the Bluetooth speaker can be determined based on the audio attributes of the wake-up audio to ensure that the prompt volume is sufficient for the user to clearly receive the message.
[0093] In some other preferred embodiments, before calculating the time difference between the control audio and the wake-up audio, the dialect similarity between the two can be determined, including: determining the dialect similarity between the control audio and the wake-up audio; when the dialect similarity is greater than the similarity threshold, the control audio is identified through the audio recognition library; otherwise, the control audio is analyzed to re-match the audio recognition library.
[0094] After receiving the control audio, the dialect similarity between it and the wake-up audio is calculated to ensure that the audio recognition library called based on the wake-up audio is suitable for the control audio. If the control audio and the wake-up audio are not in the same dialect, the audio recognition library needs to be re-determined and called based on the dialect corresponding to the control audio to improve the recognition accuracy of the control audio. In other preferred embodiments, when the dialects of the control audio and the wake-up audio are inconsistent, a reminder can also be sent to re-awaken the Bluetooth speaker.
[0095] After identifying the audio attribute that controls the audio, the attribute-volume curve in the Bluetooth speaker is retrieved and the relevant parameters of the audio attribute are input into the attribute-volume curve to obtain the optimal volume for the audio attribute. This optimal volume is then used as the feedback volume. It should be noted that the attribute-volume curve uses the audio attribute as the independent variable and its optimal volume as the dependent variable. This curve can be simulated using a large amount of data. In general scenarios, the independent variable in the attribute-volume curve is distance, and the feedback volume is determined based on the distance.
[0096] In other preferred embodiments, the independent variables in the attribute-volume curve may include two parameters: direction and distance. This type of curve is suitable for scenes with a lot of background noise or a complex spatial layout. The attribute-volume curve can self-learn the optimal volume based on the volume adjusted by daily users. When there is sufficient data, this type of attribute-variable curve can be established. Of course, it is also possible to simulate and train attribute-variable curves for various special scenarios in the laboratory and select them according to the actual scenario in which the Bluetooth speaker is located.
[0097] It is worth noting that the feedback volume can also be determined based on the audio properties of the wake-up audio, because the direction and distance between the user and the Bluetooth speaker can also be identified based on the wake-up audio. However, when the user actually uses the Bluetooth speaker, the Bluetooth speaker may be woken up while moving, and the position when the control command is issued is different from the position where the wake-up audio is issued. Therefore, this embodiment determines the feedback volume by the audio properties of the control audio. Moreover, if the user moves after issuing the control command, the indoor surveillance camera can be linked or a camera (or sensor) can be set for the Bluetooth speaker to determine the user's position. After determining the feedback information, the feedback volume corresponding to the user's current position is obtained, which can further improve the usability of the feedback volume.
[0098] After verifying the interval time and dialect consistency between the control audio and the wake-up audio, this embodiment identifies the audio attributes of the control audio, invokes the attribute-volume curve based on the scenario in which the Bluetooth speaker is located, and obtains the optimal volume corresponding to the audio attributes. This optimal volume is used as the feedback volume. Determining the optimal volume by controlling the audio in this embodiment is equivalent to determining the feedback volume based on the user's most recent location, providing the most appropriate feedback volume for mobile users. Furthermore, using direction and distance as the two parameters to determine the feedback volume improves scenario applicability and ensures that the feedback volume meets user requirements without requiring frequent adjustments.
[0099] Example 4
[0100] This embodiment, based on the third embodiment, provides a method for determining the feedback volume in a scenario where multiple control audios are being played. In the daily use of Bluetooth speakers, there may be scenarios where multiple people are sending control commands to the Bluetooth speaker at the same time. For example, after one person sends a wake-up word, multiple people send control commands at the same time. In this case, the Bluetooth speaker will receive multiple control audios. When the audio properties of multiple control audios are different, how to determine the feedback volume? This method is proposed to solve this problem. Please refer to Figure 4 , the specific steps are as follows:
[0101] When receiving several control audios, the timbre characteristics of each control audio are identified; the timbre characteristics are compared with the timbre characteristics of the preset users, the control audio with the consistent comparison is used as the target audio, and the target audio is analyzed to generate control instructions; among which the preset users have priority.
[0102] In this embodiment, when the Bluetooth speaker receives multiple control audios, it compares the timbre of each control audio with the wake-up audio, selects the control audio with the same timbre as the target audio, and generates control instructions based on the target audio. However, the feedback volume is not determined solely based on the audio properties of the target audio. It should be noted that in addition to determining the target audio based on timbre comparison, the target audio can also be determined based on the priority of the user's timbre, that is, analyzing the priority corresponding to the timbre of each control audio (this priority is set in advance) and selecting the audio with the highest priority as the target audio.
[0103] While generating control instructions based on the target audio, the feedback volume is determined based on several audio properties of the control audio, including:
[0104] Identify several audio attributes that control audio, select the audio attribute closest to the Bluetooth speaker from the several audio attributes as target attribute one, and use the audio attribute of the target audio as target attribute two;
[0105] Volume one is determined according to target attribute one and the attribute-volume curve, and volume two is determined according to target attribute two and the attribute-volume curve; and feedback volume is determined based on volume one and volume two.
[0106] Target attribute 1 corresponds to the user closest to the Bluetooth speaker, while target attribute 2 corresponds to the target user of the Bluetooth speaker's feedback information. Of course, the target user must be able to hear clearly. If the feedback volume is determined solely based on target attribute 2, the feedback volume may be higher than that of the user corresponding to target attribute 1, which may scare the user.
[0107] Therefore, the feedback volume is determined based on the volume one and the volume two, including:
[0108] Set volume 1 as the reminder volume and volume 2 as the target volume;
[0109] After setting the transition between the alert volume and the target volume, generate the feedback volume.
[0110] Use volume 1 as the reminder volume. You can use volume 1 to play reminder words or reminder music. After the reminder words or reminder music are played, use volume 2 to play feedback information. This ensures that the feedback information reaches the target user while avoiding startling other users. It should be noted that a transition volume should be set between volume 1 and volume 2 to avoid abrupt jumps between volume 1 and volume 2.
[0111] In other preferred embodiments, the control audio may be audio emitted by a pet in addition to the relevant control requirements issued by the user, that is, when the pet is close to the Bluetooth speaker, the volume is increased to remind the pet to prevent the pet from going crazy.
[0112] See also Figure 5In a second aspect, an embodiment of the present application provides a Bluetooth speaker control method based on AI voice analysis, which is applied to a Bluetooth speaker, wherein a data acquisition module and a voice analysis module are in communication and / or electrically connected, including:
[0113] Data acquisition module: used to receive wake-up audio and control audio;
[0114] Speech analysis module: used to analyze the wake-up audio and call the audio recognition library, determine the feedback volume based on the audio properties of the control audio, and issue feedback information based on the feedback volume.
[0115] The sensor is used to collect wake-up audio and control audio, and the audio playback module is used to send feedback information based on the feedback volume. Both the sensor and the audio playback module can be built into the Bluetooth speaker.
[0116] The above embodiments are only used to illustrate the technical method of the present application and are not intended to limit it. Although the present application has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical method of the present application can be modified or replaced by equivalents without departing from the spirit and scope of the technical method of the present application.
Claims
1. A Bluetooth speaker control method based on AI voice analysis, characterized in that: include: Receive a wake-up audio, recognize the wake-up audio, and call an audio recognition library; receiving control audio, identifying audio attributes of the control audio, and automatically adjusting the feedback volume of the Bluetooth speaker according to the audio attributes; wherein the audio attributes include direction and distance; and Recognize the control audio through the audio recognition library and generate a control instruction according to the recognition result; complete the operation according to the control instruction and send feedback information according to the feedback volume; When receiving a plurality of control audios, the timbre characteristics of each control audio are identified; the timbre characteristics are compared with the timbre characteristics of a preset user, the control audio with the consistent comparison is used as the target audio, and the target audio is analyzed to generate a control instruction; wherein the preset users have priority; The feedback volume is determined according to several audio properties of the control audio, including: Identify the audio attributes of the control audio, select the audio attribute closest to the Bluetooth speaker from the audio attributes as target attribute one, and select the audio attribute of the target audio as target attribute two; Determining volume one according to the target attribute one and the attribute-volume curve, and determining volume two according to the target attribute two and the attribute-volume curve; determining a feedback volume based on volume one and volume two; wherein volume one is a reminder volume, used to play reminder words or reminder music, and setting a transition volume between volume one and volume two; Determining a feedback volume based on the first volume and the second volume includes: The first volume is used as the reminder volume, and the second volume is used as the target volume; After setting a transition between the reminder volume and the target volume, a feedback volume is generated.
2. The Bluetooth speaker control method based on AI voice analysis according to claim 1, characterized in that: Before receiving the wake-up audio, setting the wake-up word of the Bluetooth speaker includes: Determine a basic word and several candidate words; and sequentially combine the basic word and the candidate words to form a candidate phrase; wherein the basic word is a required word set by the Bluetooth speaker manufacturer; An audio test is performed on the candidate phrases, and at least one candidate phrase is selected as the wake-up word according to the test results; wherein the audio test includes a wake-up rate test, a stability test, and a discrimination rate test.
3. The Bluetooth speaker control method based on AI voice analysis according to claim 2, characterized in that: Audio testing was performed on several of the candidate phrases, including: Testing the arousal rate and stability of several candidate phrases; The test obtains audio signals of different dialects corresponding to each candidate phrase, and calculates the discrimination rate of the candidate phrase according to the similarity between the audio signals corresponding to the different dialects; Based on the awakening rate, stability and discrimination rate, the alternative phrase with the best comprehensive performance is selected as the wake-up word.
4. The Bluetooth speaker control method based on AI voice analysis according to claim 1, characterized in that: The identifying and controlling audio attributes of the audio includes: Calculating a time difference between the control audio and the wake-up audio; When the time difference is greater than a set time threshold, a prompt is given to re-awaken the device; otherwise, the relative position of the control audio and the Bluetooth speaker is identified; The audio attribute of the control audio is obtained according to the relative position.
5. The Bluetooth speaker control method based on AI voice analysis according to claim 4 is characterized in that: Automatically adjusting the feedback volume of the Bluetooth speaker according to the audio attribute, including: Retrieving a property-volume curve; the property-volume curve represents the correspondence between audio properties and optimal volume, and is constructed through simulation testing; The audio attribute is input into the attribute-volume curve, an optimal volume corresponding to the audio attribute is calculated, and the optimal volume is used as the feedback volume.
6. The Bluetooth speaker control method based on AI voice analysis according to claim 5, characterized in that: Before calculating the time difference between the control audio and the wake-up audio, determining the dialect similarity between the two includes: Determining the dialect similarity between the control audio and the wake-up audio; When the dialect similarity is greater than a similarity threshold, the control audio is recognized by the audio recognition library; otherwise, the control audio is analyzed to re-match the audio recognition library.
7. A Bluetooth speaker control system based on AI voice analysis, applied to a Bluetooth speaker, for executing the Bluetooth speaker control method based on AI voice analysis according to any one of claims 1 to 6, characterized in that: include: Data acquisition module: used to receive wake-up audio and control audio; Speech analysis module: used to analyze the wake-up audio and call the audio recognition library, and determine the feedback volume according to the audio properties of the control audio, and send feedback information according to the feedback volume.
Citation Information
Patent Citations
Voice interaction device and output method thereof
CN106782544A
Speech recognition method, device and system
CN109817220A
Volume control system and method for intelligent loudspeaker box
CN112073874A
Wake word evaluation
US9275637B1