A control system and method based on a Bluetooth headset

Through the voice monitoring and lightweight model of Bluetooth headphones, the problem of insufficient wake-up word dependence and speech signal resolution capabilities in the existing technology is solved, and natural voice interaction and multi-round dialogue are realized, power consumption is reduced, user experience and system robustness are improved.

CN120148512BActive Publication Date: 2025-07-22SHANXI ZUNTE INTELLIGENT TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510615555.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-05-14
Publication Date
2025-07-22
Estimated Expiration
2045-05-14

AI Technical Summary

Technical Problem

The voice interaction control system of existing Bluetooth headphones needs to rely on fixed wake words, resulting in users needing to repeat wake words frequently, the interaction process is lengthy, and it is easy to trigger in multi-person conversation scenarios. It lacks the ability to analyze sublingual information such as tone and rhythm in the voice signal, making it difficult to achieve seamless multi-round dialogue and scenario adaptive interaction.

Method used

The voice monitoring module continuously monitors the ambient sound, uses voice activity detection and lightweight models to make voice dialogue judgments, extracts stress perception and situational perception data, and uses environment complexity to determine whether the main core is awakened, realizing natural voice command interaction.

Benefits of technology

No need to repeat wake-up words, reduce the risk of false triggering and privacy leakage, support multiple rounds of dialogue and contextual association, reduce power consumption, extend battery life, improve user experience and system robustness.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120148512B_ABST
    Figure CN120148512B_ABST
Patent Text Reader

Abstract

The present invention discloses a control system and method based on a Bluetooth headset, relating to the technical field of Bluetooth headsets. The system is composed of several functional modules, including: a voice monitoring module, which continuously listens to the sounds in the environment through the Bluetooth headset, acquires and processes audio signals, and uses voice activity detection technology to judge voice conversations; a semantic understanding module, which acquires the sounds in the environment, uses voice data processing technology to extract voice data; uses an embedded lightweight model to perform semantic understanding and decomposition on the processed voice data, and splits and generates stress perception data and scenario perception data, calculates stress evaluation values and environmental complexity respectively, and outputs the stress perception data and scenario perception data; a wake-up module, which is used to acquire the stress evaluation values and environmental complexity, and respectively compare and judge them with the preset stress evaluation values and preset environmental complexity.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of Bluetooth headsets, and specifically to a control system and method based on Bluetooth headsets. Background Technique

[0002] In recent years, Bluetooth headsets, as an important part of intelligent wearable devices, have been widely used in scenarios such as voice calls, music playback, real-time translation, etc. With the progress of technology and the change of consumer demands, the functions and performances of Bluetooth headsets are also constantly improving. Modern Bluetooth headsets increasingly integrate intelligent functions such as voice assistants, health monitoring, gesture control, etc.

[0003] For example, the current mainstream Bluetooth headset voice interaction control system needs to rely on a fixed wake word to activate the device; this design causes users to need to repeat the wake word frequently, the interaction process is lengthy, and there is a risk of privacy leakage due to accidental triggering in a multi-person conversation scenario; in addition, the wake word mechanism limits the naturalness of voice commands and it is difficult to achieve seamless multi-round conversations; traditional systems focus on speech-to-text (ASR) and simple command recognition, lacking the ability to analyze paralinguistic information such as intonation and rhythm in voice signals; for example, the existing technology cannot distinguish the intonation differences between interrogative sentences and declarative sentences, resulting in a high error rate of intention recognition; at the same time, the device fails to effectively integrate context information such as environmental noise, user location, and device status, and it is difficult to achieve situation-adaptive interaction. Summary of the Invention

[0004] (I) Technical Problems to be Solved

[0005] Aiming at the deficiencies of the existing technology, the present invention provides a control system and method based on Bluetooth headsets, so as to solve the problems in the existing technology that rely on a fixed wake word to activate the device, and at the same time, due to the wake word mechanism restricting the naturalness of voice commands, it is difficult to achieve seamless multi-round conversations, and lacking the ability to analyze paralinguistic information such as intonation and rhythm in voice signals.

[0006] (II) Technical Solutions

[0007] To achieve the above objectives, the present invention is realized through the following technical solutions:

[0008] A control system based on Bluetooth headsets includes:

[0009] A voice monitoring module, which continuously listens to the sounds in the environment through the Bluetooth headset, acquires and processes the audio signal, and uses voice activity monitoring technology to judge voice conversations;

[0010] The semantic understanding module acquires the sounds in the environment, uses speech data processing technology to extract speech data, and uses an embedded lightweight model to perform semantic understanding and decomposition on the processed speech data, and splits and generates stress perception data and scenario perception data, calculates the stress evaluation value and the environmental complexity respectively, and outputs the stress perception data and the scenario perception data.

[0011] The wake-up module compares the stress evaluation value and the environmental complexity with the preset stress evaluation value threshold and environmental complexity threshold respectively. When the condition for direct wake-up is not met, it calculates the wake-up index and determines whether to wake up.

[0012] Further, acquiring and processing the audio signal includes:

[0013] Using a multi-core processor, allocating the speech detection task to a low-power core, keeping the main core in a dormant state during standby, and using the lightweight VAD algorithm integrated in the Bluetooth headset to continuously detect whether there is speech in the environment, and judging whether the current frame contains speech by analyzing the characteristics of the audio signal.

[0014] Segment the continuous audio signal into short-time frames, window each frame signal, use the short-time energy to extract the distinguishing features between speech and non-speech, and perform smoothing processing on the short-time frames of multiple consecutive frames.

[0015] Further, the voice dialogue is judged as:

[0016] By judging the short-time frames of three consecutive frames:

[0017] When the short-time frames of three consecutive frames are all judged as speech, it is determined as a voice dialogue.

[0018] When any one of the short-time frames of three consecutive frames is judged as non-speech, it is determined as noise and continuous monitoring is carried out.

[0019] Further, speech data extraction:

[0020] Call the low-power core to process the sounds in the environment, segment the audio signal into continuous frames, estimate the noise spectrum by superimposing and calculating the spectrum and noise spectrum in the continuous frames, obtain the noise spectrum estimation value, get the noise-reduced speech spectrum, and subtract the noise spectrum estimation value.

[0021] Further, speech data processing:

[0022] Through short-time Fourier transform, the time-domain signal in the denoised speech spectrum is converted into a frequency-domain representation to obtain the complex spectrum on each time frame. Using the estimated noise spectrum value, the noise spectrum of the current frame is adjusted. Under the condition of a stable noise spectrum, the noise spectrum is updated in real time, and signal preprocessing is performed through preloading to enhance the high-frequency part of the speech signal, thereby compensating for the high-frequency loss caused by vocal tract attenuation. By applying a high-pass filter to the input signal, a relationship expression is established using the transfer function.

[0023] Further, semantic understanding and deconstruction:

[0024] Extract MFCC features from the speech data, convert the MFCC features into a format suitable for the input of MobileBERT, understand the semantic content of the speech through a lightweight model, use the MobileBERT model to perform semantic understanding on the converted features, generate semantic labels for user intent and keywords, split the semantic data into stress-aware data and context-aware data, use a lightweight BERT model, input the MFCC feature sequence after passing through the embedding layer and position encoding, output semantic labels, input the logits into a fully connected layer to output the intent category, use a sequence labeling model to extract keywords from the logits, and use a classification head to extract user intent and keywords from the output of MobileBERT.

[0025] Further, stress awareness:

[0026] Perform intonation analysis and rhythm analysis on the processed speech data to obtain stress-aware data; at the same time, based on the processed speech data, calculate the mean and variance of the fundamental frequency, judge the pitch and change of the intonation, use the fundamental frequency curve to analyze the rising, falling, and stable trends of the intonation, and extract the intonation information of the speech for intonation analysis; at the same time, extract the rhythm information of the speech through short-time energy and zero-crossing rate, calculate the change of short-time energy, judge the strong and weak rhythm of the speech, and analyze the change of the zero-crossing rate;

[0027] Intonation analysis:

[0028] There are fundamental frequency values for N time frames, denoted as , and the formula is:

[0029]

[0030] In the formula, is the average value of all time frame fundamental frequency values; N is the total number of fundamental frequency values; i is the current time frame; is the fundamental frequency value at the i-th time frame;

[0031] Calculation of the variance of the fundamental frequency:

[0032] Obtain the average value of all time frame fundamental frequency values , the formula is:

[0033]

[0034] In the formula, is the variance of the fundamental frequency value; is the difference between the fundamental frequency value of each time frame and the average fundamental frequency value;

[0035] By obtaining the average value of the fundamental frequency and the variance of the fundamental frequency, calculate the stress evaluation value, and the formula is:

[0036]

[0037] In the formula, is the stress evaluation value; is the short-time energy value; is the current intonation pattern weight; is the current harmonic mean weight; is the transfer function value.

[0038] Furthermore, the scenario awareness includes:

[0039] Through the semantic content and context information split from the speech data, where the semantic content includes time, place, and person; the context information includes: user intention; through the user intention, obtain the scenario awareness data, so as to obtain the scenario awareness data;

[0040] Calculate the environmental complexity;

[0041] By obtaining the decibel value, the average decibel value of the current environment, and the short-time energy variance of the speech signal, and combining the detected number of speakers, k1, k2, and k3 are the influence of the weight coefficients corresponding to the decibel value, short-time energy variance, and number of speakers respectively, and the formula is:

[0042]

[0043] In the formula, is the environmental complexity index; is the threshold of the short-time energy variance; is the base of the natural logarithm; is the natural logarithm; is the influence weight of the decibel value on the complexity; is the audio fluctuation variance, and as this exponential term increases, the entire fractional part decreases; is the number of speakers identified.

[0044] Further, for obtaining the environmental complexity, it is respectively compared with a preset stress evaluation value threshold and an environmental complexity threshold. When the direct wake-up condition is not met, the wake-up index is calculated to determine whether to perform wake-up, including:

[0045] Determine whether to perform direct wake-up:

[0046] By obtaining the environmental complexity index and the stress evaluation value, and respectively comparing them with 70% of the preset stress evaluation value as the stress evaluation threshold and 45% of the preset environmental complexity as the environmental complexity threshold for comparison and judgment, the judgment is as follows:

[0047] When the stress evaluation value is greater than or equal to 70% of the preset stress evaluation value and the environmental complexity is less than or equal to 45% of the preset environmental complexity directly wake up the main core to perform Bluetooth headset voice reception and voice command recognition;

[0048] When the stress evaluation value is less than 70% of the preset stress evaluation value or the environmental complexity is greater than 45% of the preset environmental complexity calculate the wake-up index and judge the wake-up index to determine whether to wake up the main core;

[0049] Judge the wake-up index:

[0050] By obtaining the stress evaluation value and the environmental complexity, calculate the wake-up index, and the formula is:

[0051]

[0052] In the formula, is the wake-up index; R is the adjustment coefficient; is the threshold of the environmental complexity; e is the base of the natural logarithm; is the natural logarithm; is the influence weight of the stress evaluation value on the wake-up index; is the influence weight of the environmental complexity on the wake-up index; G is the bias term of the set basic wake-up constant value;

[0053] By obtaining the sensitivity threshold set by the user , compare it with the calculated wake-up index for comparison and judgment, and the judgment process is as follows:

[0054] When the wake-up index is greater than the sensitivity threshold , wake up the main core;

[0055] When the wake-up index Less than or equal to the sensitivity threshold Keep the main core in standby sleep mode

[0056] A control method based on a Bluetooth headset, comprising the following steps

[0057] Step 1: Continuously monitor the ambient sound through the Bluetooth headset, acquire and process the audio signal, and use voice activity detection technology to judge voice conversations

[0058] Step 2: Acquire the ambient sound, use speech data processing technology to extract speech data; use an embedded lightweight model to perform semantic understanding and decomposition on the processed speech data, and split and generate stress perception data and scenario perception data, calculate the stress evaluation value and environmental complexity respectively, and output the stress perception data and scenario perception data

[0059] Step 3: Used to obtain the stress evaluation value and environmental complexity, compare them with the preset stress evaluation value threshold and environmental complexity threshold respectively, calculate the wake-up index under the condition that the direct wake-up condition is not met, and judge whether to wake up

[0060] (III) Beneficial effects

[0061] The present invention provides a control system and method based on a Bluetooth headset, with the following beneficial effects

[0062] (1) In this solution, users do not need to repeat the wake-up word, directly interact with the device through natural voice commands, improving the user experience; reducing the risk of false triggering and privacy leakage, especially in multi-person conversation scenarios, judging whether the environment allows voice interaction; and supporting multi-round conversations and context association, making the interaction more fluent and intelligent

[0063] (2) This solution significantly reduces the device power consumption through a low-power voice activity detection (VAD) module and a lightweight model, and a phased processing mechanism (such as waking up the main processor only when valid speech is detected), extending the battery life of the Bluetooth headset. While ensuring real-time performance, it optimizes energy consumption management and is suitable for long-term wearing and use

[0064] (3) This solution real-time optimizes speech recognition and interaction strategies, enhances the system robustness, and uses a lightweight model to achieve efficient real-time processing and low-latency response on resource-constrained Bluetooth headsets, meeting the user's need for instant interaction Description of the drawings

[0065] Figure 1 It is a schematic diagram of the system flow of the present invention

[0066] Figure 2 It is a schematic diagram of the overall method of the present invention Detailed implementation manners

[0067] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0068] Embodiment 1:

[0069] Please refer to Figure 1 , this embodiment provides a control system based on a Bluetooth headset. The warning system includes:

[0070] A voice monitoring module continuously listens to the sounds in the environment through the Bluetooth headset and uses voice activity detection technology to judge voice conversations;

[0071] Bluetooth headset:

[0072] Adopt a multi-core processor, allocate the voice detection task to the low-power core, and the main core remains dormant in the standby state; use a low-power MEMS microphone to reduce power consumption while improving the signal-to-noise ratio;

[0073] For example, one core is responsible for continuous listening, and another core processes complex tasks. During continuous monitoring, the main core stands by and sleeps to reduce power consumption; when running in the low-power mode, only wake up the main processor when voice is detected;

[0074] Voice judgment:

[0075] Use the lightweight VAD algorithm integrated in the Bluetooth headset to continuously detect whether there is voice in the environment, and judge whether the current frame contains voice by analyzing the characteristics of the audio signal;

[0076] Collect the audio signal in the environment through the MEMS microphone in the Bluetooth headset. The sampling rate is usually 8kHz or 16kHz (sufficient to cover the voice frequency range); divide the continuous audio signal into short-time frames (usually 20 - 30ms per frame), window each frame signal (such as Hamming window) to reduce the boundary effect, use a simple noise suppression algorithm (such as spectral subtraction) to reduce background noise; use short-time energy to extract features that can distinguish voice and non-voice, and smooth the short-time frames of continuous multiple frames to avoid misjudgment; judge by judging the short-time frames of three consecutive frames, and the judgment is:

[0077] When the short-time frames of three consecutive frames are all judged as voice, it is judged as a voice conversation;

[0078] When any one of the short-time frames of three consecutive frames is judged as non-voice, it is judged as noise, and the Bluetooth headset continuously listens to the sounds in the environment;

[0079] The Bluetooth headset supports keyword wake-up. It uses a lightweight keyword detection model that is deployed to run locally on the phone. Running locally means that by connecting the Bluetooth headset to a mobile phone or tablet, an accompanying installation package is installed to calculate the sound input from the Bluetooth headset; the model can be optimized for specific wake-up words to reduce the amount of calculation and response time;

[0080] For example, if the Bluetooth headset is set with the name "XX", it can be woken up by various methods such as "XX", "Call XX", "Hello XX", "Are you there XX", etc., to improve the success rate of accurately waking up the Bluetooth headset;

[0081] The semantic understanding module obtains the sound in the environment, extracts the speech data, and uses an embedded lightweight model to perform semantic understanding and decomposition on the speech data, and splits and generates stress perception data and scenario perception data, and outputs the stress perception data and scenario perception data;

[0082] The process of speech data extraction is as follows:

[0083] Call the low-power core to process the sound in the environment, split the audio signal into continuous frames, and calculate the stable spectrum in the continuous frames by superposition. The stable spectrum in the continuous frames is the noise spectrum; estimate the noise spectrum to obtain the noise spectrum estimate value, and subtract the noise spectrum estimate value to obtain the noise-reduced speech spectrum. The formula is:

[0084]

[0085] Among them, is the noise-reduced speech spectrum; is the amplitude spectrum of the noisy speech, is the estimate of the noise amplitude spectrum, is a subtraction factor (usually greater than 1) used to control the amount of noise subtracted; is the maximum spectrum value;

[0086] The process of speech data processing is as follows:

[0087] And perform short-time Fourier transform to convert the time-domain signal in the noise-reduced speech spectrum into a frequency-domain representation to obtain the complex spectrum on each time frame. Use the noise spectrum estimate value to adjust the noise spectrum of the current frame. Under the condition of stable noise spectrum, update the noise spectrum in real time to adapt to the changing noise environment, and perform signal preprocessing through preloading to enhance the high-frequency part of the speech signal, thereby compensating for the high-frequency loss caused by channel attenuation. By applying a high-pass filter to the input signal and establishing a relationship expression using the transfer function, the formula is:

[0088]

[0089] In the formula, is the transfer function value, which describes the representation of the relationship between input and output in the Z-domain; is a positive number less than 1, and the typical value is between 0.9 and 1; is the delay operator. In a discrete-time system, it represents delaying the signal at the current moment by one sampling period; and it is implemented through a difference equation, and the formula is:

[0090]

[0091] In the formula, is the input signal; is the output signal; is the previous input signal;

[0092] Semantic understanding and decomposition:

[0093] Extract MFCC features from the speech data, convert the MFCC features into a format suitable for the input of MobileBERT, understand the semantic content of the speech through a lightweight model, use the MobileBERT model to perform semantic understanding on the converted features, generate semantic labels of user intentions and keywords, split the semantic data into stress-aware data and context-aware data, and output the split speech data for subsequent processing;

[0094] For example:

[0095] Generate context information by combining device sensor data (such as time, location, user behavior); extract time from the speech (such as "3 o'clock tomorrow afternoon"); extract location from the speech (such as "a certain Beijing"), or combine GPS data; device status: combine the current status of the device (such as playing music);

[0096] MFCC feature extraction:

[0097] Use the Python_speech_features library to extract MFCC features, convert the time-domain signal into a frequency-domain signal, calculate the logarithmic energy of the Mel spectrum, and extract MFCC coefficients;

[0098] Code example:

[0099] import librosa

[0100] def extract_mfcc(audio_data, sample_rate=16000, n_mfcc=13):

[0101] mfcc_features = librosa.feature.mfcc(y=audio_data, sr=sample_rate, n_mfcc=n_mfcc)

[0102] return mfcc_features.T # Transpose to (number of time frames, MFCC dimension)

[0103] Feature transformation:

[0104] MobileBERT is a text-based NLP model. Therefore, the MFCC features need to be transformed into a format suitable for model input. Serialize the MFCC features, treat the MFCC feature sequence as "pseudo-text", use an additional embedding layer to map the MFCC features to a vector space compatible with MobileBERT, and add positional encoding to the MFCC feature sequence to preserve the temporal information;

[0105] Code example:

[0106] import torch

[0107] import torch.nn as nn

[0108] class MFCCEmbedding(nn.Module):

[0109] def __init__(self, mfcc_dim, hidden_dim):

[0110] super(MFCCEmbedding, self).__init__()

[0111] self.embedding = nn.Linear(mfcc_dim, hidden_dim)

[0112] self.position_encoding = nn.Parameter(torch.zeros(1, 1000,hidden_dim)) # Assume the maximum sequence length is 1000

[0113] def forward(self, mfcc_features):

[0114] seq_length = mfcc_features.size(1)

[0115] embeddings = self.embedding(mfcc_features) # (batch_size,seq_length, hidden_dim)

[0116] embeddings += self.position_encoding[:, :seq_length, :] # Add positional encoding

[0117] return embeddings

[0118] Perform edge computing in a mobile device connected to a Bluetooth headset. Utilize a lightweight BERT model. By inputting the MFCC feature sequence that has passed through the embedding layer and positional encoding, output semantic labels (user intent and keywords). Input the logits into a fully connected layer to output the intent category. Use a sequence labeling model (such as CRF or BiLSTM) to extract keywords from the logits, and use a classification head to extract user intent and keywords from the output of MobileBERT;

[0119] It should be noted that: MFCC features are continuous numerical data, while MobileBERT is designed for discrete text data. Use the embedding layer to map MFCC features to a vector space compatible with MobileBERT. Although MobileBERT is lightweight, it may still be limited in running on embedded devices and requires edge computing on a smartphone device connected to a Bluetooth headset;

[0120] Stress perception:

[0121] Perform intonation analysis and rhythm analysis on the processed speech data to obtain stress perception data; By calculating the mean and variance of the fundamental frequency, judge the pitch and changes of the intonation, use the fundamental frequency curve to analyze the rising, falling and stable trends of the intonation, and extract the intonation information of the speech for intonation analysis; At the same time, extract the rhythm information of the speech through short-time energy and zero-crossing rate, calculate the change of short-time energy, judge the strong and weak rhythm of the speech, and analyze the change of zero-crossing rate;

[0122] Intonation analysis:

[0123] There are fundamental frequency values for N time frames, denoted as , and the formula is:

[0124]

[0125] In the formula, is the average value of the fundamental frequency values for all time frames (i.e., the mean of the fundamental frequency), which reflects the average pitch of the entire speech segment; N is the number of total fundamental frequency values, that is, the number of time frames; i is the current time frame; is the fundamental frequency value at the i-th time frame, where represents the time point or index of the i-th time frame;

[0126] Calculation of the variance of the fundamental frequency:

[0127] Obtain the average value of the fundamental frequency values for all time frames , and the formula is:

[0128]

[0129] In the formula, is the variance of the fundamental frequency values, which measures the degree of deviation of the fundamental frequency values from their mean and reflects the variation range of the speech intonation; is the difference between the fundamental frequency value of each time frame and the fundamental frequency mean, and this difference reflects the deviation of the fundamental frequency value of this time frame from the overall average fundamental frequency value; is the square of the above difference. Squaring the difference is to ensure that all results are positive numbers and to amplify larger deviations, so as to better reflect the dispersion degree of the data;

[0130] By identifying the turning points in the fundamental frequency curve, using a pre-trained classification model, identifying different intonation patterns (such as rising tone, falling tone, flat tone, etc.) according to the fundamental frequency characteristics, and automatically annotating the intonation type according to the fundamental frequency change pattern;

[0131] Combining the fundamental frequency information and other speech features (such as energy, zero-crossing rate, etc.), identifying the stress positions in the speech. Stress usually appears as local fundamental frequency peaks or significant energy enhancements;

[0132] For example:

[0133] There is a speech signal containing "Hello, nice to meet you". The following are the specific processing steps:

[0134] By dividing the speech signal into several 25-millisecond frames and applying a Hanning window to each frame; calculating the fundamental frequency value of each frame using the autocorrelation method; performing low-pass filtering on the fundamental frequency sequence and removing significantly deviated fundamental frequency values (such as silent segments); calculating the fundamental frequency mean and variance ; observing the fundamental frequency curve, it is found that the fundamental frequency of some words (such as "nice") is relatively high and has an obvious upward trend; using the classification model to identify that the intonation pattern of this sentence is a rising tone (for example, the part "nice to"); combining the fundamental frequency and energy information, identifying the "nice" part as the stress area and analyzing its rhythm characteristics;

[0135] By obtaining the mean value of the fundamental frequency and the variance of the fundamental frequency, calculate the stress evaluation value, and the formula is:

[0136]

[0137] In the formula, is the stress evaluation value; is the short-time energy value, which is the difference between the peak of the current speech data and the underestimated energy value; is the weight of the current intonation pattern, that is, the rising, falling and stable trends, 1 for rising, 2 for falling, and 3 for stable; is the weight of the current harmonic mean;

[0138] Situational awareness:

[0139] Through the semantic content and context information split from the speech data, where the semantic content includes time, place and person; the context information includes: user intention; through the user intention, obtain the situational awareness data, so as to obtain the situational awareness data;

[0140] The lightweight speech recognition model converts speech into text and extracts time, place and person information from the text. Example code:

[0141] import spacy

[0142] # Load the model

[0143] nlp = spacy.load("en_core_web_sm")

[0144] doc = nlp("Meet Zhang at a certain place in Beijing at 3 pm tomorrow.")

[0145] # Extract entities

[0146] entities = {

[0147] "Time": [ent.text for ent in doc.ents if ent.label_ == "TIME"],

[0148] "Place": [ent.text for ent in doc.ents if ent.label_ == "GPE"],

[0149] "Person": [ent.text for ent in doc.ents if ent.label_ == "PERSON"]

[0150] }

[0151] print(entities) # Output: {'Time': ['Tomorrow at 3 PM'], 'Location': ['A certain city'], 'Person': ['Mr. Zhang']}

[0152] Use a lightweight intent classification model that takes the semantic embedding of the speech as input and the intent label as output, such as playing music, querying the weather, etc., and combines the semantic content and context information into situation-aware data;

[0153] Calculate the environmental complexity;

[0154] When the decibel value is higher, the complexity is higher. At the same time, it is judged whether there are multiple speakers through sound source localization or voice separation. The greater the short-time energy variance of the voice signal, the higher the complexity; by obtaining the average decibel value of the current environment and the short-time energy variance of the voice signal, combined with the detected number of speakers (1 represents a single speaker, greater than 1 represents multiple speakers), k1, k2, and k3 are the influence weights corresponding to the decibel value, short-time energy variance, and number of speakers respectively. The formula is:

[0155]

[0156] In the formula, is the environmental complexity index; is the threshold of the short-time energy variance, used to adjust the influence of the energy variance on the complexity; is the base of the natural logarithm (approximately equal to 2.71828); is the natural logarithm; is the influence weight of the decibel value on the complexity. The higher the decibel value, the higher the complexity; is a smoothing function used to adjust the interaction between the decibel value and the short-time energy variance. When the short-time energy variance is close to or exceeds the threshold this term approaches 1, making the influence of the decibel value more significant; conversely, if the short-time energy variance is low, this term will increase, weakening the influence of the decibel value; through this part, the influence of the short-time energy variance on the complexity can be adjusted. When is large, the exponential term approaches 0, so that the whole fractional part approaches ; when is small, the exponential term increases, so that the whole fractional part decreases; The logarithmic function is used to handle the number of speakers to ensure that even in the case of multiple speakers, the growth of complexity is gradual; is a weight coefficient used to adjust the degree of influence of the number of speakers on the complexity;

[0157] A wake-up module, which is used to obtain the stress evaluation value and the environmental complexity, and compare them with the preset stress evaluation value 70% of which is used as the stress evaluation threshold and the preset environmental complexity 45% of which is used as the environmental complexity threshold for comparison and judgment to determine whether to directly wake up. If the direct wake-up condition is not met, calculate the wake-up index and determine whether to wake up;

[0158] Determine whether to directly wake up:

[0159] By obtaining the environmental complexity index and the stress evaluation value, and comparing them with the preset stress evaluation value 70% of which is used as the stress evaluation threshold and the preset environmental complexity 45% of which is used as the environmental complexity threshold for comparison and judgment, and the judgment is:

[0160] When the stress evaluation value is greater than or equal to 70% of the preset stress evaluation value and the environmental complexity is less than or equal to 45% of the preset environmental complexity directly wake up the main core, perform Bluetooth headset voice reception, and voice command recognition;

[0161] When the stress evaluation value is less than 70% of the preset stress evaluation value or the environmental complexity is greater than 45% of the preset environmental complexity calculate the wake-up index and judge the wake-up index to determine whether to wake up the main core;

[0162] Judge the wake-up index:

[0163] By obtaining the stress evaluation value and the environmental complexity, calculate the wake-up index, and the formula is:

[0164]

[0165] In the formula, is the wake-up index; R is a regulation coefficient used to control the influence degree of the environmental complexity on the wake-up index; is the threshold of the environmental complexity, which is used to adjust the influence of the complexity on the wake-up index; e is the base of the natural logarithm (approximately equal to 2.71828); is the natural logarithm; reflects the influence weight of the stress evaluation value on the wake-up index, is a smoothing function used to adjust the interaction between the stress evaluation value and the environmental complexity; when the environmental complexity is close to or exceeds the threshold this term approaches 1, making the influence of the stress evaluation value more significant; conversely, if the environmental complexity is low, this term will increase, weakening the influence of the stress evaluation value; It directly reflects the influence weight of environmental complexity on the wake-up index; the more complex the environment, the higher the wake-up index; G is the offset term of the set basic wake-up constant value, ensuring that the system still has a certain wake-up ability even in the absence of significant accents and complex environments;

[0166] For example:

[0167] Suppose we have the following measurement data:

[0168] Accent evaluation value , environmental complexity , weight coefficient , , adjustment coefficient , threshold , offset term , and substitute them into the formula for calculation:

[0169] 1. Calculate the accent evaluation value part:

[0170]

[0171] 2. Calculate the environmental complexity part:

[0172]

[0173] 3. Final wake-up index:

[0174]

[0175] Through the above complex high-math formula, we can comprehensively consider the accent evaluation value and environmental complexity to calculate the wake-up index; such a formula can not only reflect the independent influence of each factor, but also capture the interaction between them, thus providing a more comprehensive and accurate assessment of the wake-up state;

[0176] Judge based on the wake-up index:

[0177] By obtaining the sensitivity threshold set by the user , compare it with the calculated wake-up index for comparison and judgment. The judgment process:

[0178] When the wake-up index is greater than the sensitivity threshold , wake up the main core and recognize the voice data at this time;

[0179] When the wake-up index is less than or equal to the sensitivity threshold , keep the main core in standby sleep and continuously monitor the sounds in the environment;

[0180] This has important reference value for application scenarios such as intelligent voice assistants and smart home devices; determine whether to activate the voice assistant according to the wake-up index to avoid false wake-up; adjust the response sensitivity of the device according to the wake-up index to improve the user experience; determine whether to start certain automated operations according to the wake-up index to ensure that the system responds only when necessary.

[0181] Embodiment 2:

[0182] Please refer to Figure 2 , based on Embodiment 1, this embodiment also provides a control method based on a Bluetooth headset, including the following specific steps:

[0183] Step 1: Continuously monitor the sounds in the environment through the Bluetooth headset, acquire and process the audio signals, and use voice activity detection technology to judge voice conversations;

[0184] Step 2: Acquire the sounds in the environment, use voice data processing technology to extract voice data; use an embedded lightweight model to perform semantic understanding and decomposition on the processed voice data, and split and generate stress perception data and scenario perception data, calculate the stress evaluation value and the environmental complexity respectively, and output the stress perception data and the scenario perception data;

[0185] Step 3: Compare the stress evaluation value and the environmental complexity with the preset stress evaluation value threshold and environmental complexity threshold respectively. Under the condition that the direct wake-up condition is not met, calculate the wake-up index and judge whether to wake up.

[0186] The above embodiments can be implemented in whole or in part by software, hardware, firmware or any other combination. When implemented using software, the above embodiments can be implemented in whole or in part in the form of a computer program product. Those of ordinary skill in the art can realize that the units and algorithm steps of the examples described in conjunction with the embodiments disclosed herein can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in hardware or software depends on the specific application and design constraints of the technical solution.

[0187] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units. They can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0188] The above is only the specific implementation manner of this application, but the protection scope of this application is not limited thereto. Any person skilled in the art can easily think of changes or substitutions within the technical scope disclosed in this application, and all should be covered by the protection scope of this application.

Claims

1. A control system based on a Bluetooth headset, characterized in that: The system includes: The voice monitoring module continuously monitors the sounds in the environment through Bluetooth headphones, obtains and processes audio signals, and uses voice activity monitoring technology to judge voice conversations; The semantic understanding module obtains the sounds in the environment and uses voice data processing technology to extract voice data. It uses an embedded lightweight model to perform semantic understanding and deconstruction on the processed voice data, and splits it into stress perception data and situation perception data, and calculates the stress evaluation value and environmental complexity respectively. The wake-up module compares the stress evaluation value and the environment complexity with the preset stress evaluation value threshold and the environment complexity threshold respectively. If the direct wake-up conditions are not met, the wake-up index is calculated to determine whether to wake up.

2. The control system based on a Bluetooth headset according to claim 1, characterized in that, Acquire and process audio signals, including: A multi-core processor is used to assign the voice detection task to the low-power core. The main core remains dormant in standby mode. The lightweight VAD algorithm integrated in the Bluetooth headset is used to continuously detect whether there is voice in the environment. By analyzing the characteristics of the audio signal, it is determined whether the current frame contains voice. The continuous audio signal is divided into short-time frames; each frame signal is windowed, and the distinguishing features of speech and non-speech are extracted using short-time energy, and multiple consecutive short-time frames are smoothed.

3. The control system based on a Bluetooth headset according to claim 2, wherein The voice dialogue is judged as: By judging the short time frames of three consecutive frames: When three consecutive short-time frames are all judged as speech, it is determined to be a speech conversation; When any frame in three consecutive short-time frames is judged as non-speech, it is determined as noise and continues to be monitored.

4. The control system based on a Bluetooth headset according to claim 2, characterized in that: Voice data extraction: The low-power core is called to process the sound in the environment, the audio signal is divided into continuous frames, and the noise spectrum is estimated by superimposing the spectrum and noise spectrum in the continuous frames, the noise spectrum estimation value is obtained, the speech spectrum after noise reduction is obtained, and the noise spectrum estimation value is subtracted.

5. The control system based on a Bluetooth headset according to claim 4, characterized in that: Voice data processing: Through short-time Fourier transform, the time domain signal in the denoised speech spectrum is converted into frequency domain representation to obtain the complex spectrum on each time frame. The noise spectrum of the current frame is adjusted using the noise spectrum estimation value. The noise spectrum is updated in real time under the condition that the noise spectrum is stable, and the signal is preprocessed through preloading. By applying a high-pass filter to the input signal, a relational expression is established using the transfer function.

6. The control system based on a Bluetooth headset according to claim 5, characterized in that: Semantic understanding and deconstruction: Extract MFCC features from speech data, convert MFCC features into a format suitable for MobileBERT input, understand the semantic content of speech through a lightweight model, use the MobileBERT model to perform semantic understanding on the converted features, generate semantic labels for user intent and keywords, split the semantic data into accent-aware data and context-aware data, use the lightweight BERT model, output semantic labels by inputting the MFCC feature sequence after the embedding layer and position encoding, input the logits into a fully connected layer, output the intent category, use the sequence labeling model to extract keywords from the logits, and use the classification head to extract user intent and keywords from the output of MobileBERT.

7. The control system based on a Bluetooth headset according to claim 6, characterized in that: Accent Perception: Perform intonation analysis and rhythm analysis on the processed speech data to obtain stress perception data; at the same time, based on the processed speech data, calculate the mean and variance of the fundamental frequency, judge the pitch and variation of the intonation, use the fundamental frequency curve to analyze the rising, falling and stable trends of the intonation, extract the intonation information of the speech, and perform intonation analysis; at the same time, extract the rhythm information of the speech data through short-time energy and zero-crossing rate, calculate the short-time energy change, judge the strong and weak rhythm of the speech, and analyze the zero-crossing rate change. Intonation analysis: The fundamental frequency values of N time frames are denoted as , and the formula is: Wherein, is the average value of the fundamental frequency values of all time frames; N is the number of the total fundamental frequency values; i is the current time frame; is the fundamental frequency value at the i-th time frame; Calculation of fundamental frequency variance: Obtain the average value of the fundamental frequency values for all time frames , and the formula is: In the formula, is the variance of the fundamental frequency value; is the difference between the fundamental frequency value of each time frame and the mean fundamental frequency; By obtaining the mean and variance of the fundamental frequency, calculate the stress evaluation value. The formula is: Wherein, is the stress evaluation value; is the short-time energy value; is the current intonation pattern weight; is the current harmonic mean weight; is the transfer function value.

8. The control system based on a Bluetooth headset according to claim 7, wherein: Situational awareness: Through the semantic content and context information split from the speech data, where the semantic content includes: time, place and person; the context information includes: user intention; through the user intention, obtain the situational awareness data. Calculate the environmental complexity; By obtaining the decibel value, the average decibel value of the current environment, and the short-time energy variance of the speech signal, and combining the detected number of speakers, where k1, k2, and k3 are the influence of the weight coefficients corresponding to the decibel value, short-time energy variance, and number of speakers respectively. The formula is: In the formula, is the environmental complexity index; is the threshold of short-time energy variance; is the influence weight of the decibel value on complexity; is the audio fluctuation variance; is the number of speakers identified.

9. The control system based on a Bluetooth headset according to claim 8, characterized in that: Obtain the environmental complexity, compare it with the preset stress evaluation value threshold and environmental complexity threshold respectively. Under the condition of not meeting the direct wake-up condition, calculate the wake-up index and judge whether to wake up, including: Judge whether to perform direct wake-up: By obtaining an environmental complexity index and a stress evaluation value, and respectively comparing them with 70% of a preset stress evaluation value as a stress evaluation threshold and 45% of a preset environmental complexity as an environmental complexity threshold for comparison and judgment, the judgment is: When the stress evaluation value is greater than or equal to 70% of the preset stress evaluation value and the environmental complexity is less than or equal to 45% of the preset environmental complexity directly wake up the main core to perform Bluetooth headset voice reception and voice command recognition; When the stress evaluation value is less than 70% of the preset stress evaluation value or when the environmental complexity is greater than 45% of the preset environmental complexity calculate the wake-up index and judge the wake-up index to determine whether to wake up the main core; Judge the wake-up index: By obtaining the stress evaluation value and environmental complexity, calculate the wake-up index. The formula is: In the formula, is the arousal index; R is the adjustment coefficient; is the threshold of environmental complexity; is the influence weight of the stress evaluation value on the arousal index; is the influence weight of environmental complexity on the arousal index; G is the bias term of the set basic arousal constant value; By obtaining the sensitivity threshold set by the user , compare with the calculated wake-up index for comparison and judgment. The judgment process is as follows: When the wake-up index is greater than the sensitivity threshold the main core is woken up; When the wake-up index is less than or equal to the sensitivity threshold keep the main core in standby sleep.

10. A control method based on a Bluetooth headset, characterized in that: Include the following steps: Step 1: Continuously monitor the sound in the environment through the Bluetooth headset, obtain and process the audio signal, and use the voice activity detection technology to judge the voice conversation. Step 2: Obtain the sound in the environment, use the speech data processing technology to extract the speech data; use the embedded lightweight model to perform semantic understanding and decomposition on the processed speech data, and split and generate stress perception data and situational awareness data, calculate the stress evaluation value and environmental complexity respectively, and output the stress perception data and situational awareness data. Step 3: Compare the stress evaluation value and environmental complexity with the preset stress evaluation value threshold and environmental complexity threshold respectively. Under the condition of not meeting the direct wake-up condition, calculate the wake-up index and judge whether to wake up.

Citation Information

Patent Citations

  • Automated classification of emotio-cogniton

    CA3207044A1

  • Auto speech conversion method and apparatus

    CN101359473A