Local AI wake-up interaction method and system for dual-mode earphone
Through the local AI wake-up interaction method of dual-mode headphones, the combination of Bluetooth and WiFi modules is used to solve the network latency and privacy security issues of AI wake-up in the existing technology, improve data transmission efficiency and interaction stability, and provide an efficient and secure local AI interaction experience.
Patent Information
- Application Number
- CN202510310521.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-17
- Publication Date
- 2025-05-16
AI Technical Summary
In the prior art, the AI wake-up method has problems with network delay, instability and privacy security, and the local AI wake-up technology has defects in connection mode and data transmission, resulting in slow data transmission speed, unable to meet the needs of complex command interactions, and it is prone to wake-up interrupts and data transmission abnormalities.
The local AI wake-up interaction method of dual-mode headphones is adopted to establish an initial connection with the smart device through Bluetooth low-power technology, and the wake-up signal is obtained using the built-in voice recognition module, and transmitted to the smart device through the Bluetooth module for verification. After receiving the confirmation information, the headset is encrypted and compressed, automatically starts the WiFi module to connect to the local AI server, transmits audio data for processing, and passes the results back to the headset for presentation.
It solves the wake-up problem caused by poor network, reduces the risk of privacy leakage, improves data transmission efficiency, ensures stability and fluency during connection mode switching or network fluctuations, and provides a stable and smooth local AI interactive experience.
Smart Images

Figure CN120018097A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of headset interaction technology, and in particular to a dual-mode headset local AI wake-up interaction method and system. Background Art
[0002] In today's era of rapid digital development, smart audio devices are deeply integrated into people's daily lives. Among them, headphones have become an indispensable personal equipment for the public due to their convenience and functionality. From playing soothing music to wake up in the morning, to listening to audiobooks on the way to work to isolate the outside world, to focusing on answering voice conferences while working or studying, headphones play a key role in many scenarios. At the same time, artificial intelligence technology is booming at a rapid pace, bringing disruptive changes to various fields. This wave of change has also profoundly affected the field of headphones. Users are no longer satisfied with the simple audio playback function of headphones, and their desire for intelligent interactive functions is increasing day by day. However, there are still many shortcomings in AI wake-up and function realization.
[0003] On the one hand, the traditional AI wake-up method that relies on cloud servers faces problems such as network delay, instability, and privacy security. In an environment with poor network signals, the wake-up response is slow or even fails to wake up normally, which greatly affects the user experience. On the other hand, although local AI wake-up can avoid network problems to a certain extent, existing local AI wake-up technologies often have defects in connection mode and data transmission. The common headphone connection mode is Bluetooth, which has limited transmission bandwidth and is unable to cope with large amounts of AI-related data. In addition, the single Bluetooth connection mode is difficult to meet the demand for efficient data transmission in local AI interactions. For example, when performing complex natural language processing or image recognition command interactions, the data transmission speed seriously restricts the efficiency of AI function implementation. In addition, the existing local AI wake-up interaction method is prone to wake-up interruption and data transmission abnormality when switching between different connection modes, and cannot provide users with a stable and smooth local AI interaction experience; To this end, a local AI wake-up interaction method and system for dual-mode headphones are proposed. Summary of the invention
[0004] In view of this, the embodiments of the present invention hope to provide a dual-mode headset local AI wake-up interaction method and system to solve or alleviate the technical problems existing in the prior art and at least provide a beneficial option.
[0005] In order to solve the above technical problems, a technical solution adopted in this application is: a dual-mode headset local AI wake-up interaction method, comprising the following steps: Step 1: Turn on the headset and initialize it, and establish a Bluetooth connection with the smart device based on Bluetooth low energy technology; Step 2: Based on the built-in voice recognition module of the headset, the wake-up signal is obtained by continuously monitoring specific wake-up words or operation instructions; Step 3: pre-process the acquired wake-up signal, and send the pre-processed wake-up signal to the smart device through the Bluetooth module; Step 4: After receiving the wake-up signal from the headset, the smart device verifies the wake-up signal and feeds back confirmation information to the headset; Step 5: After receiving the confirmation information, the headset encrypts and compresses the wake-up signal to generate audio data, and automatically starts the WiFi module to establish a connection with the local AI server to transmit the audio data to the local AI server; Step 6: After receiving the audio data, the local AI server analyzes and processes the received audio data based on computing power and local data resources, and generates processing results; Step 7: The local AI server transmits the generated processing results back to the headset, and the headset presents the processing results to the user in the form of voice broadcast or visualization.
[0006] As a further preferred embodiment of the present technical solution, in step six, the method for generating the processing result comprises the following steps: Step 601: decompress and decrypt the encrypted and compressed audio data to restore it to the original audio signal; Step 602: Convert the audio signal into text information based on speech recognition technology, and analyze the text information through natural language processing technology to identify the type of user instruction; Step 603: According to the user instruction type, the corresponding processing module and local data resources are called to perform in-depth processing and generate processing results; Step 604: Format and optimize the processing result according to the display and voice broadcast capabilities of the headset.
[0007] As a further preferred embodiment of the present technical solution, in step 2, the speech recognition module built into the headset runs based on a local lightweight model or a cloud lightweight model, and the specific wake-up word is set according to the user's own needs.
[0008] As a further preferred embodiment of the present technical solution, in step three, the method for preprocessing the acquired wake-up signal comprises the following steps: Step 301: Convert the wake-up signal in audio form into a digital signal; Step 302: Perform noise reduction processing on the digital signal based on a noise reduction algorithm to remove environmental noise interference; Step 303: perform feature extraction on the processed digital signal to extract key features related to the wake-up word.
[0009] As a further preferred embodiment of the present technical solution, in step four, the smart device pre-stores a verification key and verification rules agreed upon with the headset; the wake-up signal is verified by comparing the wake-up signal with the pre-stored verification key and verification rules; if the comparison result matches, the wake-up signal is determined to have been verified, and confirmation information is fed back to the headset; if the comparison result does not match, the verification is determined to have failed, no confirmation information is fed back, and an error prompt information may be sent to the headset.
[0010] As a further preferred embodiment of the present technical solution, in step five, when encrypting and compressing the wake-up signal, the AES algorithm and AAC encoding are used to encrypt and compress the wake-up signal.
[0011] As a further preferred embodiment of the present technical solution, the user instruction types include natural language processing instructions, image recognition instructions and multimedia control instructions.
[0012] In order to solve the above technical problems, another technical solution adopted by the present application is: a dual-mode headset local AI wake-up interaction system, the system comprising: a Bluetooth connection module, a signal acquisition module, a signal preprocessing module, a signal verification module, a data generation and transmission module, a data analysis and processing module and a result presentation module; The Bluetooth connection module is configured to turn on and initialize the headset and establish a Bluetooth connection with the smart device based on Bluetooth low energy technology; The signal acquisition module is configured to acquire a wake-up signal by continuously monitoring a specific wake-up word or operation instruction based on a voice recognition module built into the headset; The signal preprocessing module is configured to preprocess the acquired wake-up signal and send the preprocessed wake-up signal to the smart device through the Bluetooth module; The signal verification module is configured to verify the wake-up signal after the smart device receives the wake-up signal from the headset, and feedback confirmation information to the headset; The data generation and transmission module is configured to encrypt and compress the wake-up signal after the headset receives the confirmation information, generate audio data, and automatically start the WiFi module to establish a connection with the local AI server to transmit the audio data to the local AI server; The data analysis and processing module is configured to analyze and process the received audio data based on computing power and local data resources after the local AI server receives the audio data, and generate a processing result; The result presentation module is configured so that the local AI server transmits the generated processing results back to the headset, and the headset presents the processing results to the user in the form of voice broadcast or visualization.
[0013] As a further preferred embodiment of the present technical solution, the system further includes a connection mode collaborative control module, and the connection mode collaborative control module is used to monitor the connection status of Bluetooth and WiFi in real time and record abnormal event information.
[0014] As a further preferred embodiment of the present technical solution, the data generation and transmission module has intelligent network switching and optimization functions when automatically starting the WiFi module to establish a connection with the local AI server.
[0015] The embodiment of the present invention has the following advantages due to the adoption of the above technical solution: 1. The present invention uses a local AI server to process data. After the headset receives confirmation information, it automatically starts the WiFi module to establish a connection with the local AI server and transmits the audio data to the local AI server. Local processing avoids the wake-up problem caused by poor network, and there is no need to transmit data to the cloud, which reduces the risk of privacy leakage and effectively solves network-related technical problems. 2. The present invention uses Bluetooth low energy technology to establish an initial connection with the smart device for transmitting simple control signaling. After obtaining and verifying the wake-up signal, the WiFi module is used to connect to the local AI server to transmit data. The WiFi transmission bandwidth is large and can meet the data transmission requirements during complex command interaction, thereby improving data transmission efficiency and solving problems in connection mode and data transmission. 3. The present invention monitors the connection status of Bluetooth and WiFi in real time through the connection mode collaborative control module, and has intelligent network switching and optimization functions, ensuring that corresponding measures are automatically taken when the connection mode is switched or the network fluctuates to avoid wake-up interruptions and data transmission anomalies, providing users with a stable and smooth local AI interaction experience.
[0016] The above summary is for illustrative purposes only and is not intended to be limiting in any way. In addition to the illustrative aspects, embodiments and features described above, further aspects, embodiments and features of the present invention will be readily apparent by reference to the accompanying drawings and the following detailed description. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the drawings required for use in the embodiments or technical descriptions will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0018] Figure 1 A schematic diagram of a flow chart of a local AI wake-up interaction method for a dual-mode headset of the present invention; Figure 2 A schematic diagram of a process flow of a method for generating processing results according to the present invention; Figure 3 A schematic diagram of a flow chart of a method for preprocessing an acquired wake-up signal according to the present invention; Figure 4 This is a schematic diagram of the functional modules of a local AI wake-up interaction system for a dual-mode headset of the present invention. DETAILED DESCRIPTION
[0019] The embodiments of the present disclosure are described in detail below with reference to the accompanying drawings.
[0020] It should be clear that the following embodiments of the present disclosure are described by specific specific examples, and those skilled in the art can easily understand other advantages and effects of the present disclosure from the contents disclosed in this specification. Obviously, the described embodiments are only a part of the embodiments of the present disclosure, rather than all of the embodiments. The present disclosure can also be implemented or applied through other different specific embodiments, and the details in this specification can also be modified or changed in various ways based on different viewpoints and applications without departing from the spirit of the present disclosure. It should be noted that the following embodiments and features in the embodiments can be combined with each other in the absence of conflict. Based on the embodiments in the present disclosure, all other embodiments obtained by ordinary technicians in the field without making creative work are within the scope of protection of the present disclosure.
[0021] It should be noted that various aspects of the embodiments within the scope of the appended claims are described below. It should be apparent that the aspects described herein may be embodied in a wide variety of forms, and any specific structure and / or function described herein is merely illustrative. Based on the present disclosure, it should be understood by those skilled in the art that an aspect described herein may be implemented independently of any other aspect, and two or more of these aspects may be combined in various ways. For example, any number of aspects described herein may be used to implement the device and / or practice the method. In addition, other structures and / or functionalities other than one or more of the aspects described herein may be used to implement this device and / or practice this method.
[0022] It should also be noted that the illustrations provided in the following embodiments are only schematic illustrations of the basic concept of the present disclosure. The drawings only show components related to the present disclosure rather than being drawn according to the number, shape and size of components in actual implementation. In actual implementation, the type, quantity and proportion of each component may be changed arbitrarily, and the component layout may also be more complicated.
[0023] Additionally, in the following description, specific details are provided to facilitate a thorough understanding of the examples. However, it will be understood by those skilled in the art that the aspects described may be practiced without these specific details.
[0024] Figure 1 is a flowchart of a dual-mode headset local AI wake-up interaction method according to an embodiment of the present invention. It should be noted that if there are substantially the same results, the method of the present application is not limited to Figure 1 The process sequence shown is limited. Figure 1-Figure 3 As shown: A dual-mode headset local AI wake-up interaction method, comprising the following steps: Step 1: Turn on the headset and initialize it, and establish a Bluetooth connection with the smart device based on Bluetooth low energy technology; Specifically, when the user turns on the headset power, the power management module inside the headset first powers the Bluetooth Low Energy (BLE) chip and related circuits, and the BLE chip starts to perform hardware self-test to check whether the internal registers, memory, RF circuits, etc. are working properly. For example, it detects the signal strength, frequency stability and other parameters of the RF transmission and receiving circuits to ensure that the hardware is in good condition; After the hardware self-test passes, the BLE chip loads the built-in firmware program, which initializes the Bluetooth protocol stack, including setting the Bluetooth working mode (such as broadcast mode, scanning mode, etc.), communication parameters (such as transmission power, frequency channel, etc.), and loading the pre-stored device configuration information, such as the Bluetooth name of the headset, the device unique identifier (UUID), etc., which will be used in the subsequent broadcast and connection process; After initialization, the headset enters the broadcast mode. The BLE chip prepares broadcast data according to the preset broadcast interval (such as every 100 milliseconds). The broadcast data contains basic information of the headset, such as the device name (which can be customized by the user), device type (identified as a headset) and UUID. In order to save power, the length of the broadcast data is usually optimized and only contains necessary information; The headset transmits broadcast data at a specific frequency (such as the 2.4GHz band) through the RF antenna. The broadcast signal spreads to the surrounding space with a certain transmission power and covers a certain range. The transmission power will be adjusted according to the design and application scenarios of the headset to ensure the signal coverage while minimizing power consumption. After turning on the Bluetooth function, smart devices (such as mobile phones) will periodically start the scanning process. The Bluetooth module of the mobile phone will listen to the surrounding Bluetooth broadcast signals on a specific frequency channel. The scanning period can be adjusted according to user settings or system defaults, generally ranging from several times to dozens of times per second; When the mobile phone receives the broadcast signal of the headset, the Bluetooth module will demodulate and decode the signal and extract the broadcast data. Then, the mobile phone system will parse the extracted data and identify the device name, device type and other information. If the headset is a paired device, the mobile phone will identify the identity of the device based on the stored pairing information. If the user chooses to establish a connection with the headset on the mobile phone, the mobile phone will send a connection request to the headset. The connection request contains the device information of the mobile phone and connection parameters (such as connection interval, slave delay, etc.). The connection interval determines the time interval for data transmission between the headset and the mobile phone. The slave delay allows the headset to delay responding to the request of the mobile phone within a certain range to save power. After receiving the connection request, the headset will verify the request, including the legitimacy of the request, the identity information of the mobile phone, etc. If the verification is successful, the headset sends a connection response to the mobile phone, indicating that it agrees to establish the connection; In the connection response, the headset can adjust the connection parameters sent by the mobile phone. The two parties determine the final connection parameters through negotiation. For example, if the current battery level of the headset is low, it will request to increase the connection interval to reduce the frequency of data transmission and reduce power consumption. When both parties reach an agreement on the connection parameters, a Bluetooth low-power connection is formally established between the headset and the mobile phone. At this time, the headset and the mobile phone can transmit data through the connection, such as simple control signaling transmission, to prepare for subsequent wake-up signal interaction.
[0025] Step 2: Based on the built-in voice recognition module of the headset, the wake-up signal is obtained by continuously monitoring specific wake-up words or operation instructions; Specifically, the voice recognition module in the headset is mainly composed of a microphone, an audio processing chip and related circuits. After powering on, the microphone starts to power on and preheat to stabilize its working state to ensure that it can accurately capture the surrounding sound signals. The audio processing chip initializes its own registers, caches, etc. to prepare for subsequent audio signal processing; the voice recognition module loads pre-stored voice recognition algorithms and models, which are obtained through training with a large amount of voice data and can recognize specific voice patterns. At the same time, it loads template data of specific wake-up words and operation instructions pre-set by the user. These templates are customized based on the user's personalized needs, and the wake-up words and instructions of different users are different; The microphone continuously collects sound signals from the surrounding environment at a certain sampling frequency (such as 16kHz or 44.1kHz), and converts the analog sound signals into digital audio signals. In order to ensure the sound quality, the collected audio signals will be initially amplified and filtered to remove some high-frequency noise and interference; the audio processing chip extracts features from the collected digital audio signals, and uses feature extraction algorithms such as Mel Frequency Cepstral Coefficient (MFCC) to convert the audio signals into a series of parameter vectors that can represent speech features. These feature vectors contain information such as the pitch, timbre, and speech speed of the speech, which is the basis for subsequent speech recognition; The extracted feature vector is compared with the pre-stored template features of specific wake-up words and operation instructions in real time. Dynamic time warping (DTW) or a matching algorithm based on deep learning is used to calculate the similarity between the feature vector and the template features. When the similarity exceeds the preset threshold, it is determined that a specific wake-up word or operation instruction has been detected. When a wake-up word or operation instruction is detected, the voice recognition module will perform a secondary confirmation to avoid misrecognition, and further verify the accuracy of the recognition result by re-analyzing factors such as the context information of the audio signal and the continuity of the voice; after confirming that a valid wake-up word or operation instruction has been detected, the voice recognition module generates a corresponding wake-up signal. The wake-up signal is usually a digital signal that contains information such as the recognized instruction type and timestamp; the generated wake-up signal is sent to the control module of the headset to prepare for subsequent signal preprocessing and transmission. At the same time, in order to save power consumption, when the wake-up signal is not detected, the voice recognition module will maintain a low-power operation state and only perform necessary sound collection and simple feature extraction.
[0026] Step 3: pre-process the acquired wake-up signal, and send the pre-processed wake-up signal to the smart device through the Bluetooth module; Specifically, the wake-up signal obtained by the headset is an analog audio signal, which is first converted into digital form by the built-in ADC, discretized into a digital audio sequence according to a specific sampling frequency and quantization bit number, and then adapted to the data format and converted into the preset standard format of the headset system; then, the adaptive filtering algorithm is used to estimate the noise, and the noise is removed by the noise reduction filter to improve the signal clarity; then, the acoustic feature extraction algorithm is used to extract the speech feature parameters, and the principal component analysis and other methods are used for feature selection and dimensionality reduction; After preprocessing, the wake-up signal is encapsulated according to the Bluetooth communication protocol, a protocol header is added, and data is segmented if necessary. Before sending, the Bluetooth connection status is monitored. If it is not good, try to reconnect or adjust the transmission power, and negotiate the connection parameters. Finally, through Bluetooth low-power technology, a specific modulation method is used to convert the signal into a radio frequency signal and transmit it to the smart device. After receiving it, the smart device sends back a confirmation message. If the headset does not receive it within the specified time, it will retransmit it to ensure reliable data transmission.
[0027] Step 4: After receiving the wake-up signal from the headset, the smart device verifies the wake-up signal and feeds back confirmation information to the headset; Specifically, first, the Bluetooth module of the smart device continuously monitors, and after receiving the wake-up signal sent by the headset, it first demodulates the RF signal into a digital signal, and then extracts the wake-up signal data according to the protocol; then the wake-up signal is verified in many aspects, including checking the data integrity and structure in the format, matching the wake-up word in the content, verifying the legitimacy of the instruction, and performing encryption verification and device identity authentication in terms of security; if all verifications are passed, the wake-up signal is determined to be valid; otherwise, the verification is determined to have failed; After the verification is completed, the smart device generates a confirmation message based on the result. If the verification is successful, it contains a pass mark and status code, and if it fails, it contains a reason code. The confirmation message is encapsulated according to the Bluetooth protocol, and the Bluetooth connection status with the headset is checked. After ensuring that the connection is stable, the confirmation message is sent to the headset through Bluetooth low-power technology. If no response is received within the timeout, it will be retransmitted.
[0028] Step 5: After receiving the confirmation information, the headset encrypts and compresses the wake-up signal to generate audio data, and automatically starts the WiFi module to establish a connection with the local AI server to transmit the audio data to the local AI server; Specifically, after the headset receives the confirmation information sent by the smart device, it begins to encrypt and compress the wake-up signal. When encrypting, it uses methods such as the AES algorithm to encrypt the wake-up signal to prevent the data from being stolen or tampered with during transmission, thereby protecting user privacy and data security. For compression, it uses technologies such as AAC encoding to reduce the amount of data and improve transmission efficiency. After encryption and compression, audio data suitable for transmission is generated. Next, the headset automatically starts the WiFi module, which searches for surrounding WiFi signals and selects the optimal WiFi network corresponding to the local AI server for connection based on preset signal strength, network stability and other rules. After the connection is successful, the encrypted and compressed audio data is transmitted to the local AI server through the WiFi connection. During the transmission process, the headset will also monitor the network status in real time. If there are network fluctuations, signal weakening, etc., it will automatically take corresponding measures, such as adjusting the transmission rate, retransmitting data, etc., to ensure that the audio data is transmitted to the local AI server completely and accurately.
[0029] Step 6: After receiving the audio data, the local AI server analyzes and processes the received audio data based on computing power and local data resources, and generates processing results; Specifically, since the received audio data is encrypted and compressed, the local AI server first uses the corresponding decryption algorithm (such as the AES decryption algorithm corresponding to the headset encryption) to decrypt the data, and then uses the decompression algorithm (such as the AAC decoding algorithm corresponding to the headset compression) to restore the original audio signal. The restored audio signal has some interference or non-standard format, and the server will perform standardization processing on it, such as adjusting the audio sampling rate, number of channels and other parameters to make it meet the requirements of subsequent analysis and processing; Then, based on speech recognition technology, the server converts the audio signal into text information, and then uses natural language processing technology to perform lexical analysis, syntactic analysis and semantic understanding on the text information to identify the type of user instructions, which usually include natural language processing instructions (such as querying information, chatting, etc.), image recognition instructions (if related external devices are involved), and multimedia control instructions (such as playing music, adjusting volume, etc.); Then, according to the identified instruction type, the local AI server calls the corresponding processing module and local data resources for in-depth processing; if it is a natural language processing instruction, it uses a large-scale language model and a local knowledge base to perform semantic understanding, knowledge retrieval and logical reasoning, and integrates relevant information to generate processing results; if it is an image recognition instruction, it performs feature extraction and pattern matching on the image data, compares and analyzes it with the local image database, and obtains the recognition result; for multimedia control instructions, it interacts with local multimedia devices or software, performs corresponding control operations, and generates result information based on the operation feedback; Finally, the server formats the results according to the display and voice broadcast capabilities of the headset. For example, it adjusts the text results to a format suitable for speech synthesis, or succinctly describes the image recognition results. At the same time, it removes redundant information, highlights key content, improves the readability and usability of the results, and finally forms a complete processing result.
[0030] Step 7: The local AI server transmits the generated processing results back to the headset, and the headset presents the processing results to the user in the form of voice broadcast or visualization; Specifically, first, after the local AI server completes the processing of the audio data and generates the results, it will use the established WiFi connection to transmit the processing results back to the headset; before transmitting back, the server will encapsulate the processing results and add the necessary protocol header information, including the data type, length, and the identification of the headset device, etc., to ensure the accuracy and integrity of the data during the transmission process; if the amount of data in the processing result is large and exceeds the limit of a single WiFi transmission, the server will segment the data and add a sequence number and identification to each segment to facilitate the headset to reassemble it after receiving it; Next, the headset receives the processing result data from the local AI server through the WiFi module. After receiving the data, the headset first decapsulates the data and extracts the processing result information; if the processing result is in voice form, the headset directly calls the built-in speech synthesis engine to convert the voice data into a sound signal, and plays it to the user through the speaker to realize the voice broadcast function; if the processing result is visual information, such as text, charts, etc., the headset will transmit this information to the visual display module. If the headset itself has a display screen, it will be displayed directly on the screen; if the headset has no display screen, it will be connected to a smart device (such as a mobile phone) and visualized with the help of the screen of the smart device, and finally the processing result will be intuitively displayed to the user.
[0031] In one embodiment, specifically: in step six, the method for generating a processing result includes the following steps: Step 601: decompress and decrypt the encrypted and compressed audio data to restore it to the original audio signal; Specifically, the local AI server identifies the compression algorithm used for the received data. For example, if the headset previously used AAC encoding for compression, the server will call the corresponding AAC decoding algorithm; during the decompression process, the server reads the compression identifier, data structure and other information in the compressed data, and gradually decompresses it according to the algorithm rules, restores the encoded audio samples, frequency information, etc. in the compressed data, and restores the original time series and amplitude information of the audio data to form the initially decompressed audio data; data verification is also performed during the decompression process to ensure the integrity and accuracy of the decompressed data. If data errors or losses are found, the error handling mechanism will be triggered, requiring the headset to retransmit the data; After decompression is completed, the server calls the corresponding AES decryption algorithm for decryption according to the encryption method agreed upon in advance with the headset, such as AES algorithm encryption; the server inputs the key used for encryption and converts the encrypted audio data into plain text through a specific decryption operation; during the decryption process, the decryption result will be verified, such as checking whether the header information and checksum of the decrypted data are correct, to ensure the reliability of the decrypted data; After decompression and decryption, the data obtained by the server still needs to undergo some format conversion and repair operations before it can be restored to the original audio signal; for example, adjusting the audio sampling rate, number of channels and other parameters to make them consistent with the original audio signal; repairing existing transmission errors or interference, removing abnormal audio samples or noise signals, and finally obtaining the original audio signal that can be used for subsequent speech recognition and command analysis.
[0032] Step 602: Convert the audio signal into text information based on speech recognition technology, and analyze the text information through natural language processing technology to identify the type of user instruction; Specifically, the server uses a special audio feature extraction algorithm to extract key features from the original audio signal; for example, the Mel Frequency Cepstral Coefficient (MFCC) algorithm can simulate the human auditory system's perception of sound frequency and convert the audio signal into a set of parameters that can reflect the characteristics of speech. These parameters cover information such as the frequency, amplitude, and formant of the speech, and constitute the basic data for speech recognition. The extracted audio features are compared with the pre-trained acoustic model in the server. The acoustic model is trained with a large amount of speech data and contains the probability relationship between various speech feature combinations and corresponding pronunciation units (such as phonemes). When the input audio features have a high degree of match with a certain pronunciation unit pattern in the acoustic model, the pronunciation corresponding to the audio can be preliminarily determined. In order to improve the accuracy of speech recognition, the server will introduce a language model, which is trained based on a large amount of text corpus. It can predict the next pronunciation based on the recognized pronunciation, thereby correcting the errors that occur in the acoustic model matching process; for example, when the acoustic model recognizes the two pronunciations of "shui guo", the language model will judge that it is "fruit" based on common vocabulary combinations, rather than other homophone combinations; by continuously combining the results of the acoustic model and the language model, the server gradually converts the audio signal into coherent text information; for the converted text information, the server first performs lexical analysis; it will divide the text into independent words or morphemes, and determine the part of speech of each word, such as noun, verb, adjective, etc.; for example, for the sentence "play Zhou Moumou's songs", the lexical analysis will split it into "play" (verb), "Zhou Moumou" (noun), "of" (particle), "song" (noun); based on the lexical analysis, the server performs syntactic analysis, constructs the grammatical structure tree of the text, and determines the subject, predicate, object and other components of the sentence by analyzing the grammatical relationship between words. For the above sentence, "Play" is the predicate, "song" is the object, and "Zhou's" is the attributive modifying "song"; combining the results of lexical and syntactic analysis, the server uses semantic understanding technology to understand the actual meaning of the text, and identifies the type of user instructions by matching it with pre-set instruction patterns and semantic knowledge bases. If the text contains words related to multimedia operations such as "play", "pause", and "next song", and involves multimedia objects such as music and video, it can be judged as a multimedia control instruction; if the text is about asking for information, seeking knowledge answers, etc., such as "What will the weather be like tomorrow", it is classified as a natural language processing instruction; if the instruction involves the processing or recognition of an image, such as "Identify the animals in this picture", it is an image recognition instruction. Through these steps, the server can accurately identify the type of user instruction and provide a basis for subsequent in-depth processing.
[0033] Step 603: According to the user instruction type, the corresponding processing module and local data resources are called to perform in-depth processing and generate processing results; Specifically, if the user's command is to ask for information, such as "Which is the highest mountain in the world?", the server will call the local knowledge base (including knowledge in multiple fields such as geography, history, and science) for retrieval. The server uses natural language processing technology to perform semantic analysis on the question, extract key information (such as "the highest mountain in the world"), and then search for matching answers in the knowledge base. For some complex questions, the server also needs to reason and integrate information. For example, if you ask about "the impact of a certain historical event," the server needs to integrate the relevant background, process, and other information of the event in the knowledge base, and use logical reasoning to generate a comprehensive answer. When the command is a chat dialogue, the server will call the dialogue management module, which generates appropriate replies based on the dialogue history and current user input. It will use sentiment analysis technology to determine the user's emotional tendency (positive, negative, neutral, etc.) and adjust the tone and content of the reply based on the emotion. At the same time, it uses context understanding technology to ensure that the reply is logically coherent with the entire dialogue. If the user's instruction involves image recognition, such as "identify the animal in this picture", the server will call the image recognition processing module; first, the image is preprocessed, such as resizing and normalizing; then, the image features, such as texture, shape, color, etc., are extracted using algorithms such as convolutional neural networks (CNN); the extracted features are matched with feature templates in the local image database to find the image category with the highest similarity, thereby determining the object in the picture; after identifying the object in the image, the server will further analyze its attributes. For example, if it is identified as a cat, the server will analyze the cat's breed, coat color, posture and other attributes, and combine local image annotation data and knowledge to generate a detailed image description and recognition results; For multimedia control instructions, such as "play music" and "pause video", the server will call the multimedia control module, which interacts with the local multimedia device or software and performs corresponding operations according to the instructions, for example, sending a play instruction to the music player to control it to start playing music; or sending a pause instruction to the video player software to pause it; in the process of executing the operation, the server will obtain the status information of the multimedia device or software, for example, the name and progress of the currently playing music, the resolution and playing time of the video, etc. This information is integrated into the processing results so as to provide feedback to the user; During the processing process, the server will continuously optimize and adjust the processing strategy, and update and improve the local data resources and processing modules according to the accuracy of the processing results and user feedback to improve the processing effect and user experience.
[0034] Step 604: formatting and optimizing the processing result according to the display and voice broadcasting capabilities of the headset; Specifically, if the processing result is in text form, the server will first check the text content. For complex sentence structures and professional terms, the server will simplify and convert them, replace professional vocabulary with easy-to-understand expressions, and split complex long sentences into simple short sentences; for "Please binarize the image", it will be converted to "Change the image to black and white". At the same time, the grammatical structure of the text will be adjusted to make it more in line with daily spoken expression habits and enhance user understanding; Taking into account the speech synthesis effects of different headphones, the server will adjust the speech synthesis parameters according to the characteristics of the speech synthesis engine of the headphones. For some headphones with good speech synthesis effects, a higher speech rate and rich intonation changes can be set to improve the efficiency and fun of the broadcast; while for headphones with limited speech synthesis capabilities, the speech rate will be appropriately reduced to avoid unclear speech and ensure that users can clearly hear the broadcast content; To help users better understand the voice content, the server will add pauses and prompt sounds at appropriate locations, add appropriate pauses between sentences and paragraphs, simulate the rhythm of human speech, and add specific prompt sounds for important information, such as reminders, key command results, etc., to attract users' attention and ensure that users do not miss important content; If the headset has a display screen, the server will obtain information such as the resolution and size of the headset screen. For the text content in the processing result, the server will reasonably adjust the font size, line spacing, and paragraph layout according to the screen width and height to avoid the text being too large to be displayed completely, or too small to be read easily. For image content, the server will scale the image according to the screen size to ensure that the image can be fully displayed on the screen while maintaining the clarity and quality of the image. The processing results are converted into a format suitable for the headset display. If it is tabular data, the server will convert it into a concise list or chart format so that it can be clearly displayed on a small screen. For multiple types of data (such as text, pictures, and charts), the server will perform reasonable layout optimization and arrange them according to importance and logical relationships to make the interface more beautiful and concise, allowing users to quickly obtain information. Due to the limited space on the headset display, the server will streamline the processing results, remove redundant information, and retain only the core content. For key information, such as important values, event names, etc., they are highlighted by changing the font color, making them bold, and increasing the font size, so that users can quickly focus on the key content in a short time.
[0035] In one embodiment, specifically: in step 2, the speech recognition module built into the headset runs based on a local lightweight model or a cloud lightweight model; Among them, the operation mode and wake-up word setting of the built-in voice recognition module of the headset have certain flexibility. The voice recognition module can operate based on a local lightweight model or a cloud lightweight model; If a local lightweight model is used, its advantage is that it does not rely on the network, can work offline, has a fast response speed, and can recognize the user's voice in a timely manner. The local lightweight model has been carefully optimized, occupies less headset memory, has low calculation volume, and can effectively reduce headset power consumption and extend battery life. It is deployed in the headset in advance, and after receiving the voice signal, it can quickly complete the recognition processing locally; The cloud-based lightweight model uses the powerful computing power and rich data resources of the cloud to achieve more accurate speech recognition. Although it relies on network connection, it can accurately recognize complex speech when the network conditions are good. When encountering special speech, dialects or new words that are difficult for local lightweight models to handle, the cloud-based lightweight model can provide better recognition results through continuous updating and learning. The specific wake-up word is set according to the user's own needs; regarding the specific wake-up word, the user can freely set it according to his own needs. Each user has different language habits and usage scenarios. You can set your own familiar and easy-to-pronounce words as the wake-up word. For example, if the user likes personalization, you can set the wake-up word to your own exclusive words; if the usage scenario is relatively noisy, you can set clear and highly recognizable words to ensure that the headset can be accurately woken up even in a complex environment. This personalized setting enhances the convenience and comfort of user interaction with the headset.
[0036] In one embodiment, specifically: in step three, the method for preprocessing the acquired wake-up signal includes the following steps: Step 301: Convert the wake-up signal in audio form into a digital signal; Specifically, the wake-up signal obtained by the headset is usually in the form of an analog audio signal. In order to facilitate subsequent digital signal processing, analog-to-digital conversion is required. The built-in analog-to-digital converter (ADC) of the headset plays a role and works according to the set sampling frequency and quantization bit number. Common sampling frequencies are 16kHz, 44.1kHz, etc. The higher the sampling frequency, the more precise the capture of the sound signal and the higher the degree of restoration. The quantization bit number is 16 bits or 24 bits. The higher the quantization bit number, the higher the accuracy of the sound amplitude that can be represented. The ADC discretizes the continuous analog audio signal in time and amplitude, and converts it into a series of digital samples to form a digital audio sequence, so that the subsequent digital signal processing algorithm can operate on it.
[0037] Step 302: Perform noise reduction processing on the digital signal based on a noise reduction algorithm to remove environmental noise interference; Specifically, in actual usage scenarios, the wake-up signal is often accompanied by environmental noise, such as the sound of vehicles on the street, the sound of conversations indoors, etc. In order to improve the quality and accuracy of the wake-up signal, noise reduction processing is required; the headphones use the least mean square error (LMS) algorithm or the recursive least squares (RLS) algorithm in the adaptive filtering algorithm. These algorithms adjust the parameters of the filter in real time based on the statistical characteristics of the signal and historical data, and estimate the environmental noise. The noise reduction purpose is achieved by subtracting the estimated noise component from the noisy digital signal. The frequency domain filtering method is also used to attenuate the frequency components of the noise according to the frequency characteristics of the noise, so as to make the wake-up signal clearer and reduce the interference of noise on subsequent processing.
[0038] Step 303: extract features from the processed digital signal to extract key features related to the wake-up word; Specifically, the digital signal that has been processed with noise reduction still contains a large amount of information. In order to improve recognition efficiency and accuracy, it is necessary to extract key features related to the wake-up word. The headset uses an acoustic feature extraction algorithm, such as the Mel-frequency cepstral coefficient (MFCC) algorithm. The MFCC algorithm simulates the human auditory system's perception of sound frequencies and converts the audio signal into a set of parameters that can reflect speech characteristics. These parameters include information such as the pitch, timbre, and speaking speed of the speech, and can effectively distinguish different speech contents. By performing MFCC feature extraction on the processed digital signal, feature vectors related to the wake-up word are obtained. These feature vectors serve as the key basis for subsequent recognition of the wake-up word, reducing the amount of data and improving the speed and accuracy of wake-up word recognition.
[0039] In one embodiment, specifically: in step 4, the smart device pre-stores a verification key and a verification rule agreed upon with the headset; the wake-up signal is verified by comparing the wake-up signal with the pre-stored verification key and verification rule; if the comparison result matches, it is determined that the wake-up signal verification is passed, and confirmation information is fed back to the headset; if the comparison result does not match, it is determined that the verification fails, no confirmation information is fed back, and an error prompt information may be sent to the headset; The verification key is a special encrypted code, and the verification rule specifies how to verify the wake-up signal, such as the signal format and content requirements. When the smart device receives the wake-up signal sent by the headset, it uses the verification key to compare the wake-up signal according to the established verification rules; if the identity information, encryption features, etc. carried by the wake-up signal completely match the verification key and rules stored in the smart device, the smart device determines that the wake-up signal verification has passed; at this time, it will feedback confirmation information to the headset. After receiving the confirmation information, the headset knows that the wake-up signal has been correctly received, and can then perform subsequent encryption compression, connection to the local AI server and other operations; If the comparison results do not match, the smart device will determine that the verification has failed; in this case, it will not feedback confirmation information to the headset to prevent the headset from performing incorrect operations; the smart device can also choose to send error prompt information to the headset. This information includes the reason for the verification failure, such as "key error" and "signal format abnormality". After receiving the error prompt, the headset can remind the user of the problem or re-perform the wake-up operation to ensure that the entire interaction process is accurate.
[0040] In one embodiment, specifically: in step 5, when the wake-up signal is encrypted and compressed, the wake-up signal is encrypted and compressed using the AES algorithm and AAC encoding; Among them, the Advanced Encryption Standard (AES) is a symmetric encryption algorithm that uses the same key for encryption and decryption operations. When the headset encrypts the wake-up signal, it will generate a key of a specific length (such as 128 bits, 192 bits or 256 bits), divide the wake-up signal according to a certain data block size (usually 128 bits), and apply the encryption round function of the AES algorithm to each data block in turn. These round functions include byte replacement, row shift, column mixing and round key addition operations. Through multiple rounds of complex transformations, the original wake-up signal data is converted into ciphertext, making it difficult to steal and crack the data during transmission. The AES algorithm has a high degree of security. Its complex encryption process and long key length can effectively resist various common attack methods, such as brute force cracking, differential attack and linear attack. In actual applications, even if the attacker intercepts the encrypted wake-up signal, due to the lack of the correct key, it is impossible to restore the original wake-up signal content, thereby ensuring the privacy of users and the security of data. Advanced Audio Coding (AAC) is an efficient audio coding standard. It is based on perceptual audio coding technology. By analyzing the characteristics of the human auditory system, it removes redundant information in the audio signal that is difficult for the human ear to perceive. When AAC encodes the wake-up signal, the audio signal is first converted from the time domain to the frequency domain, and the audio signal is decomposed into components of different frequencies. Then, according to the masking effect of human hearing, the different frequency components are quantized and encoded to retain the audio information that is important to the human ear perception, while greatly reducing the amount of data. AAC encoding can better preserve the audio quality while ensuring a high compression ratio. Compared with traditional audio encoding formats, AAC encoding can provide higher sound quality at the same bit rate, or use a lower bit rate at the same sound quality requirements, thereby effectively reducing the amount of data in the wake-up signal, which is very important for headphones to transmit data via WiFi. A smaller amount of data can reduce transmission time and network bandwidth occupancy, and improve transmission efficiency. By combining the encryption of the AES algorithm and the compression of the AAC encoding, the headset can improve the efficiency of data transmission while ensuring the security of the wake-up signal, laying a good foundation for subsequent interaction with the local AI server.
[0041] In one embodiment, specifically: the user instruction types include natural language processing instructions, image recognition instructions and multimedia control instructions; Natural language processing instructions: This type of instruction mainly revolves around language-related interactions and information processing; users use voice input to let the device perform operations such as knowledge questions and answers, text translation, and semantic understanding. For example, if a user asks "What is the diameter of the earth", the device will use its built-in knowledge base and natural language processing technology to understand the question and search for relevant information, and then give an accurate answer; or if the user says "Translate this English paragraph into Chinese", the device will analyze and translate the input English text and feedback the results to the user; this type of instruction relies on powerful language models and semantic analysis capabilities to achieve natural and fluent language communication with users; Image recognition instructions: When the user issues such instructions, the device will process image-related tasks, including analyzing and identifying the image content of photos, screenshots, etc. For example, if the user says "Identify the type of animal in this picture", the device will use the image recognition algorithm to extract and match the features of the objects in the picture, and then determine the type of animal and inform the user; or "Find all the cars in this picture", the device will scan the picture, find objects that meet the characteristics of cars, and present them to the user in the form of tags or text descriptions; image recognition instructions are widely used in security monitoring, smart album management, object detection and other fields; Multimedia control instructions: These instructions are mainly used to control the operation of multimedia devices or software. Users can use voice commands to control music playback, video playback, volume adjustment, etc. For example, if a user says "play Zhou's songs", the device will search for Zhou's songs in its multimedia library and start playing; or "turn up the volume", the device will increase the volume accordingly. In addition, operations such as pause, fast forward, and fast rewind can be performed to achieve convenient control of multimedia content and provide users with a more relaxed entertainment experience.
[0042] In summary, an embodiment of the present invention provides a local AI wake-up interaction method for dual-mode headphones. When the headphones are turned on, an initial connection is established with a smart device based on Bluetooth low energy technology. A built-in voice recognition module is used to continuously monitor specific wake-up words or operation instructions based on a local or cloud lightweight model to obtain a wake-up signal. The wake-up signal is pre-processed, verified, encrypted and compressed in sequence and then transmitted to a local AI server. The server decompresses, decrypts, voice recognizes, analyzes the command type, performs deep processing on the audio data based on computing power and local data resources, and formats the optimization results. Finally, the processing results are transmitted back to the headphones to be presented to the user in the form of voice broadcast or visualization. This method ensures the security, accuracy and efficiency of the interaction and improves the user experience.
[0043] Figure 4 : is a functional module diagram of a dual-mode headset local AI wake-up interaction system according to an embodiment of the present application, such as Figure 4 As shown, a dual-mode headset local AI wake-up interaction system, the system includes: a Bluetooth connection module, a signal acquisition module, a signal preprocessing module, a signal verification module, a data generation and transmission module, a data analysis and processing module and a result presentation module; The Bluetooth connection module is configured to turn on and initialize the headset and establish a Bluetooth connection with the smart device based on Bluetooth low energy technology; The signal acquisition module is configured to acquire a wake-up signal by continuously monitoring a specific wake-up word or operation instruction based on a voice recognition module built into the headset; The signal preprocessing module is configured to preprocess the acquired wake-up signal and send the preprocessed wake-up signal to the smart device through the Bluetooth module; The signal verification module is configured to verify the wake-up signal after the smart device receives the wake-up signal from the headset, and feedback confirmation information to the headset; The data generation and transmission module is configured to encrypt and compress the wake-up signal after the headset receives the confirmation information, generate audio data, and automatically start the WiFi module to establish a connection with the local AI server to transmit the audio data to the local AI server; The data analysis and processing module is configured to analyze and process the received audio data based on computing power and local data resources after the local AI server receives the audio data, and generate a processing result; The result presentation module is configured so that the local AI server transmits the generated processing results back to the headset, and the headset presents the processing results to the user in the form of voice broadcast or visualization.
[0044] In one embodiment, specifically: the system further includes a connection mode collaborative control module, and the connection mode collaborative control module is used to monitor the connection status of Bluetooth and WiFi in real time and record abnormal event information.
[0045] In one embodiment, specifically: the data generation and transmission module has intelligent network switching and optimization functions when automatically starting the WiFi module to establish a connection with the local AI server.
[0046] In summary, an embodiment of the present invention provides a dual-mode headset local AI wake-up interaction system, which establishes an initial connection with a smart device based on Bluetooth low energy technology through a Bluetooth connection module when the headset is turned on. The signal acquisition module uses the built-in voice recognition module of the headset to monitor specific wake-up words or operation instructions to obtain a wake-up signal. The signal preprocessing module preprocesses the wake-up signal and sends it to the smart device via Bluetooth. The signal verification module verifies the signal and feeds back confirmation information. The data generation and transmission module encrypts and compresses the wake-up signal after the headset receives the confirmation information, and automatically starts the WiFi module to establish a connection with the local AI server and transmit audio data. The module also has intelligent network switching and optimization functions. The data analysis and processing module analyzes and processes the audio data to generate results. The result presentation module transmits the results back to the headset and presents them to the user in voice or visual form. In addition, the system is also provided with a connection mode collaborative control module for real-time monitoring of the Bluetooth and WiFi connection status and recording abnormal event information, thereby realizing an efficient, stable and secure interaction process as a whole.
[0047] For other details about the technical solutions for implementing each module in a local AI wake-up interaction system of a dual-mode headset in the above-mentioned embodiment, please refer to the description of a local AI wake-up interaction method of a dual-mode headset in the above-mentioned embodiment, which will not be repeated here.
[0048] It should be noted that each embodiment in this specification is described in a progressive manner, and each embodiment focuses on the differences from other embodiments, and the same or similar parts between the embodiments can be referred to each other. For system embodiments, since they are basically similar to method embodiments, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment.
[0049] The basic principles of the present disclosure are described above in conjunction with specific embodiments. However, it should be noted that the advantages, strengths, effects, etc. mentioned in the present disclosure are only examples and not limitations, and it cannot be considered that these advantages, strengths, effects, etc. are required by each embodiment of the present disclosure. In addition, the specific details disclosed above are only for the purpose of illustration and ease of understanding, and are not limitations. The above details do not limit the present disclosure to the necessity of adopting the above specific details to be implemented.
[0050] In the present disclosure, relational terms such as first and second, etc. are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply that there is any such actual relationship or order between these entities or operations. The block diagrams of the devices, devices, equipment, and systems involved in the present disclosure are only illustrative examples and are not intended to require or imply that they must be connected, arranged, and configured in the manner shown in the block diagram. As will be appreciated by those skilled in the art, these devices, devices, equipment, and systems can be connected, arranged, and configured in any manner. Words such as "including", "comprising", "having", etc. are open words, referring to "including but not limited to", and can be used interchangeably with them. The words "or" and "and" used here refer to the words "and / or" and can be used interchangeably with them, unless the context clearly indicates otherwise. The words "such as" used here refer to the phrase "such as but not limited to", and can be used interchangeably with them.
[0051] Additionally, as used herein, "or" used in a list of items beginning with "at least one" indicates a separate list, so that, for example, a list of "at least one of A, B, or C" means A or B or C, or AB or AC or BC, or ABC (i.e., A and B and C). Furthermore, the word "exemplary" does not mean that the example described is preferred or better than other examples.
[0052] It should also be noted that in the system and method of the present disclosure, each component or each step can be decomposed and / or recombined. Such decomposition and / or recombination should be regarded as equivalent solutions of the present disclosure.
[0053] Various changes, substitutions, and modifications of the techniques described herein may be made without departing from the teachings defined by the appended claims. Furthermore, the scope of the claims of the present disclosure is not limited to the specific aspects of the processes, machines, manufactures, compositions of events, means, methods, and actions described above. Currently existing or later to be developed processes, machines, manufactures, compositions of events, means, methods, or actions that perform substantially the same functions or achieve substantially the same results as the corresponding aspects described herein may be utilized. Thus, the appended claims include such processes, machines, manufactures, compositions of events, means, methods, or actions within their scope.
[0054] The above description of the disclosed aspects is provided to enable any person skilled in the art to make or use the present disclosure. Various modifications to these aspects will be readily apparent to those skilled in the art, and the general principles defined herein may be applied to other aspects without departing from the scope of the present disclosure. Therefore, the present disclosure is not intended to be limited to the aspects shown herein, but rather to the widest scope consistent with the principles and novel features disclosed herein.
[0055] The above description has been given for the purpose of illustration and description. In addition, this description is not intended to limit the embodiments of the present disclosure to the forms disclosed herein. Although multiple example aspects and embodiments have been discussed above, those skilled in the art will recognize certain variations, modifications, changes, additions and sub-combinations thereof.
Claims
1. A dual-mode headset local AI wake-up interaction method, characterized in that: The following steps are involved: The headset is turned on and initialized, and a Bluetooth connection is established with the smart device based on Bluetooth low energy technology; Based on the built-in voice recognition module of the headset, the wake-up signal is obtained by continuously monitoring specific wake-up words or operation instructions; Preprocess the acquired wake-up signal, and send the preprocessed wake-up signal to the smart device through the Bluetooth module; After receiving the wake-up signal from the headset, the smart device verifies the wake-up signal and feeds back confirmation information to the headset; After receiving the confirmation information, the headset encrypts and compresses the wake-up signal to generate audio data, and automatically starts the WiFi module to establish a connection with the local AI server to transmit the audio data to the local AI server; After receiving the audio data, the local AI server analyzes and processes the received audio data based on computing power and local data resources, and generates processing results; The local AI server transmits the generated processing results back to the headset, and the headset presents the processing results to the user in the form of voice broadcast or visualization.
2. A dual-mode headset local AI wake-up interaction method according to claim 1, characterized in that: The method for generating a processing result comprises the following steps: Decompress and decrypt the encrypted and compressed audio data to restore it to the original audio signal; Convert audio signals into text information based on speech recognition technology, and analyze the text information through natural language processing technology to identify the type of user instructions; According to the user instruction type, call the corresponding processing module and local data resources for deep processing and generate processing results; The processing results are formatted and optimized according to the display and voice broadcast capabilities of the headset.
3. A dual-mode headset local AI wake-up interaction method according to claim 1, characterized in that: The built-in speech recognition module of the headset runs based on a local lightweight model or a cloud lightweight model, and the specific wake-up word is set according to the user's own needs.
4. A dual-mode headset local AI wake-up interaction method according to claim 1, characterized in that: The method for preprocessing the acquired wake-up signal comprises the following steps: Converting the wake-up signal in the form of audio into a digital signal; Perform noise reduction processing on digital signals based on noise reduction algorithms to remove environmental noise interference; Perform feature extraction on the processed digital signal to extract key features related to the wake-up word.
5. A dual-mode headset local AI wake-up interaction method according to claim 1, characterized in that: The smart device pre-stores a verification key and verification rules agreed upon with the headset; the wake-up signal is verified by comparing the wake-up signal with the pre-stored verification key and verification rules; if the comparison result matches, the wake-up signal is determined to have been verified, and confirmation information is fed back to the headset; if the comparison result does not match, the verification is determined to have failed, no confirmation information is fed back, and an error prompt information may be sent to the headset.
6. A dual-mode headset local AI wake-up interaction method according to claim 1, characterized in that: When the wake-up signal is encrypted and compressed, the AES algorithm and AAC encoding are used to encrypt and compress the wake-up signal.
7. A dual-mode headset local AI wake-up interaction method according to claim 2, characterized in that: The user instruction types include natural language processing instructions, image recognition instructions and multimedia control instructions.
8. A dual-mode headset local AI wake-up interaction system, applied to a dual-mode headset local AI wake-up interaction method according to any one of claims 1-7, characterized in that: The system includes: a Bluetooth connection module, a signal acquisition module, a signal preprocessing module, a signal verification module, a data generation and transmission module, a data analysis and processing module and a result presentation module; The Bluetooth connection module is configured to turn on and initialize the headset and establish a Bluetooth connection with the smart device based on Bluetooth low energy technology; The signal acquisition module is configured to acquire a wake-up signal by continuously monitoring a specific wake-up word or operation instruction based on a voice recognition module built into the headset; The signal preprocessing module is configured to preprocess the acquired wake-up signal and send the preprocessed wake-up signal to the smart device through the Bluetooth module; The signal verification module is configured to verify the wake-up signal after the smart device receives the wake-up signal from the headset, and feedback confirmation information to the headset; The data generation and transmission module is configured to encrypt and compress the wake-up signal after the headset receives the confirmation information, generate audio data, and automatically start the WiFi module to establish a connection with the local AI server to transmit the audio data to the local AI server; The data analysis and processing module is configured to analyze and process the received audio data based on computing power and local data resources after the local AI server receives the audio data, and generate a processing result; The result presentation module is configured so that the local AI server transmits the generated processing results back to the headset, and the headset presents the processing results to the user in the form of voice broadcast or visualization.
9. A dual-mode headset local AI wake-up interaction system according to claim 8, characterized in that: The system also includes a connection mode collaborative control module, which is used to monitor the connection status of Bluetooth and WiFi in real time and record abnormal event information.
10. A dual-mode headset local AI wake-up interaction system according to claim 8, characterized in that: The data generation and transmission module has intelligent network switching and optimization functions when automatically starting the WiFi module to establish a connection with the local AI server.
Citation Information
Cited By
Voice control method and electronic equipment
CN121506126A