Intelligent interaction method and device of TWS earphone and storage medium
By combining acoustic pre-activation and distributed voiceprint verification with adaptive learning, the problems of slow response and insufficient privacy protection in TWS earphones are solved, achieving fast response and intelligent offline experience optimization.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-29
- Publication Date
- 2026-04-10
AI Technical Summary
Existing voice interaction solutions for TWS earphones suffer from slow response, inefficient use of computing resources, insufficient granularity of privacy protection, and a significant decline in interactive experience in offline environments.
Through acoustic pre-activation, distributed voiceprint verification, and adaptive learning, a multi-microphone array is used for sound source localization. IMU sensor data is combined to assist in judging pre-activation commands, distributed voiceprint verification dynamically determines command paths, and the offline skill library is updated through adaptive learning.
It enables rapid response in noisy environments, improves the response speed of voice interaction, optimizes privacy protection and offline experience, and ensures rapid response and intelligent updates for high-frequency personalized needs.
Smart Images

Figure CN121838751A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of electric digital data processing, in particular to a smart interaction method and device of TWS earphones and a storage medium. BACKGROUND
[0002] With the deep integration of artificial intelligence technology and wearable devices, True Wireless Stereo (TWS) earphones have evolved from simple audio playback devices to important artificial intelligence interaction portals. Users expect to interact conveniently and intelligently with earphones through natural speech to complete complex tasks such as information query, device control, and personal assistance. This trend has put high demands on the real-time performance, accuracy, privacy security, and reliability of TWS earphones in voice interaction under different network conditions.
[0003] To improve the interaction experience and protect user privacy, existing technologies propose some local AI processing solutions, which specifically include: establishing an initial connection with a smart device through Bluetooth for wake-up verification, and then starting a WiFi module to transmit audio data to a local AI server for processing. This solution aims to use a local server to avoid network delays and privacy risks caused by cloud processing, and by using Bluetooth and WiFi dual-mode collaboration, it takes into account low-power standby and high-speed data transmission.
[0004] However, the present inventors have found through in-depth research that the above-mentioned existing technical solutions still have the following defects: slow response or even failure, low utilization efficiency of computing resources, insufficient privacy protection granularity, and a significant decline in interaction experience in a network-free environment. SUMMARY
[0005] The present application provides a smart interaction method, device and storage medium of TWS earphones, which realizes fast response, hierarchical privacy protection and offline experience optimization through acoustic pre-activation, distributed voiceprint verification and adaptive learning.
[0006] In one aspect, the present application provides a smart interaction method of TWS earphones, the method comprising: continuously collecting environmental sound through a multi-microphone array of the earphones, identifying acoustic features in the environmental sound based on a sound source positioning algorithm, and determining whether to generate a pre-activation instruction accordingly; in response to the pre-activation instruction, locking a wake-up word candidate segment through a first-stage keyword detection model at the earphone end, and extracting voiceprint features of the wake-up word candidate segment; sending the voiceprint features and the wake-up word candidate segment to a smart device through a low-latency wireless channel, performing distributed voiceprint verification by a second-stage verification model of the smart device in combination with a pre-stored voiceprint model, and generating a voiceprint confidence score; dynamically decide an execution path of the user instruction based on the voiceprint confidence score and a noise level in the real-time monitored environmental sound, wherein the execution path comprises at least one of a local processing unit, a smart device processing unit, or a cloud processing unit, and the decision follows a distribution principle of localizing high-privacy instructions and clouding complex instructions; after the user instruction is executed, the analysis module analyzes an interaction scenario and a frequency of the instruction through an adaptive learning engine, and solidifies, through an incremental learning technique, a recognition and execution logic of a high-frequency instruction to a local cache to update an offline skill library of the earphone.
[0007] In another aspect, the present application provides a smart interaction device of a TWS earphone, the device comprising: The recognition module is configured to continuously collect environmental sound through a multi-microphone array of the earphone, and determine whether to generate a pre-activation instruction to improve an algorithm power priority of a local processing unit and adjust a power consumption mode based on a sound source positioning algorithm. The extraction module is configured to, in response to the pre-activation instruction, lock a wake-up word candidate segment through a first-stage keyword detection model at the earphone end, and extract a voiceprint feature of the wake-up word candidate segment. The generation module is configured to send the voiceprint feature and the wake-up word candidate segment to a smart device through a low-latency wireless channel, and generate a voiceprint confidence score through a second-stage verification model of the smart device in combination with a pre-stored voiceprint model for distributed voiceprint verification. The decision module is configured to dynamically decide an execution path of the user instruction based on the voiceprint confidence score and a noise level in the real-time monitored environmental sound, wherein the execution path comprises at least one of a local processing unit, a smart device processing unit, or a cloud processing unit, and the decision follows a distribution principle of localizing high-privacy instructions and clouding complex instructions. The analysis module is configured to, after the user instruction is executed, analyze an interaction scenario and a frequency of the instruction through an adaptive learning engine, and solidify, through an incremental learning technique, a recognition and execution logic of a high-frequency instruction to a local cache to update an offline skill library of the earphone.
[0008] In a third aspect, the present application provides a device, which comprises a memory, a processor, and a computer program stored in the memory and executable on the processor, and the processor implements the steps of the technical solution of the smart interaction method of the TWS earphone when executing the computer program.
[0009] In a fourth aspect, the present application provides a storage medium, which stores a computer program, and the computer program implements the steps of the technical solution of the smart interaction method of the TWS earphone when executed by a processor.
[0010] From the above technical solutions provided by the present application, on the one hand, through the pre-activation mechanism based on sound source positioning, the system computing power priority and power consumption mode can be improved in advance before the user actually speaks the wake-up word by detecting the signs of preparing to make a sound (for example, specific change trend of sound source direction and energy), which effectively overcomes the inherent delay problem of the traditional passive response mode, so that fast wake-up and response can be realized even in a noisy environment; on the other hand, a distributed voiceprint verification and decision mechanism is adopted, the extraction of voiceprint features, fast comparison and complex verification are distributed and completed cooperatively at the earphone and the smart device, and the execution path (local, device end or cloud end) of the instruction is dynamically decided based on the high confidence score generated by the verification and the environmental noise level, which breaks the dependence on a single local server, uses the voiceprint verification result as a key signal for dynamic routing, so that high-privacy-sensitive instructions can be preferentially allocated to local or trusted device end processing, and complex non-sensitive tasks are allocated to cloud computing processing, which not only reduces the data amount transmitted to the cloud and reduces the risk of privacy leakage, but also realizes precise adaptation and efficient use of computing resources; thirdly, an adaptive learning engine is introduced after the execution of the instruction, which can continuously analyze the user's interaction scene and instruction frequency, and through incremental learning technology, the high-frequency and mature instruction recognition and execution logic is solidified into the local cache, so that the offline skill library of the earphone can grow and update continuously, so that the user's high-frequency personalized needs can also be quickly and accurately responded to in a network-free environment, which makes the earphone more and more intelligent, effectively solving the problem of great gap between offline experience and online experience in the prior art. In summary, the technical solutions of the present application realize fast response, hierarchical privacy protection and offline experience optimization through acoustic pre-activation, distributed voiceprint verification and adaptive learning. BRIEF DESCRIPTION OF DRAWINGS
[0011] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the drawings needed to be used in the embodiments or prior art description will be briefly introduced. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor.
[0012] Figure 1 is a flowchart of the intelligent interaction method of the TWS earphone provided by the embodiments of the present application; Figure 2 is a structural schematic diagram of the intelligent interaction device of the TWS earphone provided by the embodiments of the present application; Figure 3 is a structural schematic diagram of the device provided by the embodiments of the present application. DETAILED DESCRIPTION
[0013] With reference to the drawings and in the embodiments herein, the technical solutions in the embodiments will be clearly and completely described, obviously, the described embodiments are only part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative labor fall within the scope of protection of the present application.
[0014] In the present specification, adjectives such as first and second can be used merely to differentiate one element or action from another element or action, without necessarily implying any actual such relationship or order. Where the context permits, reference to an element or component or step (etc.) by the indefinite article "a" or "an" does not exclude the possibility that more than one element, component, or step is present.
[0015] In the present specification, for the convenience of description, the sizes of various parts shown in the drawings are not drawn in accordance with the actual proportional relationship.
[0016] To improve the interactive experience and protect user privacy, for voice interaction of true wireless stereo (TWS), the prior art proposes some local AI processing schemes, specifically including: establishing an initial connection with a smart device through Bluetooth to perform wake-up verification, and then starting a WiFi module to transmit audio data to a local AI server for processing. This scheme aims to use a local server to avoid network delay and privacy risks caused by cloud processing, and by using Bluetooth and WiFi dual-mode cooperation, it takes into account low-power standby and high-speed data transmission. However, the present inventors have found through in-depth research that the above-mentioned prior art scheme still has several defects that need to be solved. First, its interactive process is a passive response mode of "detection - transmission - verification", and the entire link delay from wake-up word detection to server processing is high, especially in noisy environments, the wake-up word detection itself is easily disturbed, resulting in slow response or even failure; second, its voiceprint verification and subsequent processing completely depend on the local AI server, which is a centralized processing architecture, and all voice data needs to be transmitted to the server, not only increasing the risk of privacy leakage of data transmission, but also the processing path is fixed and rigid, and cannot be flexibly and dynamically allocated according to the privacy sensitivity of the instructions, resulting in low utilization efficiency of computing resources and insufficient privacy protection granularity; finally, this scheme lacks self-evolution ability and cannot optimize offline capabilities according to user usage habits, resulting in a significant decline in interactive experience in a network-free environment.
[0017] In view of the above problems of the prior art, the present application proposes a smart interaction method for TWS earphones, the flow chart of which is shown in FIG. 1, mainly including steps S101-S105, which are described in detail as follows: Figure 1 Step S101: continuously collect environmental sound through the multi-microphone array of the earphone, identify the acoustic characteristics in the environmental sound based on a sound source positioning algorithm, and determine whether to generate a pre-activation instruction accordingly.
[0018] In a noisy environment, the cold start time of a voice interaction system such as an earphone from a low-power sleep state to a full wake-up and effective detection is too long, resulting in a poor user experience. To solve this problem, one existing solution is continuous full-power monitoring, that is, the earphone is always in a high-performance mode and continuously performs wake-up word detection. Another solution is simple threshold triggering, that is, only when the environmental sound intensity exceeds a fixed threshold, the power consumption of the voice interaction system such as the earphone is increased. For the continuous full-power monitoring solution, the earphone battery life is dramatically shortened, which cannot meet the basic requirement of low power consumption for TWS earphones. The simple threshold triggering solution has poor anti-interference ability. Specifically, in a noisy environment (for example, a subway, a coffee shop, etc.), the environmental noise itself can easily exceed the threshold, resulting in frequent false activation. Not only cannot it save power, but also increases additional power consumption due to frequent state switching, and seriously affects the user experience.
[0019] Therefore, the present application continuously collects environmental sound through the multi-microphone array of the earphone, identifies the acoustic characteristics in the environmental sound based on a sound source positioning algorithm, and determines whether to generate a pre-activation instruction to improve the algorithmic priority of the local processing unit and adjust the power consumption mode. On the one hand, this solution can more accurately predict user intent by analyzing more specific acoustic characteristics such as sound source direction and energy change trend, greatly reducing the false activation rate. On the other hand, the pre-activation instruction is a preparation signal generated internally by the system when it detects signs that the user may soon speak (rather than having spoken a specific wake-up word). This signal is used to mobilize computing resources in advance, allowing the system to enter a "warm-up" state from a low-power monitoring state, thereby shortening the delay of subsequent formal responses. Therefore, the voice interaction system such as the earphone is ready before the user speaks the wake-up word, and can respond instantly when it is really needed, achieving a smooth transition from cold start to hot standby and improving the response speed of the voice interaction system.
[0020] As an embodiment of the present application, identifying the acoustic characteristics in the environmental sound based on a sound source positioning algorithm and determining whether to generate a pre-activation instruction can be: tracking the change amplitude of the sound source direction angle in real time through the beamforming data of the multi-microphone array; when the sound source direction angle changes continuously towards the earphone wearer and the audio energy aggregation rate exceeds a threshold value within a preset time window, it is determined that the acoustic characteristics meet the pre-activation characteristics and a pre-activation instruction is generated.
[0021] Specifically, the amplitude of the change in the direction angle of the sound source can be tracked in real time by beamforming data of a multi-microphone array through steps S1011 to S1016: step S1011, synchronously collecting environmental audio signals through each microphone in the multi-microphone array to obtain a plurality of synchronous original audio data streams; step S1012, preprocessing the plurality of synchronous original audio data streams, including analog-to-digital conversion, sampling rate alignment, and pre-filtering, to eliminate hardware differences between channels and suppress out-of-band noise; step S1013, applying a beamforming algorithm to the preprocessed multi-channel audio data, adjusting the phase and amplitude weight of each channel signal to form a directional receiving beam in space, and calculating the spatial power spectrum of the beamformer in different directions; step S1014, scanning the spatial power spectrum of the beamformer in different directions to identify the azimuth angle corresponding to the power spectrum peak, and determining the azimuth angle as the estimated value of the direction angle of the sound source at the current time; step S1015, repeating steps S1011 to S1014 on consecutive time frames to obtain a time series of the estimated value of the direction angle of the sound source; and step S1016, performing smoothing filtering processing (for example, using Kalman filtering or moving average algorithm) on the time series of the estimated value of the direction angle of the sound source to suppress transient jitter, and calculating the amplitude and trend of the change in the direction angle between adjacent time frames.
[0022] Further, in order to overcome the unreliability problem caused by relying only on acoustic signals for sound source positioning and intention judgment in complex real environments, the above sound source positioning algorithm can further include: fusing the IMU sensor data built-in the earphone, assisting to verify the authenticity of the change in the direction of the sound source by analyzing the user's head turning or throat muscle vibration signals, to reduce the probability of false activation.
[0023] Specifically, the fusion of the inertial measurement unit (IMU) data built-in the earphone, through the analysis of the user's head turning or throat muscle vibration signal, can assist in verifying the authenticity of the sound source direction change. The authenticity of the sound source direction change can be achieved in the following way: real-time reading of inertial measurement data from the IMU built-in the earphone, which at least contains three-axis accelerometer and three-axis gyroscope data; processing the inertial measurement data, and calculating the attitude change information of the earphone and the rotational angular velocity of the user's head through sensor fusion algorithm; synchronously acquiring the sound source direction change data; time-space alignment and correlation analysis of the rotational angular velocity of the head and the sound source direction change data, and calculating the coupling degree between them; based on the coupling degree, the authenticity of the sound source direction change is verified: when the head rotational angular velocity and the sound source direction change are highly consistent in direction and time sequence, it is determined that the sound source direction change is caused by the user's active head movement, and the authenticity is high; when the two are inconsistent or the head does not rotate, it is determined that the sound source direction change may be caused by environmental interference, and the authenticity is low; the authenticity verification result is output to the pre-activation judgment logic as a weighted factor, which is used to reduce the false activation probability caused by environmental interference.
[0024] The technical solution of the above embodiment has the following aspects: on the one hand, the IMU directly measures the physical movement of the earphone body, i.e. the user's head turning, nodding and other actions, which is the first physical evidence about the user's intention and is not affected by the acoustic environment. When the acoustic sensor detects the sound source direction change, the IMU data can provide evidence of whether the user's head has a matching physical movement. If the two are highly consistent in time and space, the credibility of the pre-activation signal is improved, otherwise, the acoustic signal should be treated with suspicion, thereby reducing the false activation probability caused by acoustic illusion; on the other hand, in a noisy and reverberation serious environment, the performance of the acoustic positioning algorithm will decrease, while the IMU data provides an independent and reliable information source. Even if the acoustic signal quality is poor, the system can still partially rely on the user's active behavior pattern (e.g. habitually turning the head to align the microphone) captured by the IMU to assist in decision-making, thereby maintaining the usability and stability of the system in complex scenarios and improving the lower limit of user experience; on the third aspect, analyzing the vibration signal of the throat muscle (which can be indirectly captured by a high-precision IMU sensor) is a more advanced vocalization intention indicator than the acoustic signal. It can capture the subtle activity of the user's throat before preparing to speak, combined with the acoustic direction change and head turning, it constitutes a complete intention judgment evidence chain from "physiological preparation" to "physical orientation" to "acoustic performance", making the pre-activation judgment have higher accuracy and confidence from the physiological signal level.
[0025] Step S102: in response to the pre-activation instruction, the first-stage keyword detection model at the earphone end locks the wake-up word candidate segment, and extracts the voiceprint features of the wake-up word candidate segment.
[0026] Step S103: send the voiceprint feature of the wake-up word candidate segment and the wake-up word candidate segment to the smart device through the low-delay wireless channel, and generate a voiceprint confidence score by the second-stage verification model of the smart device in combination with a pre-stored voiceprint model.
[0027] In related technologies, centralized processing of identity verification procedures often brings privacy risks, delays, and insufficient flexibility, and many other problems. For example, a pure cloud verification scheme uploads complete audio data to a cloud server for voiceprint recognition, which has significant network delay and high risk of privacy leakage of transmitted complete voice data; for another example, a pure local (earphone end) verification scheme stores a voiceprint model in the earphone and completes all verification calculations, which is limited by the limited computing power and storage space of the earphone end, and the complexity and accuracy of the verification model are limited, making it difficult to guarantee security and accuracy, and the device end cannot take advantage of stronger computing power.
[0028] To this end, the scheme adopted by the present application is to lock the wake-up word candidate segment through the first-stage keyword detection model of the earphone end in response to a pre-activation instruction, extract the voiceprint feature of the wake-up word candidate segment, send the voiceprint feature of the wake-up word candidate segment and the wake-up word candidate segment to the smart device through the low-delay wireless channel, and generate a voiceprint confidence score by the second-stage verification model of the smart device in combination with a pre-stored voiceprint model. The above-mentioned scheme of the present application places lightweight feature extraction on the earphone end and complex model verification on the smart device, which takes advantage of the computing power of the device end and avoids remote transmission of sensitive data. The wake-up word candidate segment refers to a segment of continuous audio data that the earphone end initially determines may contain a preset wake-up word (such as "Xiao X Xiao X"), which is a "possible object" that needs to be further verified subsequently. The voiceprint confidence score refers to a quantitative similarity value (for example, between 0 and 1) obtained by comparing the current speaker's voice feature with the pre-registered owner's voice model. This score is used to represent the likelihood of "the current speaker being the owner". The generated voiceprint confidence score provides a quantitative and reliable basis for subsequent dynamic decision-making, rather than just a binary judgment of "pass" or "fail".
[0029] In an embodiment of the present application, the first-stage keyword detection model adopts a Connectionist Temporal Classification (CTC) model to realize fast rough screening of the wake-up word, the second-stage verification model adopts a U2-KWS attention mechanism to perform fine-grained verification on the candidate segment, and the outputs of the two-stage models jointly participate in the joint decision of the voiceprint features. Specifically, the first-stage keyword detection model adopting a Connectionist Temporal Classification (CTC) model to realize fast rough screening of the wake-up word can be: the continuous audio signal collected by the microphone is preprocessed by framing and windowing, and the acoustic features (such as Mel Frequency Cepstral Coefficients MFCC or Filter Bank Features FBANK) of each frame of audio are extracted to form an acoustic feature sequence; the acoustic feature sequence is input into a pre-trained CTC model, and the output layer of the CTC model defines phoneme or character units related to the wake-up word; the CTC model encodes and decodes the input acoustic feature sequence, automatically learns the alignment relationship between the acoustic features and the output units, does not require pre-accurate phoneme boundary information, and directly outputs the probability distribution of each type of phoneme / character corresponding to each time step; the probability distribution sequence output by the CTC model is decoded (such as using greedy decoding or beam search decoding), the path with the highest probability is combined to generate a possible character or phoneme sequence; the sequence obtained by decoding is matched with the pre-defined wake-up word text, and when the sequence matching degree exceeds a preset first threshold, it is determined that the audio in this time period is a wake-up word candidate segment, and the start and end time stamps of the segment are output.
[0030] As for the second-stage verification model adopting a U2-KWS attention mechanism to perform fine-grained verification on the candidate segment, and the outputs of the two-stage models jointly participating in the joint decision of the voiceprint features, it can specifically be: the U2-KWS model uses its internal attention mechanism to calculate the importance weight of all time frames of the wake-up word candidate segment, thereby focusing on the most critical speech frames for wake-up word discrimination and suppressing irrelevant or noise frames; based on the attention weight, the features of the speech frames are weighted and summed to obtain a deep feature vector that can represent the core information of the candidate segment and has stronger robustness; at the same time, a voiceprint feature vector is extracted from the wake-up word candidate segment; the deep feature vector (reflecting content accuracy) output by the U2-KWS model and the voiceprint feature vector (reflecting speaker identity) are fused (for example, by vector splicing or feature addition) to form a comprehensive feature representation; based on the comprehensive feature after fusion, the final joint decision is made: if the matching degrees of the comprehensive feature with the target wake-up word and the target voiceprint are both high, it is determined to be an effective wake-up and identity confirmation; otherwise, it is rejected.
[0031] As an embodiment of the present application, the second-stage verification model of the intelligent device combines the pre-stored voiceprint model to perform distributed voiceprint verification, and the generation of the voiceprint confidence score can be: using the U2-KWS attention mechanism to perform attention weighting on the received wake-up word candidate segment, focusing on the key speech frames; fusing the weighted speech frame features with the voiceprint feature vector received from the earphone end, and then performing similarity matching with the encrypted and stored user voiceprint template to generate the voiceprint confidence score. It should be noted that after the weighted speech frame features are fused with the voiceprint feature vector received from the earphone end, the more the similarity value is, the higher the matching degree is, and the system will map it to a higher voiceprint confidence score (for example, close to 1.0 or 100%), indicating that "the current speaker is most likely the machine owner"; the smaller the similarity value is, the lower the matching degree is, and the lower the voiceprint confidence score is (for example, close to 0), indicating that "the current speaker is most likely not the machine owner". It should be noted that the voiceprint confidence score is a continuous quantitative value, rather than a simple "yes" or "no" judgment, which provides a more fine and flexible basis for subsequent dynamic decision-making.
[0032] In the above embodiment, the voiceprint features of the wake-up word candidate segment and the wake-up word candidate segment are transmitted to the intelligent device through a low-delay wireless channel, specifically, the voiceprint features of the wake-up word candidate segment and the wake-up word candidate segment are transmitted from the earphone end to the intelligent device with a delay of less than 30 ms through the star flash technology. Further, the low-delay wireless channel supports dual-mode dynamic switching of star flash and Bluetooth, that is, when the star flash signal strength of the star flash technology is lower than -70 dBm, it is automatically switched to Bluetooth connection; when transmitting high-bandwidth audio data, the star flash channel of the star flash technology is preferentially enabled, and the link stability is monitored through the heartbeat packet; during the dual-mode dynamic switching process, the audio data that has not been transmitted is temporarily stored in the local buffer after differential encoding compression in the switching gap, and is preferentially retransmitted after the connection is restored.
[0033] Further, Figure 1 The method of the example also includes a privacy protection mechanism, that is, after the voiceprint verification is passed, only the encrypted hash value of the instruction text and the voiceprint signature are uploaded to the cloud; the intermediate data in the local processing process is immediately cleared after the instruction execution is completed, and the clearing operation is performed through a hardware-level security chip.
[0034] Step S104: Based on the voiceprint confidence score and the noise level in the real-time monitored environmental sound, dynamically deciding the execution path of the user instruction, wherein the execution path includes at least one of the local processing unit, the intelligent device processing unit or the cloud processing unit, and the decision follows the allocation principle of localizing high-privacy instructions and clouding complex instructions.
[0035] The dynamic decision execution path refers to the system selecting which one of the earphone chip (local), the connected mobile phone / tablet (smart device end), or the Internet server (cloud end) to process the user instruction in real time and flexibly according to the voiceprint confidence, environmental noise, instruction content, and other factors. The core principle is that instructions involving privacy are processed locally or on the device end as much as possible, and instructions requiring complex calculations can be handed over to the cloud. If all user instructions are fixedly sent to the local server or the cloud for processing, or simple routing is based on a single factor, such as only network conditions or instruction type (for example, instructions such as "play music" are allocated to the local, and instructions such as "query the weather" are allocated to the cloud), the former lacks flexibility, that is, the fixed cloud path has privacy risks, the fixed local path cannot process complex instructions, and cannot achieve the balance of performance, privacy, and power consumption; the decision granularity of the simple routing scheme based on a single factor is rough, for example, only the instruction type may not be able to distinguish between low-privacy instructions (such as "play music") and high-privacy instructions (such as "play my bank balance"), and ignoring the environmental noise level will also lead to a decline in voice recognition quality in noisy environments. In order to convert the results of the previous steps (including pre-activation state, voiceprint confidence, environmental information, etc.) into the optimal business processing strategy, and improve the overall efficiency and user experience of the system, the present application can dynamically decide the execution path of the user instruction based on the voiceprint confidence score and the noise level in the real-time monitored environmental sound, wherein the execution path includes at least one of the local processing unit, the smart device processing unit, or the cloud processing unit, and the decision follows the allocation principle of localizing high-privacy instructions and clouding complex instructions. In this way, by intelligently allocating tasks to the most suitable processing unit, fine-grained, context-aware computing resource scheduling is achieved, and at the same time, the principle of data minimization is also implemented, with high-sensitivity instructions being maximally localized, thereby improving the overall security of the system.
[0036] Specifically, as an embodiment of the present application, based on the voiceprint confidence score and the noise level in the real-time monitored environmental sound, the dynamic decision execution path of the user instruction can be: when the voiceprint confidence score is higher than a first threshold, the user's identity is determined to be in a high-trust state, and subsequent non-sensitive instructions are preferentially allocated to the local processing unit or the smart device processing unit for execution; when the voiceprint confidence score is lower than the first threshold but higher than a second threshold, a secondary verification mechanism is triggered, requiring the user to confirm the identity through voice or action; resource adaptive adjustment is performed in combination with the environmental noise level: when the environmental signal-to-noise ratio is lower than a first signal-to-noise ratio threshold, the high-performance mode of the local noise reduction algorithm and the voice processing module is maintained, and when the environmental signal-to-noise ratio is higher than a second signal-to-noise ratio threshold, the power consumption of the non-core module is reduced to optimize the endurance. The resource adaptive adjustment is realized through a dynamic voltage frequency adjustment module, the chip working frequency of the dynamic voltage frequency adjustment module is positively correlated with the noise level, and the power consumption strategy is updated immediately after the instruction allocation decision is completed.
[0037] As described above, the decision follows the allocation principle of localizing high privacy instructions and clouding complex instructions. In the embodiment of the present application, the sensitivity classification of user instructions is based on the following rules: user instructions involving payment and location privacy are defined as high-sensitivity instructions, which are forced to be allocated to the cloud processing unit and transmitted through the TLS 1.3 encrypted channel; multimedia control or user instructions cached in the local skill library are defined as low-sensitivity instructions, which are preferentially allocated to the local processing unit or the smart device processing unit for execution. The sensitivity classification of the above user instructions is realized in real time through natural language processing technology, including: parsing the text content of the user instruction, matching the pre-defined keyword table and semantic rules; dynamically labeling the sensitivity level of the instruction according to the matching result, and triggering the corresponding execution path allocation strategy.
[0038] Step S105: After the execution of the user instruction is completed, the adaptive learning engine is used to analyze the interaction scene and the instruction frequency, and the recognition and execution logic of high-frequency instructions are fixed to the local cache through incremental learning technology to update the offline skill library of the earphone.
[0039] The existing voice interaction system such as earphone has problems of static, rigidity, lack of evolution ability, etc. If a static skill library solution is adopted, i.e. the offline function of the earphone is fixed after the earphone is shipped, the offline experience is poor, the dynamic changing needs of the user cannot be met, and the earphone capacity devalues over time. If a full model update solution is adopted, i.e. the offline capacity is updated by firmware upgrade or downloading a complete new model package, the defects are low efficiency, large resource consumption, the need to download a large amount of data, occupation of valuable storage space, and long update cycle, which cannot adapt to the user habits in real time. In view of the problems of the above related technologies, the solution of the present application is that after the execution of the user instruction is completed, the adaptive learning engine is used to analyze the interaction scene and the instruction frequency, and the recognition and execution logic of high-frequency instructions are fixed to the local cache through incremental learning technology to update the offline skill library of the earphone. This solution enables the earphone to have the ability of continuous learning and self-optimization, which can continuously adapt to the individualized preferences of the user, and through the incremental learning technology, the offline skill library is updated in a lightweight and efficient manner, ensuring seamless connection and consistency between online and offline experiences.
[0040] Specifically, as an embodiment of the present application, by analyzing the interaction scene and instruction frequency through the adaptive learning engine, the recognition and execution logic of high-frequency instructions can be fixed to the local cache through incremental learning technology: record the instruction sequence, execution result and environmental noise data of the user in different scenes to form an interaction log; based on the interaction log, identify high-frequency instruction combinations through clustering algorithm and calculate the feasibility score of offline execution; the high-frequency instruction model with a feasibility score of offline execution higher than the threshold is distilled from the cloud model through incremental learning technology to a lightweight version and stored in the local cache. The incremental learning technology in the above embodiment adopts a combination of gradient clipping and knowledge distillation, and only synchronizes the incremental parameters required for model updating to avoid resource overhead caused by full model transmission.
[0041] In the embodiments of the present application, the instruction response time supported by the offline skill library is less than 100 ms, and when network interruption is detected, it automatically switches to a pure local decision mode and only executes the instructions already cached in the skill library.
[0042] From the above description of the accompanying drawings Figure 1The smart interaction method of the example TWS earphone can know that, on the one hand, through the pre-activation mechanism based on sound source positioning, the system algorithm priority can be raised in advance and the power consumption mode can be adjusted through detecting the signs of the user preparing to speak (for example, specific change trend of sound source direction and energy) before the user actually speaks the wake-up word, the pre-advance resource scheduling and state preparation effectively overcome the inherent delay problem of the traditional passive response mode, so that quick wake-up and response can be realized in a noisy environment; on the other hand, a distributed voiceprint verification and decision mechanism is adopted, the extraction of voiceprint features, fast comparison and complex verification are distributed and completed cooperatively at the earphone and the smart device, and the execution path (local, device end or cloud end) of the instruction is dynamically decided based on the high confidence score generated by the verification and the environmental noise level, this mechanism breaks the dependence on a single local server, takes the voiceprint verification result as a key signal for dynamic routing, so that high-privacy-sensitive instructions can be preferentially allocated to local or trusted device end processing, and complex non-sensitive tasks are allocated to cloud computing processing, which not only reduces the data amount transmitted to the cloud and reduces the risk of privacy leakage, but also realizes precise adaptation and efficient utilization of algorithm resources; the third aspect is that an adaptive learning engine is introduced after the execution of the instruction, which can continuously analyze the user's interaction scene and instruction frequency, and through incremental learning technology, the high-frequency and mature instruction recognition and execution logic is solidified into the local cache, so that the offline skill library of the earphone can grow and update continuously, so that the high-frequency personalized needs of the user can also be quickly and accurately responded to in a network-free environment, which makes the earphone more intelligent, and effectively solves the problem of great gap between offline experience and online experience in the prior art. In summary, the technical scheme of the present application realizes fast response, hierarchical privacy protection and offline experience optimization through acoustic pre-activation, distributed voiceprint verification and adaptive learning.
[0043] Please refer to the accompanying Figure 2 The smart interaction device of the TWS earphone provided by the embodiment of the present application can include an identification module 201, an extraction module 202, a generation module 203, a decision module 204 and an analysis module 205, which are described in detail as follows: The identification module 201 is configured to continuously collect environmental sound through the multi-microphone array of the earphone, and determine whether to generate a pre-activation instruction to raise the algorithm priority of the local processing unit and adjust the power consumption mode based on a sound source positioning algorithm; The extraction module 202 is configured to, in response to the pre-activation instruction, lock the wake-up word candidate segment through the first-stage keyword detection model at the earphone end, and extract the voiceprint features of the wake-up word candidate segment; The generation module 203 is configured to send the voiceprint features of the wake-up word candidate segment and the wake-up word candidate segment to the smart device through a low-delay wireless channel, and generate a voiceprint confidence score by a second-stage verification model of the smart device in combination with a pre-stored voiceprint model through distributed voiceprint verification. The decision module 204 is configured to dynamically decide an execution path of the user instruction based on the voiceprint confidence score and the noise level in the real-time monitored environmental sound, wherein the execution path comprises at least one of the local processing unit, the smart device processing unit or the cloud processing unit, and the decision follows the allocation principle of localizing high-privacy instructions and clouding complex instructions; The analysis module 205 is configured to, after the user instruction is executed, analyze the interaction scene and the instruction frequency by the adaptive learning engine, and solidify the recognition and execution logic of the high-frequency instruction to the local cache by the incremental learning technology, so as to update the offline skill library of the earphone.
[0044] From the above description of the accompanying drawings Figure 2 It can be known from the smart interaction device of the example TWS earphone that, on the one hand, through the pre-activation mechanism based on the sound source positioning, the system computing power priority and the power consumption mode can be raised in advance by detecting the sign of the user's preparation to speak (for example, the specific change trend of the sound source direction and energy) before the user actually speaks the wake-up word, and the pre-prepared resources effectively overcome the inherent delay problem of the traditional passive response mode, so that the fast wake-up and response can be realized even in a noisy environment. On the other hand, the distributed voiceprint verification and decision mechanism is adopted, the extraction, rapid comparison and complex verification of the voiceprint features are distributed and completed cooperatively at the earphone and the smart device, and the execution path (local, device end or cloud end) of the instruction is dynamically decided based on the high-confidence score generated by the verification and the noise level of the environment. This mechanism breaks the dependence on a single local server, uses the voiceprint verification result as a key signal for dynamic routing, so that high-privacy-sensitive instructions can be preferentially allocated to local or trusted device end processing, and complex non-sensitive tasks can be allocated to cloud computing power processing. Not only is the data amount transmitted to the cloud reduced, and the privacy leakage risk is reduced, but also the precise adaptation and efficient use of computing resources are realized. Thirdly, the adaptive learning engine is introduced after the instruction execution, which can continuously analyze the user's interaction scene and instruction frequency, and solidify the high-frequency and mature instruction recognition and execution logic to the local cache by the incremental learning technology, so that the offline skill library of the earphone can grow and update continuously, so that the high-frequency personalized needs of the user can also be quickly and accurately responded to in a network-free environment. This makes the earphone more and more intelligent, and effectively solves the problem of great gap between offline experience and online experience in the prior art. In summary, the technical scheme of the present application realizes fast response, hierarchical privacy protection and offline experience optimization through acoustic pre-activation, distributed voiceprint verification and adaptive learning.
[0045] Figure 3 is a structural schematic diagram of the device provided by an embodiment of the present application. As Figure 3As shown, the device 3 of the embodiment mainly comprises a processor 30, a memory 31, and a computer program 32 stored in the memory 31 and executable on the processor 30, such as a program of the smart interaction method of the TWS earphone. The processor 30 implements the steps in the above-mentioned embodiment of the smart interaction method of the TWS earphone when executing the computer program 32, for example Figure 1 As shown, the steps S101 to S105. Alternatively, the processor 30 implements the functions of each module / unit in the above-mentioned various device embodiments when executing the computer program 32, for example Figure 2 As shown, the functions of the identification module 201, the extraction module 202, the generation module 203, the decision module 204, and the analysis module 205.
[0046] Exemplarily, the computer program 32 of the smart interaction method of the TWS earphone mainly includes: continuously collecting environmental sound through the multi-microphone array of the earphone, identifying the acoustic characteristics in the environmental sound based on a sound source positioning algorithm, and determining whether to generate a pre-activation instruction accordingly; in response to the pre-activation instruction, locking the wake-up word candidate segment through the first-stage keyword detection model at the earphone end, and extracting the voiceprint features of the wake-up word candidate segment; sending the voiceprint features of the wake-up word candidate segment and the wake-up word candidate segment to the smart device through a low-delay wireless channel, performing distributed voiceprint verification by the second-stage verification model of the smart device in combination with a pre-stored voiceprint model, and generating a voiceprint confidence score; based on the voiceprint confidence score and the noise level in the real-time monitored environmental sound, dynamically deciding the execution path of the user instruction, wherein the execution path includes at least one of a local processing unit, a smart device processing unit or a cloud processing unit, and the decision follows the allocation principle of localizing high-privacy instructions and clouding complex instructions; after the user instruction is executed, analyzing the interaction scene and the instruction frequency through an adaptive learning engine, and solidifying the recognition and execution logic of high-frequency instructions to the local cache through incremental learning technology to update the offline skill library of the earphone. The computer program 32 can be divided into one or more modules / units, which are stored in the memory 31 and executed by the processor 30 to complete the present application. One or more modules / units can be a series of computer program instruction segments capable of completing a specific function, which are used to describe the execution process of the computer program 32 in the device 3.For example, the computer program 32 can be divided into the functions of the identification module 201, the extraction module 202, the generation module 203, the decision module 204 and the analysis module 205 (modules in the virtual device), and the specific functions of each module are as follows: the identification module 201 is configured to continuously collect environmental sound through a multi-microphone array of the earphone, and determine whether to generate a pre-activation instruction to improve the algorithm priority of the local processing unit and adjust the power consumption mode based on a sound source positioning algorithm; the extraction module 202 is configured to, in response to the pre-activation instruction, lock a wake-up word candidate segment through a first-stage keyword detection model at the earphone end, and extract a voiceprint feature of the wake-up word candidate segment; the generation module 203 is configured to send the voiceprint feature of the wake-up word candidate segment and the wake-up word candidate segment to the smart device through a low-delay wireless channel, and generate a voiceprint confidence score by a second-stage verification model of the smart device in combination with a pre-stored voiceprint model for distributed voiceprint verification; the decision module 204 is configured to dynamically decide an execution path of a user instruction based on the voiceprint confidence score and a noise level in the environmental sound monitored in real time, wherein the execution path includes at least one of the local processing unit, the smart device processing unit or the cloud processing unit, and the decision follows a distribution principle that high-privacy instructions are localized and complex instructions are clouded; and the analysis module 205 is configured to, after the user instruction is executed, analyze an interaction scene and an instruction frequency through an adaptive learning engine, and solidify the recognition and execution logic of high-frequency instructions to a local cache through an incremental learning technology to update an offline skill library of the earphone.
[0047] The device 3 can include but is not limited to a processor 30 and a memory 31. Those skilled in the art can understand that, Figure 3 The device 3 is only an example and does not constitute a limitation on the device 3, and can include more or fewer components than those shown, or combine certain components, or different components, for example, the device can also include an input / output device, a network access device, a bus, etc.
[0048] The processor 30 can be a central processing unit (CPU), and can also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs) or other programmable logic devices, discrete gates or transistor logic components, discrete hardware components, etc. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor.
[0049] The memory 31 can be an internal storage unit of the device 3, such as a hard disk or a memory of the device 3. The memory 31 can also be an external storage device of the device 3, such as a plug-in hard disk, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, and the like equipped on the device 3. Further, the memory 31 can also include both the internal storage unit and the external storage device of the device 3. The memory 31 is used to store computer programs and other programs and data required by the device. The memory 31 can also be used to temporarily store data that has been output or will be output.
[0050] It can be clearly understood by those skilled in the art that, for the convenience and brevity of description, only the above-mentioned division of each functional unit and module is exemplified, and in actual application, the above-mentioned functions can be completed by different functional units and modules according to needs, that is, the internal structure of the device is divided into different functional units or modules to complete all or part of the functions described above. Each functional unit and module in the embodiment can be integrated in one processing unit, or each unit can exist physically, or two or more units can be integrated in one unit, and the integrated unit can be realized in the form of hardware or in the form of software functional unit. In addition, the specific names of each functional unit and module are only for the convenience of mutual distinction, and do not limit the protection scope of the present application. The specific working process of the units and modules in the above device can refer to the corresponding process in the foregoing method embodiments, which will not be described here.
[0051] In the above embodiments, the description of each embodiment has its own emphasis, and the parts not described or recorded in detail in a certain embodiment can be referred to the relevant description of other embodiments.
[0052] Those of ordinary skill in the art can appreciate that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be realized in electronic hardware or a combination of computer software and electronic hardware. Whether the functions are executed in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of the present application.
[0053] In the embodiments of the present application, it should be understood that the disclosed apparatuses / devices and methods can be implemented in other manners. For example, the embodiments of the apparatuses / devices described above are merely illustrative. For example, the division of the modules or units is merely logical function division. There can be another division manner for the actual implementation. For example, a plurality of units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the displayed or discussed mutual couplings or direct couplings or communication connections can be indirect couplings or communication connections through some interfaces, devices or units, and can be in electrical, mechanical or other forms.
[0054] The units described as separated components can or can not be physically separated, and the components displayed as units can or can not be physical units, i.e., can be located in one place, or can be distributed on a plurality of network units. Some or all of the units can be selected according to actual needs to achieve the purposes of the embodiments of the present application.
[0055] In addition, each functional unit in the embodiments of the present application can be integrated in a processing unit, or each unit can exist physically as a separate unit, or two or more units can be integrated in one unit. The integrated unit can be implemented in the form of hardware or in the form of a software functional unit.
[0056] The integrated modules / units, if implemented in the form of software functional units and sold or used as independent products, can be stored in a storage medium. Based on such understanding, all or part of the processes in the above-mentioned embodiment methods can also be completed by a computer program instructing related hardware, and the computer program of the smart interaction method of the TWS earphone can be stored in a storage medium. When the computer program is executed by a processor, the steps of each method embodiment can be implemented, that is, the environmental sound is continuously collected by the multi-microphone array of the earphone, the acoustic characteristics in the environmental sound are identified based on a sound source positioning algorithm, and it is determined whether to generate a pre-activation instruction; in response to the pre-activation instruction, the wake-up word candidate segment is locked by the first-stage keyword detection model at the earphone end, and the voiceprint features of the wake-up word candidate segment are extracted; the voiceprint features of the wake-up word candidate segment and the wake-up word candidate segment are sent to the smart device through a low-delay wireless channel, distributed voiceprint verification is performed by the second-stage verification model of the smart device in combination with a pre-stored voiceprint model to generate a voiceprint confidence score; based on the voiceprint confidence score and the noise level of the environmental sound monitored in real time, a dynamic decision is made on an execution path of a user instruction, wherein the execution path includes at least one of a local processing unit, a smart device processing unit, or a cloud processing unit, and the decision follows the allocation principle of localizing high-privacy instructions and clouding complex instructions; after the user instruction is executed, the interaction scene and the instruction frequency are analyzed by an adaptive learning engine, the recognition and execution logic of high-frequency instructions are fixed to the local cache through incremental learning technology, and the offline skill library of the earphone is updated. The computer program includes computer program code, which can be in the form of source code, object code, executable files, or some intermediate forms, etc. The storage medium can include any entity or device capable of carrying computer program code, recording medium, U disk, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signal, telecommunication signal, and software distribution medium, etc. It should be noted that the content included in the storage medium can be appropriately increased or decreased according to the requirements of legislation and patent practice in the jurisdiction, for example, in some jurisdictions, according to legislation and patent practice, the storage medium does not include electrical carrier signals and telecommunication signals.
[0057] The above examples are only used to illustrate the technical solutions of the present application, but not limit the same; although the present application has been described in detail with reference to the foregoing examples, those skilled in the art should understand that the technical solutions recorded in the foregoing examples can still be modified, or some technical features can be replaced by equivalents; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application, and should be included in the protection scope of the present application. The specific embodiments described above further illustrate the purpose, technical solutions and beneficial effects of the present application. It should be understood that the above description is only a specific embodiment of the present application, and is not used to limit the protection scope of the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application should be included in the protection scope of the present application.
Claims
1. A smart interaction method for TWS earphones, characterized in that, The method includes: The system continuously collects ambient sound through the multi-microphone array of the headphones, identifies the acoustic features in the ambient sound based on the sound source localization algorithm, and determines whether to generate a pre-activation command accordingly. In response to the pre-activation instruction, the first-stage keyword detection model on the earphone end is used to lock the wake word candidate segment and extract the voiceprint features of the wake word candidate segment; The voiceprint features and wake word candidate fragments are sent to the smart device via a low-latency wireless channel. The second-stage verification model of the smart device combines the pre-stored voiceprint model to perform distributed voiceprint verification and generate a voiceprint confidence score. Based on the voiceprint confidence score and the noise level in the real-time monitored ambient sound, the execution path of the user command is dynamically determined. The execution path includes at least one of the local processing unit, the smart device processing unit, or the cloud processing unit, and the decision follows the allocation principle of localizing high privacy commands and cloudifying complex commands. After the user command is executed, the interaction scenario and command frequency are analyzed by the adaptive learning engine. The recognition and execution logic of high-frequency commands are solidified into the local cache through incremental learning technology to update the offline skill library of the headphones.
2. The method according to claim 1, characterized in that, The step of identifying acoustic features in the ambient sound based on the sound source localization algorithm and determining whether to generate a pre-activation instruction includes: The beamforming data from a multi-microphone array is used to track the change in the direction angle of the sound source in real time. When the change in the direction angle of the sound source is continuously directed towards the headphone wearer and the rate at which the audio energy accumulates exceeds a threshold within a preset time window, the acoustic feature is determined to meet the pre-activation feature and a pre-activation instruction is generated.
3. The method according to claim 1, characterized in that, The first-stage keyword detection model uses the CTC model to achieve rapid coarse screening of wake words, while the second-stage verification model uses the U2-KWS attention mechanism to perform fine-grained verification of candidate segments. The outputs of the two stages jointly participate in the joint decision-making of voiceprint features.
4. The method according to claim 3, characterized in that, The distributed voiceprint verification, which combines the second-stage verification model of the smart device with a pre-stored voiceprint model, generates a voiceprint confidence score, including: The U2-KWS attention mechanism is used to apply attention weights to the received wake word candidate segments, focusing on key speech frames; The weighted speech frame features are fused with the voiceprint feature vector received from the headset, and then similarity matching is performed with the encrypted user voiceprint template to generate a voiceprint confidence score.
5. The method according to claim 1, characterized in that, The dynamic decision-making of the user command execution path based on the voiceprint confidence score and the noise level in the real-time monitored ambient sound includes: When the voiceprint confidence score is higher than the first threshold, the user's identity is determined to be in a high-trust state, and subsequent non-sensitive instructions are preferentially assigned to the local processing unit or the smart device processing unit for execution. When the voiceprint confidence score is lower than the first threshold but higher than the second threshold, a secondary verification mechanism is triggered, requiring the user to confirm their identity through voice or action. Resource adaptive adjustment based on ambient noise level: When the ambient signal-to-noise ratio is lower than the first signal-to-noise ratio threshold, maintain the high-performance mode of the local noise reduction algorithm and voice processing module; when the ambient signal-to-noise ratio is higher than the second signal-to-noise ratio threshold, reduce the power consumption of non-core modules to optimize battery life.
6. The method according to claim 5, characterized in that, The resource adaptive adjustment is achieved through a dynamic voltage and frequency adjustment module. The chip operating frequency of the dynamic voltage and frequency adjustment module is positively correlated with the noise level, and the power consumption strategy is updated immediately after the instruction allocation decision is completed.
7. The method according to claim 1, characterized in that, The step of analyzing interaction scenarios and command frequencies using an adaptive learning engine, and then embedding the identification and execution logic of high-frequency commands into a local cache through incremental learning technology, includes: Record user command sequences, execution results, and environmental noise data in different scenarios to form an interaction log; Based on the interaction logs, a clustering algorithm is used to identify high-frequency instruction combinations and calculate their feasibility score for offline execution. The high-frequency instruction models with feasibility scores above the threshold are distilled from the cloud model into a lightweight version using incremental learning technology and stored in the local cache.
8. A smart interactive device for TWS earphones, characterized in that, The device includes: The recognition module is used to continuously collect ambient sound through the multi-microphone array of the headphones, and determine whether to generate a pre-activation instruction based on the sound source localization algorithm to improve the computing power priority of the local processing unit and adjust the power consumption mode. The extraction module is used to respond to the pre-activation instruction, lock the wake word candidate segment through the first-stage keyword detection model on the earphone end, and extract the voiceprint features of the wake word candidate segment; The generation module is used to send the voiceprint features and wake word candidate fragments to the smart device through a low-latency wireless channel. The second-stage verification model of the smart device combines the pre-stored voiceprint model to perform distributed voiceprint verification and generate a voiceprint confidence score. The decision module is used to dynamically determine the execution path of user instructions based on the voiceprint confidence score and the noise level in the real-time monitored ambient sound. The execution path includes at least one of a local processing unit, a smart device processing unit, or a cloud processing unit, and the decision follows the allocation principle of localizing high privacy instructions and cloudifying complex instructions. The analysis module is used to analyze the interaction scenario and command frequency through an adaptive learning engine after the user command is executed, and to solidify the identification and execution logic of high-frequency commands into the local cache through incremental learning technology in order to update the offline skill library of the headphones.
9. An apparatus comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the steps of the method as described in any one of claims 1 to 7.
10. A storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method as described in any one of claims 1 to 7.