Bluetooth speaker multi-user personalization configuration management system based on user voiceprint recognition

Through deep collaboration between voiceprint recognition and context-aware decision-making modules, the Bluetooth speaker enables personalized configuration management for multiple users, solving the problems of identity recognition and device linkage in multi-user scenarios, and improving user experience and device utilization.

CN122640751APending Publication Date: 2026-08-25SHENZHEN ROYQUEEN AUDIO TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610840085.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-11
Publication Date
2026-08-25

AI Technical Summary

Technical Problem

Existing Bluetooth speakers suffer from limited personalized configuration options in multi-user shared scenarios, inaccurate identification when multiple users' voiceprints conflict, and an inability to manage external Bluetooth environments, resulting in response delays and identity recognition issues.

Method used

The system employs a multi-user personalized configuration management system based on user voiceprint recognition. Through deep collaboration between the voiceprint registration and management module, the context-aware decision-making module, and the comprehensive configuration module, it achieves one-time registration, seamless switching, and full-dimensional configuration. It combines sound source separation, Bluetooth signal strength, and physical button status to determine identity and adjust audio parameters and peripheral devices in a coordinated manner.

Benefits of technology

It enables accurate identification of the current user on Bluetooth speakers, automatic switching of audio parameters and synchronous adjustment of peripheral devices, improving the personalized experience and ease of operation in multi-user scenarios, avoiding identity misjudgment and configuration flickering, and ensuring seamless human-computer interaction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122640751A_ABST
    Figure CN122640751A_ABST
Patent Text Reader

Abstract

The application discloses a Bluetooth audio box multi-user personalized configuration management system based on user voiceprint recognition, relates to the technical field of Bluetooth audio equipment in smart home, and comprises a voiceprint registration and management module, extracts voiceprint features and generates a voiceprint digital fingerprint bound with the unique identity of the user; a context awareness decision module, which is connected with the voiceprint registration and management module, a microphone array of the Bluetooth audio box and a Bluetooth communication unit, respectively, performs weighted fusion decision based on voiceprint matching scores and signal receiving strength indication, and outputs the identity of the current active user; and a comprehensive configuration module, which is connected with the context awareness decision module, an audio digital signal processor of the Bluetooth audio box and the Bluetooth communication unit. The Bluetooth audio box multi-user personalized configuration management system based on user voiceprint recognition can realize audio parameter switching based on user identity and external Bluetooth device state linkage configuration in the scene where multiple users share a Bluetooth audio box.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of Bluetooth audio device technology in smart homes, specifically to a multi-user personalized configuration management system for Bluetooth speakers based on user voiceprint recognition. Background Technology

[0002] With the rapid development of Bluetooth communication technology and MEMS microphone arrays, Bluetooth speakers are no longer just simple music playback terminals, but also important interactive gateways for the Internet of Things in the home. Currently, smart Bluetooth speakers on the market have voiceprint wake-up and voice interaction functions. For example, some devices can realize personalized music recommendations and sound effect switching for specific users through built-in voiceprint recognition modules, and use cloud servers to verify user voiceprints to push related messages. Some existing technologies have also disclosed adaptive tuning methods based on artificial intelligence, which identify users by recognizing voiceprints to automatically load preset tuning configurations, thereby meeting the personalized listening pursuits of different users.

[0003] However, the personalized configurations of some existing Bluetooth speakers mainly focus on audio content recommendations or basic sound effect adjustments. There is still room for improvement in areas such as active user identification in multi-user scenarios, audio parameter switching, and external Bluetooth device linkage configuration. The configuration dimensions are limited: Most of the aforementioned smart Bluetooth speakers or smart voice-controlled home appliances are limited to simple audio content recommendations, playlist pushes, or basic EQ adjustments, lacking comprehensive coverage of all user scenarios. In actual home or office environments, different users not only want to hear their preferred sound effects but also expect the speaker to be able to adjust the color temperature of indoor lights, the room temperature setting of air conditioners, and the opening and closing of curtains, among other ambient lighting devices. Current technology has not yet proposed a solution that can simultaneously meet the needs of audio parameter switching. Bluetooth speaker systems that integrate with the overall atmosphere management of Bluetooth peripheral devices suffer from several shortcomings. Firstly, their robustness in handling multiple user voiceprint conflicts is insufficient. Existing technologies typically only activate accurately when a single user speaks. In complex acoustic scenarios with multiple users, the mixed voiceprint features captured by the microphone array interfere with each other, making it difficult for the system to identify the active user currently using the device. This often results in identity recognition confusion or misconfiguration and intermittent disconnections. Secondly, there is a delay in wake-up processing. Existing technologies require the user to utter a specific wake-up word to trigger the speaker before voiceprint recognition and subsequent configuration loading begin. This wake-up-after-processing significantly delays system response when the user has already actively interacted with the speaker, such as pressing a physical button or entering Bluetooth coverage area, impacting the user's seamless experience.

[0004] Therefore, how to propose a system that can achieve seamless multi-user identification on Bluetooth speakers and simultaneously perform unified configuration of audio parameters based on user identity and peripheral devices in multiple dimensions is a technical problem that urgently needs to be solved by those skilled in the art. Summary of the Invention

[0005] The purpose of this invention is to provide a Bluetooth speaker multi-user personalized configuration management system based on user voiceprint recognition, so as to solve the defects of existing Bluetooth speakers mentioned in the background art, such as limited personalized configuration dimensions in multi-user shared scenarios, inaccurate recognition when multiple voiceprints conflict, and inability to manage the external Bluetooth environment in conjunction with each other.

[0006] To achieve the above objectives, the present invention provides the following technical solution: a Bluetooth speaker multi-user personalized configuration management system based on user voiceprint recognition, comprising: a voiceprint registration and management module, used to collect registration voice samples of at least one user, extract voiceprint features and generate a voiceprint digital fingerprint bound to the user's unique identity. The context-aware decision-making module is connected to the voiceprint registration and management module, the microphone array of the Bluetooth speaker, and the Bluetooth communication unit. This module intercepts the mixed ambient sound collected by the microphone array in real time, performs sound source separation to extract valid user voice frames, extracts the voiceprint features of the voice frames and matches them with the voiceprint digital fingerprint, and at the same time obtains the signal reception strength indication of at least one terminal device that has established a Bluetooth connection with the Bluetooth speaker. Based on the voiceprint matching score and the signal reception strength indication, a weighted fusion decision is made to output the identity identifier of the currently active user. The integrated configuration module is connected to the context-aware decision module, the audio digital signal processor of the Bluetooth speaker, and the Bluetooth communication unit. This module contains an audio parameter adaptive engine and a peripheral device status linkage engine. The audio parameter adaptive engine pre-stores audio equalizer parameters, dynamic range compression thresholds, and virtual surround sound parameters corresponding to different user identities, and sends a switching command to the audio digital signal processor based on the received current active user identity. The peripheral device status linkage engine pre-stores a snapshot of the operating status of at least one external Bluetooth device corresponding to different user identities, and sends a status switching command to the external Bluetooth device through the Bluetooth communication unit based on the received current active user identity. Among them, the voiceprint registration and management module, the context-aware decision-making module, and the comprehensive configuration module work together to enable the same Bluetooth speaker to switch the corresponding audio parameters according to the identity of the currently active user and send the running status switching command to the corresponding external Bluetooth device.

[0007] Furthermore, the context-aware decision-making module employs an attention-based sound source separation algorithm to separate the independent speech frames of each speaker from the mixed sound, and uses a dual-threshold competition and sliding time window dynamic fusion mechanism to determine active users. When the average fusion score of a candidate user within the sliding time window is higher than the first threshold, the candidate user is directly locked as the current active user. When the average fusion score is greater than the second threshold and less than or equal to the first threshold, the historical connection frequency of the candidate user's corresponding terminal device and the current physical button operation state of the Bluetooth speaker are introduced as auxiliary weighting factors to resolve identity recognition conflicts when multiple people speak at the same time.

[0008] Furthermore, the peripheral device status linkage engine has a built-in scenario mapping table, which associates user identity identifiers, multiple preset scenario tags, and corresponding peripheral device dynamic configuration logic. When the current active user identity identifier output by the context-aware decision module changes, and the current time falls within the activation time period corresponding to the scenario tag, the peripheral device status linkage engine automatically sends a scenario mode switching command to the associated external Bluetooth device.

[0009] Furthermore, the integrated configuration module also includes a scene recognition submodule, which is connected to the microphone array and used to analyze the environmental noise spectrum characteristics and environmental signal-to-noise ratio. When the environmental noise contains vehicle engine characteristics, the audio parameter adaptive engine automatically loads the vehicle audio equalizer preset. When the environmental noise is below the silence threshold and the current voiceprint matching user is an elderly user, the elderly user tag comes from the user attribute tag entered during user registration, or from the user category tag pre-set by the administrator in the user permission hierarchical management module, and automatically triggers the dialogue enhancement and high-frequency compensation algorithm.

[0010] Furthermore, the voiceprint feature extraction and matching calculation of the voiceprint registration and management module are all completed in the embedded neural network processing unit on the Bluetooth speaker. The real-time recognition process does not rely on the cloud network. The system also includes a user permission hierarchical management module. This module controls the open level of the Bluetooth speaker host function based on the user identity identifier output by the context-aware decision module, including access permissions for adult content, maximum volume limit, and system setting modification permissions.

[0011] Furthermore, the expression for the weighted fusion decision made by the context-aware decision module is as follows: ,in, For the normalized voiceprint matching score, This is a normalized value for the signal reception strength indication of the terminal device bound to the user. ,and and Dynamically adjusts based on the current noise level in the environment.

[0012] Furthermore, the peripheral device status linkage engine establishes group communication with external Bluetooth devices through a Bluetooth Low Energy Mesh network. When multiple external Bluetooth devices are involved in the configuration snapshot corresponding to the current active user identity, the engine sends control commands in parallel and sends a configuration completion signal to the integrated configuration module after receiving confirmation frames from all devices.

[0013] Furthermore, the voiceprint registration and management module supports incremental registration and dynamic updates: During system use, when the context-aware decision module identifies the same user three or more times with a fusion score higher than the preset fusion confidence threshold, and the voiceprint similarity between the user's corresponding voice frame and the stored voiceprint digital fingerprint is lower than the template update threshold, the system will use the corresponding voice frame as an incremental training sample to fine-tune and update the user's voiceprint digital fingerprint during the device's idle period.

[0014] Furthermore, the integrated configuration module also includes a learning and prediction submodule. This submodule records each user's manual adjustment behavior at different time periods and under different external device connection states, and generates a personalized configuration prediction model. When the context-aware decision module identifies the user's identity, the learning and prediction submodule actively pushes the predicted configuration parameters to the audio parameter adaptive engine and the peripheral device status linkage engine to achieve advanced configuration.

[0015] Furthermore, the system establishes a non-real-time data synchronization channel with the cloud server that does not participate in the real-time recognition process; the voiceprint registration and management module uploads the de-identified voiceprint feature hash value, anonymized configuration preference information, and local model update parameters to the cloud server. Based on the anonymized configuration preference information and local model update parameters of multiple users, the cloud server performs federated learning optimization on the general audio equalizer preset and distributes the optimized preset template to the comprehensive configuration module for new users to use as an initial configuration reference when registering.

[0016] Compared with the prior art, the beneficial effects of the present invention are: This Bluetooth speaker multi-user personalized configuration management system based on user voiceprint recognition achieves an intelligent experience of one-time registration, seamless switching, and full-dimensional configuration on Bluetooth speakers through deep collaboration of the voiceprint registration and management module, context-aware decision-making module, and comprehensive configuration module. The system can not only accurately identify the current user's identity and automatically switch audio parameters, but also synchronously link and adjust peripheral Bluetooth devices such as lights, air conditioners, and curtains, upgrading the traditional Bluetooth speaker from a single music player to a multi-user smart home hub, significantly improving the personalized experience, ease of operation, and device utilization in shared device scenarios.

[0017] 1. Furthermore, this invention breaks through the limitations of existing technologies that can only switch sound effects or recommend songs. The integrated configuration module has a built-in audio parameter adaptive engine and a peripheral device status linkage engine. When different users are identified, it not only automatically loads the user's preferred equalizer, dynamic range compression threshold, virtual surround sound and other auditory parameters, but also controls multiple external devices in parallel through Bluetooth Low Energy Mesh network. This multi-dimensional linkage configuration of hearing and environment creates an immersive and highly customized usage scenario for users, significantly improving the integrated experience of smart homes.

[0018] 2. Furthermore, the context-aware decision module of the present invention continuously intercepts the ambient audio stream and performs voiceprint separation and matching in real time without relying on the wake word. Combined with the Bluetooth signal reception strength and physical button status, it can complete the identity determination in advance or instantly. The user does not even need to speak or operate. As long as the user is within the Bluetooth coverage range of the speaker, the system can predict the user's identity and prepare the corresponding configuration, truly realizing the seamless integration of human-computer interaction.

[0019] 3. Furthermore, this invention employs an attention-based sound source separation algorithm to decompose mixed sound into independent speech frames, and introduces a dual-threshold competition and sliding time window dynamic fusion mechanism. When the voiceprint matching score is between the high and low thresholds, the system automatically introduces auxiliary weighting factors such as the user terminal device's historical connection frequency and current physical button operation status for secondary judgment. This mechanism effectively avoids identity misjudgment and configuration flicker caused by brief interference, ensuring that the current active user can still be stably and accurately locked in complex acoustic scenarios.

[0020] Furthermore, this invention completes all voiceprint feature extraction, storage, and matching calculations within the embedded neural network processing unit on the Bluetooth speaker itself, without relying on the cloud network. This fundamentally eliminates the risk of leakage of user biometric data. At the same time, the system supports incremental registration and dynamic updates: when a user's voice undergoes natural changes that cause a drop in the matching score, the system automatically uses high-confidence voice frames to fine-tune the voiceprint model locally, without requiring the user to manually re-register. In addition, the learning and prediction submodule records each user's historical adjustment behavior and proactively pushes predictive configurations, achieving a continuous evolution capability that becomes more intuitive and understanding over time. Attached Figure Description

[0021] Figure 1 This is a schematic diagram of the overall hardware architecture of the Bluetooth speaker multi-user personalized configuration management system based on user voiceprint recognition according to the present invention. Figure 2 The flowchart shows the algorithm logic for the context-aware decision module of the present invention to perform multi-person voiceprint separation and user weighted scoring decision. Figure 3This is a timing data flow diagram showing how the audio parameters and the status of the external Bluetooth device are synchronously adjusted when the integrated configuration module of this invention switches user identities. Detailed Implementation

[0022] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0023] This invention provides a multi-user personalized configuration management system for Bluetooth speakers based on user voiceprint recognition. The system can be logically divided into a data acquisition layer, a decision-making layer, and a configuration execution layer. Its core lies in the deep collaboration of the voiceprint registration and management module, the context-aware decision-making module, and the comprehensive configuration module, which realizes an intelligent experience of one-time registration, seamless switching, and full-dimensional configuration on the Bluetooth speaker. The system can not only accurately identify the current user's identity and automatically switch audio parameters, but also synchronously adjust peripheral Bluetooth devices, upgrading the traditional Bluetooth speaker from a single music player to a multi-user smart home hub. It significantly improves the personalized experience, ease of operation, and device utilization in shared device scenarios, and has outstanding substantive features and significant technological progress.

[0024] Example 1: To better understand the above technical solution, the following will provide a detailed description of the technical solution in conjunction with the accompanying drawings and specific implementation methods. (Refer to...) Figure 1 As shown in the figure, this is a schematic diagram of the overall hardware architecture of a Bluetooth speaker multi-user personalized configuration management system based on user voiceprint recognition. The Bluetooth speaker multi-user personalized configuration management system based on user voiceprint recognition includes the following modules: The voiceprint registration and management module is used to collect registration voice samples from at least one user, extract voiceprint features, and generate a voiceprint digital fingerprint bound to the user's unique identity. Specifically, the voiceprint feature extraction and matching calculation of the voiceprint registration and management module are completed in the embedded neural network processing unit on the Bluetooth speaker itself, without relying on the cloud network. Specifically, the module uses the built-in microphone array of the Bluetooth speaker, with at least two microphones, to support beamforming to collect registration voice samples read aloud by the user according to prompts. Each user needs to read 2 to 4 random short sentences of 2 to 3 seconds each, such as number strings or fixed phrases, to cover different pronunciation states. The collected voice signal undergoes echo cancellation, noise suppression, and automatic gain control to obtain a clean registration voice segment.The system also includes a user permission hierarchical management module. This module controls the open level of the Bluetooth speaker's main functions based on the user identity identifier output by the context-aware decision module, including access permissions for adult content, maximum volume limits, and system setting modification permissions. The voiceprint registration and management module supports incremental registration and dynamic updates: specifically, it extracts the Mel-frequency cepstral coefficients (MFCC) of each speech segment (40-dimensional), the Log-Mel spectrogram (64-dimensional), and the d-vector (256-dimensional) based on the ECAPA-TDNN deep neural network to form a multi-feature fusion vector. All feature extraction and embedding vector calculations are completed offline in the Bluetooth speaker's local embedded neural network processing unit (NPU), without relying on any cloud server, ensuring that the user's biometric privacy is not uploaded. The feature vectors of multiple speech samples from the same user are averaged or clustered to generate a unique voiceprint digital fingerprint for that user, which is then bound to a unique identifier (UID) assigned by the system or set by the user and stored in the speaker's encrypted non-volatile memory. During system use, if the context-aware decision module registers a voiceprint with a value higher than the preset fusion confidence threshold three or more times, the system will take action. When the fusion score identifies the same user, and the voiceprint similarity between the user's corresponding voice frame and the stored voiceprint digital fingerprint is lower than the template update threshold, the system uses the corresponding voice frame as an incremental training sample to fine-tune and update the user's voiceprint digital fingerprint during device idle periods. Specifically, the system has a built-in user voiceprint database that supports the registration of at least 10 different users, each with a corresponding record, including: UID, voiceprint digital fingerprint, user-defined tags such as father, mother, permission level, child / adult / administrator, and historical matching confidence logs. The user permission hierarchical management module is a sub-module of the system. Based on the active user UID output in real time by the context-aware decision module, it controls the openness level of the Bluetooth speaker host's functions: Child users: limit the maximum volume to no more than 75dB, block music / podcast sources containing adult content, and prohibit modification of system time and Bluetooth pairing information; Adult users: open the entire volume range and content access, and allow modification of EQ presets; Administrator users: additionally allow adding / deleting other users' voiceprint registrations, restoring factory settings, and modifying permission levels.

[0025] Incremental registration and dynamic update mechanism: The module supports incremental registration: In the existing user database, new voice samples can be added to existing UIDs via voice commands or physical buttons to update the voiceprint fingerprint to adapt to the natural changes in the user's voice with age, cold, and other conditions; The module also has automatic dynamic update capability: During use, if the context-aware decision module identifies a user's UID with high confidence and a matching score >0.85 more than 3 times in a row, but the similarity between the matching score and the stored voiceprint fingerprint is lower than the registration threshold, such as lower than 0.7, the system determines that the user's voice has drifted to an acceptable degree. At this point, the system automatically caches the voice frames from this interaction as incremental training samples. When the device is idle, with no audio playback or user interaction for more than 30 seconds and sufficient battery power, the system fine-tunes the voiceprint model on the local NPU and updates the user's voiceprint digital fingerprint without re-executing the complete registration process. Output interface: Provides the registered user's voiceprint digital fingerprint database and corresponding UID to the context-aware decision module and responds to matching requests in real time. Input interface: Receives high-confidence, low-score events from the context-aware decision module, triggering a dynamic update process; Receives query requests from the user permission hierarchical management module and returns the permission level for the specified UID.

[0026] The role of the voiceprint registration and management module in the overall solution: This module serves as the identity source for the entire system. It pre-converts different users in the physical world into machine-recognizable voiceprint fingerprints in the digital world, providing a unique matching benchmark for the subsequent context-aware decision-making module. Without the registration data from this module, the system cannot distinguish who is speaking, making subsequent personalized configuration impossible. By performing voiceprint feature extraction, storage, and matching calculations entirely on the local NPU without relying on cloud networks, this module fundamentally avoids the risk of user voiceprint data leakage or server misuse during upload. Furthermore, its offline capability allows the system to operate fully in environments without internet access, such as outdoor camping or in-vehicle scenarios, enhancing the solution's versatility and security. Combined with the built-in user permission hierarchical management submodule, this module not only identifies who the user is but also determines what actions are permitted for that user. This enables the Bluetooth speaker to automatically recognize different identities such as children, visitors, and family members. The dynamic adjustment of the openness of the function effectively prevents minors from accessing inappropriate content or misoperating the core settings of the device, improving the family-friendliness and security of the product. Through incremental registration and dynamic update mechanism, this module overcomes the problem of decreased experience caused by the one-time registration and lifelong immutability of traditional voiceprint recognition systems. For example, the recognition rate drops sharply when a user's voice changes after a cold. This mechanism allows the voiceprint fingerprint to be dynamically fine-tuned according to the user's natural voice changes, always maintaining a high recognition accuracy. At the same time, it avoids the tedious operation of frequent manual re-registration, reflecting the system's intelligence and user-unobtrusive maintenance characteristics. The voiceprint digital fingerprint generated by this module is the first key to start the entire personalized configuration chain. After the context-aware decision module completes the matching, the comprehensive configuration module can load the corresponding audio parameters and peripheral device status according to the UID. It can be said that the accuracy, security and adaptability of the voiceprint registration and management module directly determine whether the entire recognition, decision-making and configuration closed loop can operate reliably and smoothly.

[0027] refer to Figure 2As shown, the context-aware decision-making module is connected to the voiceprint registration and management module, the microphone array of the Bluetooth speaker, and the Bluetooth communication unit. Specifically: Input end: connected to the microphone array of the Bluetooth speaker (at least two microphones, supporting beamforming connection) to receive the original ambient mixed audio stream in real time; connected to the voiceprint registration and management module to obtain the voiceprint digital fingerprint database of registered users, UID-feature vector; connected to the Bluetooth communication unit to obtain the device identifiers of all terminal devices currently connected to the speaker via Bluetooth, such as mobile phones and tablets, and their real-time signal reception strength indicators (RSSI); Output end: outputs the currently determined active user identity identifier to the comprehensive configuration module, and can simultaneously... The selected location feeds back matching confidence and high-confidence low-match-score events to the voiceprint registration and management module to trigger incremental updates. This module intercepts mixed ambient sound collected by the microphone array in real time, performing sound source separation to extract valid user speech frames. Specifically, the module continuously intercepts mixed sound collected by the microphone array when the Bluetooth speaker is in standby or playback mode, without relying on wake-word triggers, achieving zero-latency response. A U-Net architecture based on an attention mechanism is used for real-time sound source separation: multi-channel time-frequency domain features are input into the U-Net encoder, and the independent mask of each sound source is extracted through a cross-channel attention module, separating the independent speech frames of each speaker from the mixed spectrum. This algorithm can run on the local NPU with a latency of less than 20ms and supports the simultaneous separation of up to three speakers.

[0028] The voiceprint features of the speech frames are extracted and matched with the voiceprint digital fingerprints. Simultaneously, the signal reception strength indication of at least one terminal device currently connected to the Bluetooth speaker is obtained. Specifically, for each isolated independent speech frame, a 40-dimensional MFCC + 64-dimensional Log-Mel spectrogram is extracted and input into a lightweight ECAPA-TDNN network to generate a 256-dimensional voiceprint embedding vector. The embedding vector is then used to calculate the cosine similarity between the cosine similarity and the voiceprint digital fingerprints of all registered users provided by the voiceprint registration and management module to obtain the original matching score for each candidate user. The scores were normalized to obtain The formula is: ,in, The set of scores for all candidate users.

[0029] A weighted fusion decision is made based on voiceprint matching score and signal reception strength index (RSSI) to output the identity of the currently active user. Specifically, the terminal device bound to each candidate user is obtained simultaneously, and the associated RSSI value is recorded via Bluetooth pairing. If a user is not currently connected to any device, then... If there are multiple devices, take the largest RSSI and normalize the RSSI: The expression for weighted fusion decision-making by the context-aware decision module is: ,in, For the normalized voiceprint matching score, This is a normalized value for the signal reception strength indication of the terminal device bound to the user. ,and and The fusion score is dynamically adjusted based on the noise level in the current environment. Equal to the normalized value of voiceprint matching score With weight The product of the two values, plus the normalized value of the received signal strength indicator. With weight The product of these factors, along with the context-aware decision-making module employing an attention-based sound source separation algorithm, separates the independent speech frames of each speaker from the mixed sound. Dynamic weight adjustment is also implemented: the module incorporates a noise level estimator to calculate the environmental signal-to-noise ratio (SNR) in real time. When SNR < 5dB, in a high-noise environment, the reliability of voiceprint matching decreases. Automatically reduced to 0.4. Increase to 0.6; when SNR>15dB, in a quiet environment, Increase to 0.8, Reduced to 0.2; default value used under moderate noise. =0.65, =0.35.

[0030] The system utilizes a dual-threshold competition and sliding time window dynamic fusion mechanism to determine active users. When the average fusion score of a candidate user within the sliding time window is higher than the first threshold, the candidate user is directly locked as the current active user. When the average fusion score is greater than the second threshold and less than or equal to the first threshold, the historical connection frequency of the candidate user's corresponding terminal device and the current physical button operation state of the Bluetooth speaker are introduced as auxiliary weighting factors to resolve identity recognition conflicts when multiple people speak simultaneously. Specifically, two confidence thresholds are set: a high threshold and a low threshold. low threshold Time window: Length of the sliding window In seconds, with a step of 0.3 seconds, calculate the average fusion score for each candidate user within each window. .

[0031] Decision logic: If a user exists within the window... ,That If the condition persists for more than 0.5 seconds, the user is directly locked as the currently active user, and the following output is displayed: If all users are below However, there is a certain user of lie in Within the interval, an auxiliary weighting factor is introduced for secondary determination: historical connection frequency. The number of times the user's terminal device established a Bluetooth connection with the speaker in the past 5 minutes, normalized to [0,1]; physical button operation status. The physical button operation status For candidate users Confirmed, when a physical button on the Bluetooth speaker is detected to have been triggered within the last second, and the candidate user... When the rate of change of RSSI of the bound terminal device is the highest among all candidate users, Assign a value of 1; otherwise, Assigning a value of 0: If a physical button on the speaker, such as play / pause or volume +, is pressed within the last second, a higher weight is given to the user currently approaching the device. This is typically determined by the RSSI rate of change, and is used as an auxiliary scoring formula. ,in ,default , The final overall score is The user with the highest score is selected as the active user. The default value is 0.7; if all users... If no active user or unknown user is found, an empty identifier is output, and the system retains its current configuration.

[0032] Collaboration with other modules: Feedback to the voiceprint registration and management module: When the same user is identified in three consecutive windows but the matching score is lower than [a certain value], [the following occurs]. However, if the value is higher than 0.5, an incremental update request is triggered, informing the module that the user's voiceprint needs fine-tuning; the integrated configuration module is then notified that once an active user is locked, an interrupt signal is immediately sent via the internal bus, along with... And the confidence level, high / medium, requires the comprehensive configuration module to perform configuration switching.

[0033] The role of the context-aware decision-making module in the overall solution: The context-aware decision-making module is the brain of the entire system. It does not simply use voiceprint matching results as the sole criterion, but integrates auditory information, voiceprint, spatial distance information, Bluetooth RSSI, historical behavior, connection frequency, physical interaction, and button operation to achieve multimodal and multi-dimensional comprehensive reasoning. This design solves the problem of single biometric recognition easily failing in noisy environments, significantly improving the robustness and accuracy of identity determination. Traditional smart speakers require users to say a wake word to start recognition. The module in this invention continuously intercepts and analyzes ambient sound, without relying on any wake word. Once a valid voice frame is captured, the system can complete identity determination in milliseconds. The user is not even aware that the speaker is thinking, but personalized sound effects and environmental configurations have already been quietly set up. This seamless switching greatly improves the smoothness of the user experience. In home or office scenarios, multiple people often speak at the same time or multiple people are close to the speaker. This module decomposes mixed sound into independent voice frames through sound source separation, and then combines dual thresholds. The sliding time window and auxiliary weighting factor effectively avoid configuration flickering caused by brief interference, such as frequent switching between user A and user B. For example, when user A speaks from a distance and user B holds their phone close to the speaker and presses the play button, the module can correctly lock onto user B by combining the strong RSSI signal and physical button operation. By dynamically adjusting the voiceprint weight and RSSI weight, the module can adapt to different acoustic environments: in noisy kitchens or street scenes, it relies more on distance information RSSI rather than easily contaminated voiceprint matching; in quiet bedrooms, it returns to high-precision recognition dominated by voiceprint. This dynamic adjustment mechanism enables the system to maintain stable decision-making performance in changing environments. All personalized behaviors of the integrated configuration module, such as sound effect switching, lighting adjustment, and access control, depend on the active user UID output by the module. If the identity determination is incorrect or the delay is too large, the entire personalized configuration system will lose its meaning. Therefore, the accuracy, real-time performance, and anti-interference ability of the context-aware decision module directly determine the overall intelligence level and user satisfaction of this invention.

[0034] refer to Figure 3As shown, the integrated configuration module is connected to the context-aware decision module, the audio digital signal processor of the Bluetooth speaker, and the Bluetooth communication unit. This module internally includes an audio parameter adaptive engine and a peripheral device status linkage engine. Specifically, the peripheral device status linkage engine has a built-in scene mapping table that associates user identity identifiers, multiple preset scene tags, and corresponding peripheral device dynamic configuration logic. When the currently active user identity identifier output by the context-aware decision module changes, and the current time falls within the activation time period corresponding to the scene tag, the peripheral device status linkage engine automatically sends a scene mode switching command to the associated external Bluetooth device. Further, the module connection relationships are as follows: Bottom: Input end: Connects to the context-aware decision module to receive the identity identifier (UID) and confidence level of the currently active user; indirectly connects to the microphone array through the scene recognition submodule to obtain environmental audio features; connects to the audio digital signal processor (DSP) of the Bluetooth speaker for sending audio parameters; connects to the Bluetooth communication unit for sending control commands to external Bluetooth devices; Internal components: Includes an audio parameter adaptive engine, a peripheral device status linkage engine, a scene recognition submodule, and a learning and prediction submodule; Output end: Sends audio parameter configuration commands to the DSP, sends status switching commands to peripheral devices through the Bluetooth communication unit, and provides feedback on the configuration completion status to the context-aware decision module.

[0035] The audio parameter adaptive engine pre-stores audio equalizer parameters, dynamic range compression thresholds, and virtual surround sound parameters corresponding to different user identities. Based on the received currently active user identity, it sends a switching command to the audio digital signal processor. Specifically: Pre-stored content: A complete set of audio parameter configurations is independently stored for each registered user UID, including but not limited to: Multi-band equalizer (EQ) parameters: such as the center frequency, gain, and Q value of a 10-band graphic equalizer or parametric equalizer; Dynamic range compression (DRC) thresholds: compressor start threshold, compression ratio, attack time, and release time; Virtual surround sound parameters... Parameters: sound field width, reverberation intensity, spatialization algorithm enable flag; loudness normalization target value: used for volume consistency adjustment of different input sources; dialogue enhancement and high-frequency compensation: gain enhancement parameters for speech segments, stored separately, for triggering by the scene recognition submodule; switching mechanism: upon receiving the UID sent by the context-aware decision module, the engine immediately reads the corresponding parameters from the pre-stored configuration and sends a configuration switching command to the audio DSP via the I2C or SPI interface. If the DSP supports real-time parameter smooth transition, such as 100ms linear interpolation, the engine will add a gradual duration parameter to avoid popping or abruptness during switching.

[0036] The peripheral device status linkage engine pre-stores runtime status snapshots of at least one external Bluetooth device corresponding to different user identities. Based on the received currently active user identity, it sends status switching commands to the external Bluetooth devices via the Bluetooth communication unit. Specifically, the peripheral device status linkage engine establishes group communication with the external Bluetooth devices through a Bluetooth Low Energy Mesh network. When multiple external Bluetooth devices are involved in the configuration snapshot corresponding to the currently active user identity, the engine sends control commands in parallel and sends a configuration completion signal to the integrated configuration module after receiving confirmation frames from all devices. Specifically: Pre-stored content: Independently stores runtime status snapshots of external Bluetooth devices for each user. The snapshots include: device type (smart bulb, air conditioner, curtains, aroma diffuser, air purifier, etc.); device Bluetooth MAC address and binding key; desired status parameters, such as: bulb color temperature 2700K / 5000K, brightness 80%; air conditioner mode cooling / heating, set temperature 24℃; curtain opening / closing degree 50%, etc.; scene mapping table: A built-in multi-dimensional mapping table maps UID, scene tags, and time periods to a set of dynamic configuration logic. Scene tags may include clear... For scenarios such as morning wake-up, late-night video playback, commuting, and outdoor camping, users can pre-set device states for different scenarios via the app or voice. When the UID output by the context-aware decision module changes, and the current system time falls within the activation time period corresponding to a certain scenario tag (e.g., 22:00–06:00 corresponds to late-night video playback), the engine automatically finds the user's pre-set device state combination for that scenario and generates control commands. Communication method: The peripheral device status linkage engine establishes group communication with multiple external Bluetooth devices via Bluetooth Low Energy (BLE) and a Mesh network. The specific process is as follows: The engine generates a set of parallel control commands based on snapshots. Each command includes the target device address, operation code, and parameter value. Utilizing the broadcast or multi-connection capability of the Bluetooth communication unit, all commands are sent in parallel instead of waiting serially. A timeout timer is set, with a default of 2 seconds, to wait for an ACK frame from each device. When all expected devices return ACKs or the timeout period ends, the engine sends a configuration completion signal to the main control logic of the integrated configuration module, indicating that the external environment is ready. If a device fails to respond three times consecutively, the engine marks the device as offline and may optionally provide a voice prompt to the user via a speaker.

[0037] The integrated configuration module also includes a scene recognition submodule, which connects to the microphone array and is used to analyze environmental noise spectrum features and environmental signal-to-noise ratio (SNR). When the environmental noise contains vehicle engine features, the audio parameter adaptive engine automatically loads the vehicle audio equalizer preset. When the environmental noise is below the silence threshold and the current voiceprint matching user is an elderly user, the elderly user label comes from the user attribute label entered during user registration or from the user category label pre-set by the administrator in the user permission hierarchical management module, automatically triggering dialogue enhancement and high-frequency compensation algorithms. Specifically, the submodule is directly connected to the microphone array, receives environmental audio streams in real time, shares the same audio front end with the context-aware decision module, but processes it independently. This submodule runs a lightweight environmental sound classification model based on a convolutional neural network and local NPU inference, mainly analyzing: Environmental noise spectrum features: extracting frequency domain energy distribution and identifying specific acoustic patterns, such as engine roar, air conditioner hum, baby crying, rain sounds, etc.; Environmental signal-to-noise ratio (SNR): estimating the difference between speech and background noise in the current environment. The scene noise energy ratio; trigger action: when the environmental noise contains vehicle engine characteristics, such as low-frequency roar of a car engine and tire friction sound, the vehicle engine characteristics are composed of low-frequency continuous roar, periodic speed change sound pattern and broadband noise characteristics formed by tire friction with the road surface, and the speaker is placed in vehicle mode and can be detected via Bluetooth connection to the vehicle power supply, the scene recognition submodule automatically notifies the audio parameter adaptive engine to load the vehicle audio equalizer preset, which is usually to enhance bass and suppress road noise frequency bands; when the environmental noise is detected to be below the silence threshold, such as SNR>20dB and background sound pressure level<30dB, and the user tag corresponding to the active user identity identifier output by the current context perception decision module is an elderly user, which is preset by the user permission hierarchical management module, the dialogue enhancement and high frequency compensation algorithm is automatically triggered to improve the gain of the 2kHz–8kHz frequency band by 3–6dB, helping the elderly to hear voice content, such as podcasts and calls, more clearly; other preset scenarios, such as automatically turning on wind noise suppression in outdoor windy environments, can be expanded as needed.

[0038] The comprehensive configuration module also includes a learning and prediction submodule. This submodule records each user's manual adjustment behavior at different times and under different external device connection states, and generates a personalized configuration prediction model. When the context-aware decision module identifies the user, the learning and prediction submodule proactively pushes the predicted configuration parameters to the audio parameter adaptive engine and the peripheral device status linkage engine to achieve advanced configuration. Specifically: Data recording: Continuously tracks each user's interaction behavior with the speaker, recording the following dimensions: different time periods, with a granularity of half an hour, user's preferred sound effect configuration, such as liking to turn on virtual surround sound at night and turn it off during the day; configuration adjustments under different external device connection states, such as automatically turning off the loudness compensation of the speaker's external EQ when headphones are connected; user manual intervention records: when the user manually modifies EQ, volume, or peripheral device status via physical buttons or the APP. After the device status is determined, the submodule binds and stores the operation with the current user, time, and scene. The prediction model utilizes a lightweight temporal prediction model, such as LSTM or a simple Markov chain, running on the local NPU to train each user's personalized configuration. The model input includes the current user's UID, current time, day of the week + hour, current Bluetooth device connection list, and current ambient noise category. The output is the predicted combination of configuration parameters, including EQ, DRC, and scene modes. For proactive configuration, once the context-aware decision module identifies the user, this submodule proactively pushes the predicted configuration parameters to the audio parameter adaptive engine and the peripheral device status linkage engine, without waiting for user commands or entering a preset time period. If the user subsequently manually adjusts the configuration, the system records the adjusted parameters as new training samples and periodically updates the model incrementally. This mechanism achieves a proactive experience where the configuration is ready before the user speaks.

[0039] Internal Collaboration Process: When the integrated configuration module receives a new UID from the context-aware decision module, the internal processing flow is as follows: Parallel Trigger: The audio parameter adaptive engine begins reading the user's audio configuration. Simultaneously, the learning and prediction submodule generates a predicted configuration based on the current context; the scene recognition submodule analyzes environmental noise in parallel; Fusion Decision: The main control logic performs a weighted fusion of the pre-stored configuration and the predicted configuration. When the prediction confidence is high, the predicted value is adopted, and specific parameters are overridden with reference to the scene recognition results, such as forcibly overriding EQ in vehicle mode; Execution: The audio parameters are sent to the DSP via the engine; the peripheral device status linkage engine generates instructions based on the scene mapping table and user snapshot, and distributes them in parallel via BLEMesh; Feedback and Recording: After configuration, the learning and prediction submodule records the final parameters of this switch for subsequent model training.

[0040] The role of the integrated configuration module in the overall solution: The integrated configuration module is the limbs of the system, directly translating user identity into perceptible auditory and atmospheric changes: from subtle differences in EQ to the warm and cool transitions of light color temperature, from the opening and closing of curtains to the activation and deactivation of aromatherapy. Without this module, voiceprint recognition and decision-making would be mere abstractions. It determines whether users can truly experience the intelligent feeling of the speaker recognizing them and adapting accordingly. This module not only manages the speaker's own sound parameters but also coordinates all Bluetooth-enabled smart devices in the home through a peripheral device status linkage engine. It upgrades the Bluetooth speaker from an isolated player to a lightweight central controller for the home IoT, eliminating the need for expensive smart gateways or cloud platforms and enabling cross-device scene linkage solely through local BLEMesh. Through the scene recognition submodule, the system can automatically adjust the audio based on the current acoustic environment: in a car / quiet indoor / noisy outdoor environment. The ability to adjust parameters, and even provide targeted compensation for specific groups such as elderly users, makes personalized configuration no longer static and rigid, but can be optimized in real time as the environment changes, significantly improving the listening experience in different usage scenarios. The learning and prediction submodule gives the system memory and predictive capabilities. By recording the user's historical adjustment behavior, it proactively predicts the configuration that the user may want, achieving proactive configuration before the action occurs. This makes the user almost unaware of the switching delay. With long-term use, the system will understand you better and better, forming a tacit understanding. The parallel sending and ACK confirmation mechanism of the peripheral device status linkage engine ensures that multiple peripheral devices can be set to the desired state simultaneously and reliably. This avoids the long waiting time caused by serial configuration or the inconsistency caused by device offline. At the same time, the feedback of the configuration completion signal allows the system to confirm itself or promptly remind the user when problems occur.

[0041] The voiceprint registration and management module, context-aware decision-making module, and integrated configuration module work together to enable the same Bluetooth speaker to switch corresponding audio parameters based on the currently active user's identity and send a running status switching command to the corresponding external Bluetooth device. The Bluetooth speaker multi-user personalized configuration management system based on user voiceprint recognition establishes a non-real-time data synchronization channel with the cloud server. Specifically, this includes: data desensitization and uploading: The voiceprint registration and management module uses an irreversible hash function, such as SHA-256, to generate a fixed-length desensitized feature hash value for each user's locally stored voiceprint digital fingerprint, removing all timestamps and device identifiers that can be directly associated with the user's identity. In addition to name tags, only statistical dimensions such as age range and gender tags are retained, which are voluntarily provided by users during registration. When the device is idle and connected to Wi-Fi, the system encrypts and uploads the anonymized hash values ​​and corresponding user anonymization configuration preferences, including EQ preset indexes and frequently used peripheral device scenario tags, to the cloud server in batches. The upload frequency can be configured to be once every 7 days or triggered once every 50 cumulative recognition events. Federated learning optimization: The cloud server aggregates anonymized data from a large number of different user devices and uses a federated learning framework for group preference analysis. Specifically, the server does not directly obtain any original voiceprints or specific configurations, but... The server receives model gradient updates or anonymized statistical features uploaded from various devices, such as the average preference values ​​for low-frequency gain among users of different age groups. It then trains a general-purpose audio equalizer preset model in the cloud. This model takes user group labels, such as youth – pop music preference, and elderly – conversation enhancement preference, as input, and outputs a set of recommended EQ parameters and DRC thresholds. During training, the local data of each participating device remains on the device; only encrypted model updates are uploaded. The server aggregates these updates and distributes the optimized global preset template. The preset template distribution and initial configuration: The cloud server distributes the optimized preset template, which includes multiple initial EQ curves for different user types, virtual... The proposed surround sound switch uses compressed signatures and is pushed to the Bluetooth speaker's integrated configuration module in a non-real-time manner. When a new user registers their voiceprint on the speaker, the integrated configuration module first determines the user's age or preference tags, which can be obtained through a short voice Q&A. If a category matches, the preset configuration for that category is automatically used as the new user's initial audio parameters and displayed on the mobile app for fine-tuning. If the user does not provide any tags, the system's default equalizer template is used as the starting point. Privacy and security mechanisms: The cloud server does not store any raw data that can identify specific devices or users, and all communication uses TLS1.The system employs three encryption methods, and users can disable cloud synchronization locally at any time. In this case, the system relies solely on the device's built-in general presets for new user configuration initialization, without affecting local core recognition and configuration capabilities. The voiceprint registration and management module uploads the anonymized voiceprint feature hash value to the cloud server. The cloud server then performs federated learning optimization on the general audio equalizer preset based on the group preferences of multiple users and distributes the optimized preset template to the integrated configuration module for initial configuration reference during new user registration.

[0042] The above describes the functions of the cloud server collaboration module, or cloud optimization and synchronization subsystem, in the system. It plays a crucial role in the overall solution as follows: A single Bluetooth speaker can only learn the preferences of a limited number of household users. By aggregating a large amount of anonymized data through the cloud server and performing federated learning, the system can uncover statistical patterns at the group level. For example, users aged 30-40 generally prefer bass enhancement and a wide soundstage, while users over 60 prefer prominent mid-range frequencies and stable loudness. This group knowledge can be fed back to each device, allowing new users to receive a more reasonable initial configuration upon registration, rather than manually adjusting from scratch, significantly improving the out-of-the-box experience. Through anonymized hashing and federated learning, the user's original voiceprint characteristics and specific configuration content remain on the local device. The cloud only accesses statistical gradients or hash fingerprints that cannot be reversed, allowing the system to continuously optimize the recommendation model within legal and compliant boundaries, while avoiding privacy risks. Real-time decisions regarding voiceprint recognition and personalized configuration are made entirely locally, with cloud synchronization only occurring when the device is idle and connected to the internet, without interfering with normal use. This ensures the system can still function fully without an internet connection, downgrading to using only local universal presets, while gradually gaining the benefits of group optimization after the network is restored, balancing offline reliability with online evolution capabilities. Manufacturers can periodically release new preset templates via cloud servers, such as for newly discovered music genres or accessibility needs, without forcing users to update firmware. This mechanism gives the product a self-evolving characteristic; as the user base expands, the accuracy of configuration recommendations will continuously improve, forming a technological barrier and user stickiness. For ordinary households, users may not know how to adjust EQ or set up lighting linkages. The group optimization templates provided by the cloud can automatically generate an initial configuration that is likely to suit the preferences of new users, reducing the tediousness of manual settings and improving the ease of use and satisfaction of the product.

[0043] The contents not described in detail in this specification are existing technologies known to those skilled in the art.

[0044] Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions described in the foregoing embodiments or make equivalent substitutions for some of the technical features. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.

Claims

1. A Bluetooth speaker multi-user personalized configuration management system based on user voiceprint recognition, characterized in that, include: The voiceprint registration and management module is used to collect registration voice samples from at least one user, extract voiceprint features, and generate a voiceprint digital fingerprint that is bound to the user's unique identity. The context-aware decision-making module is connected to the voiceprint registration and management module, the microphone array of the Bluetooth speaker, and the Bluetooth communication unit. This module intercepts the mixed ambient sound collected by the microphone array in real time, performs sound source separation to extract valid user voice frames, extracts the voiceprint features of the voice frames and matches them with the voiceprint digital fingerprint, and at the same time obtains the signal reception strength indication of at least one terminal device that has established a Bluetooth connection with the Bluetooth speaker. Based on the voiceprint matching score and the signal reception strength indication, a weighted fusion decision is made to output the identity identifier of the currently active user. The integrated configuration module is connected to the context-aware decision module, the audio digital signal processor of the Bluetooth speaker, and the Bluetooth communication unit. The module contains an audio parameter adaptive engine and a peripheral device status linkage engine. The audio parameter adaptive engine pre-stores audio equalizer parameters, dynamic range compression thresholds, and virtual surround sound parameters corresponding to different user identities, and sends a switching command to the audio digital signal processor based on the received current active user identity. The peripheral device status linkage engine pre-stores a snapshot of the operating status of at least one external Bluetooth device corresponding to different user identities, and sends a status switching command to the external Bluetooth device through the Bluetooth communication unit based on the received current active user identity. Among them, the voiceprint registration and management module, the context-aware decision-making module, and the comprehensive configuration module work together to enable the same Bluetooth speaker to switch the corresponding audio parameters according to the identity of the currently active user and send the running status switching command to the corresponding external Bluetooth device.

2. The Bluetooth speaker multi-user personalized configuration management system based on user voiceprint recognition according to claim 1, characterized in that: The context-aware decision-making module uses an attention-based sound source separation algorithm to separate the independent speech frames of each speaker from the mixed sound, and uses a dual-threshold competition and sliding time window dynamic fusion mechanism to determine active users. When the average fusion score of a candidate user within the sliding time window is higher than the first threshold, the candidate user is directly locked as the current active user. When the average fusion score is greater than the second threshold and less than or equal to the first threshold, the historical connection frequency of the candidate user's corresponding terminal device and the current physical button operation status of the Bluetooth speaker are introduced as auxiliary weighting factors to resolve identity recognition conflicts when multiple people speak at the same time.

3. The Bluetooth speaker multi-user personalized configuration management system based on user voiceprint recognition according to claim 1, characterized in that: The peripheral device status linkage engine has a built-in scenario mapping table, which associates user identity identifiers, multiple preset scenario tags, and corresponding peripheral device dynamic configuration logic. When the current active user identity identifier output by the context awareness decision module changes and the current time falls within the activation time period corresponding to the scenario tag, the peripheral device status linkage engine automatically sends a scenario mode switching command to the associated external Bluetooth device.

4. The Bluetooth speaker multi-user personalized configuration management system based on user voiceprint recognition according to claim 1, characterized in that: The integrated configuration module also includes a scene recognition submodule, which is connected to the microphone array and used to analyze the environmental noise spectrum characteristics and environmental signal-to-noise ratio. When the environmental noise contains vehicle engine characteristics, the audio parameter adaptive engine automatically loads the vehicle audio equalizer preset. When the environmental noise is below the silence threshold and the current voiceprint matching user is an elderly user, the elderly user tag comes from the user attribute tag entered during user registration, or from the user category tag pre-set by the administrator in the user permission hierarchical management module, and automatically triggers the dialogue enhancement and high-frequency compensation algorithm.

5. The Bluetooth speaker multi-user personalized configuration management system based on user voiceprint recognition according to claim 1, characterized in that: The voiceprint feature extraction and matching calculation of the voiceprint registration and management module are completed in the embedded neural network processing unit on the Bluetooth speaker. The real-time recognition process does not rely on the cloud. The system also includes a user permission hierarchical management module. This module controls the open level of the Bluetooth speaker host function based on the user identity identifier output by the context-aware decision module, including access permissions for adult content, maximum volume limit, and system setting modification permissions.

6. The Bluetooth speaker multi-user personalized configuration management system based on user voiceprint recognition according to claim 1, characterized in that: The expression for weighted fusion decision-making by the context-aware decision module is: ,in, For the normalized voiceprint matching score, This is a normalized value for the signal reception strength indication of the terminal device bound to the user. ,and and Dynamically adjusts based on the current noise level in the environment.

7. The Bluetooth speaker multi-user personalized configuration management system based on user voiceprint recognition according to claim 1, characterized in that: The peripheral device status linkage engine establishes group communication with external Bluetooth devices through a Bluetooth Low Energy Mesh network. When multiple external Bluetooth devices are involved in the configuration snapshot corresponding to the current active user identity, the engine sends control commands in parallel and sends a configuration completion signal to the integrated configuration module after receiving confirmation frames from all devices.

8. The Bluetooth speaker multi-user personalized configuration management system based on user voiceprint recognition according to claim 1, characterized in that: The voiceprint registration and management module supports incremental registration and dynamic updates: During system use, when the context-aware decision module identifies the same user three or more times with a fusion score higher than the preset fusion confidence threshold, and the voiceprint similarity between the user's corresponding voice frame and the stored voiceprint digital fingerprint is lower than the template update threshold, the system will use the corresponding voice frame as an incremental training sample to fine-tune and update the user's voiceprint digital fingerprint during the device's idle period.

9. The Bluetooth speaker multi-user personalized configuration management system based on user voiceprint recognition according to claim 1, characterized in that: The comprehensive configuration module also includes a learning and prediction submodule. This submodule records each user's manual adjustment behavior at different time periods and under different external device connection states, and generates a personalized configuration prediction model. When the context-aware decision module identifies the user's identity, the learning and prediction submodule actively pushes the predicted configuration parameters to the audio parameter adaptive engine and the peripheral device status linkage engine to achieve advanced configuration.

10. The Bluetooth speaker multi-user personalized configuration management system based on user voiceprint recognition according to claim 1, characterized in that: The system establishes a non-real-time data synchronization channel with the cloud server that does not participate in the real-time identification process; The voiceprint registration and management module uploads the anonymized voiceprint feature hash value, anonymized configuration preference information, and local model update parameters to the cloud server. Based on the anonymized configuration preference information and local model update parameters of multiple users, the cloud server performs federated learning optimization on the general audio equalizer preset and distributes the optimized preset template to the comprehensive configuration module for new users to use as an initial configuration reference when registering.