Method, system and device for determining volume of Bluetooth headset and Bluetooth headset
By acquiring ambient audio data and motion sensor data, combining it with the audio stream played by Bluetooth headsets to generate content type vectors, and using machine learning algorithms and cluster analysis, the problem of Bluetooth headset volume adjustment relying on manual settings is solved, and intelligent, personalized and dynamic volume adjustment is achieved, thereby improving the user experience.
Patent Information
- Application Number
- CN202510665320.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-22
- Publication Date
- 2025-09-19
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Bluetooth headset volume adjustment mainly relies on manual settings by the user and cannot intelligently determine the recommended volume, resulting in a poor user experience.
By obtaining the ambient audio data and motion sensor data of the environment in which the Bluetooth headset is located, and combining it with the audio stream played by the Bluetooth headset to generate a content type vector, the recommended volume is automatically determined using machine learning algorithms and cluster analysis.
It realizes intelligent, personalized and dynamic adjustment of the volume of Bluetooth headsets to meet the needs of diverse scenarios and improve user experience and dynamic adaptability.
Smart Images

Figure CN120676281A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of volume of Bluetooth headsets, and more specifically, to a method, system, device, and Bluetooth headset for determining the volume of a Bluetooth headset. Background Art
[0002] With the rapid development of smart audio devices, Bluetooth headsets have become an indispensable part of people's daily lives. As an essential component of the user experience, volume adjustment directly impacts user satisfaction and comfort with the device. However, in practical applications, Bluetooth headset volume adjustment still faces many challenges.
[0003] First, traditional volume adjustment methods rely primarily on manual settings or fixed preset values. While simple and intuitive, this approach is inflexible in dynamic and ever-changing real-world usage scenarios. For example, users frequently need to manually adjust the volume to suit different environments or content requirements. Fixed volume presets fail to meet individual needs and may result in the volume being too high or too low. Second, as users' expectations for smart devices continue to rise, volume adjustment is also gradually moving towards intelligence. An ideal volume adjustment method should be able to automatically sense the user's usage scenario and dynamically adjust the volume based on the actual situation. Third, future Bluetooth headset volume adjustment methods should place greater emphasis on intelligence and adaptability. By introducing advanced signal processing technologies and machine learning algorithms, more precise volume adjustment can be achieved, thereby improving the user experience.
[0004] It can be seen that the development of Bluetooth headset volume control technology is at a critical stage of transitioning from manual adjustment to intelligent dynamic adjustment. Future research will focus more on how to improve the intelligence level of volume control through technological innovation to meet users' growing personalized needs.
[0005] Regarding the related technology, the volume adjustment method of Bluetooth headsets mainly relies on manual settings by users and cannot intelligently determine the recommended volume, resulting in a low user experience. No effective solution has been proposed yet.
[0006] Therefore, it is necessary to improve the related technology to overcome the above-mentioned defects in the related technology. Summary of the Invention
[0007] The embodiments of the present application provide a method, system, device, and Bluetooth headset for determining the volume of a Bluetooth headset, so as to at least solve the problem that the volume adjustment method of the Bluetooth headset mainly relies on manual settings by the user and cannot intelligently determine the recommended volume, resulting in a low user experience.
[0008] According to one aspect of an embodiment of the present application, a method for determining the volume of a Bluetooth headset is provided, comprising: obtaining ambient audio data of an environment in which the Bluetooth headset is located, and determining an environmental state based on the ambient audio data, wherein the environmental state is used to reflect the noise level of the environment in which the Bluetooth headset is located; obtaining motion sensor data from a target device, and determining a user state of a user using the Bluetooth headset based on the motion sensor data, wherein the target device is a device that establishes a communication connection with the Bluetooth headset; generating a content type vector based on an audio stream played by the Bluetooth headset, wherein the content type vector is used to characterize the speech rate characteristics, pitch characteristics, and frequency distribution characteristics of the audio stream; and determining a recommended volume for the Bluetooth headset based on the environmental state, the user state, and the content type vector.
[0009] In an exemplary embodiment, determining the environmental state based on the environmental audio data includes: digitizing the environmental audio data according to a preset sampling rate to obtain an original audio sequence, and analyzing the original audio sequence using a fast Fourier transform to obtain frequency distribution data, wherein the frequency distribution data includes amplitude information and phase information in different frequency ranges; grouping the frequency components in the frequency distribution data using a two-means clustering algorithm to obtain noise frequency component clusters and signal frequency component clusters; constructing an environmental feature vector based on the noise frequency component cluster, wherein the environmental feature vector includes: the noise decibel value corresponding to the noise frequency component cluster, the frequency range corresponding to the noise frequency component cluster, and the noise proportion; the noise proportion is equal to the ratio of the number of frequency components in the noise frequency component cluster to the number of frequency components in the frequency distribution data; and determining the environmental state based on the environmental feature vector through a preset noise classification model.
[0010] In an exemplary embodiment, determining the user state of a user using the Bluetooth headset based on the motion sensor data includes: parsing the motion sensor data through the Bluetooth communication protocol between the Bluetooth headset and the target device to obtain a target data set, wherein the target data set includes the three-axis acceleration of the target device with different timestamps and the heart rate of the user; processing the target data set using a time series analysis algorithm and a Fourier transform algorithm to obtain a target feature set, wherein the target feature set includes a time period identifier, and acceleration features and cadence features within the time period corresponding to the time period identifier; in the event that heart rate fluctuations exist within the time period corresponding to the time period identifier, determining the heart rate fluctuation rate within the time period corresponding to the time period identifier, and generating a user state vector based on the heart rate fluctuation rate and the target feature set, wherein the user state vector includes the time period identifier, the acceleration features, cadence features, and heart rate fluctuation rate within the time period corresponding to the time period identifier; and determining the user state based on the user state vector using a preset state classification model.
[0011] In an exemplary embodiment, a content type vector is generated based on the audio stream played by the Bluetooth headset, including: using an audio decoder to parse the audio stream to obtain a metadata set, wherein the metadata set includes a timestamp, encoding format, and bit rate corresponding to the audio stream; using a fast Fourier transform algorithm, based on the metadata set, the audio stream is analyzed to obtain a reference feature set, wherein the reference feature set includes a speech rate feature and a pitch feature corresponding to the audio stream; determining the content type corresponding to the audio stream according to the reference feature set; generating a content type vector according to the reference feature set, a frequency distribution feature corresponding to the audio stream, and a first weight, a second weight, and a third weight corresponding to the content type, wherein the content type vector includes evaluation values of the audio stream on the speech rate feature, the pitch feature, and the frequency distribution feature, respectively, the first weight being the weight corresponding to the speech rate feature, the second weight being the weight corresponding to the pitch feature, and the third weight being the weight corresponding to the frequency distribution feature.
[0012] In an exemplary embodiment, before determining the recommended volume of the Bluetooth headset according to the environmental state, the user state and the content type vector, the method further includes: obtaining a plurality of historical environmental feature vectors, historical user state vectors and historical content type vectors, wherein the historical environmental feature vector is an environmental feature vector obtained after processing historical environmental audio data, and the historical user state vector is a user state vector obtained after processing historical motion sensor data; normalizing the plurality of historical environmental feature vectors, historical user state vectors and historical content type vectors, and clustering the plurality of vectors obtained after normalization using a K-means clustering algorithm to obtain K clusters, wherein the K clusters correspond one-to-one to K scenes; generating a scene transition probability matrix based on the K clusters through a Markov chain method, wherein the scene transition probability matrix includes any of the K scenes. The method comprises the following steps: first, determining the recommended volume of the Bluetooth headset according to the environmental state, the user state, the content type vector and the scene similarity matrix; second, determining the recommended volume of the Bluetooth headset according to the environmental state, the user state, the content type vector and the scene similarity matrix; third, determining the recommended volume of the Bluetooth headset according to the environmental state, the user state, the content type vector and the scene similarity matrix; fourth, determining the recommended volume of the Bluetooth headset according to the environmental state, the user state, the content type vector and the scene similarity matrix; fifth ... sixth, determining the recommended volume of the Bluetooth headset according to the environmental state, the user state, the content type vector and the scene similarity matrix; sixth, determining the recommended volume of the Bluetooth headset according to the environmental state, the user state, the content type vector and the scene similarity matrix;
[0013] In an exemplary embodiment, the recommended volume of the Bluetooth headset is determined based on the environmental state, the user state, the content type vector and the scene similarity matrix, including: determining a target scene based on the environmental state, the user state and the content type vector; determining N reference scenes that are similar to the target scene based on the scene similarity matrix, where N is an integer greater than or equal to 0 and less than or equal to K-1; determining the recommended volume of the Bluetooth headset based on a target weight value corresponding to the target scene, N weight values corresponding to the N reference scenes, a volume recommendation value corresponding to the target scene, and N volume recommendation values corresponding to the N reference scenes.
[0014] In an exemplary embodiment, after determining the recommended volume of the Bluetooth headset based on the environmental state, the user state and the content type vector, the method further includes: when the difference between the current volume of the Bluetooth headset and the recommended volume is less than a second preset threshold, directly adjusting the volume of the Bluetooth headset to the recommended volume; when the difference between the current volume of the Bluetooth headset and the recommended volume is greater than the second preset threshold, determining the volume increase / deceleration rate based on the user's displacement trajectory and posture change rate, and dynamically adjusting the volume of the Bluetooth headset to the recommended volume based on the volume increase / deceleration rate.
[0015] According to another aspect of an embodiment of the present application, a system for determining the volume of a Bluetooth headset is also provided, including: a first determination subsystem, used to obtain ambient audio data of the environment in which the Bluetooth headset is located, and determine the environmental state based on the ambient audio data, wherein the environmental state is used to reflect the noise level of the environment in which the Bluetooth headset is located; a second determination subsystem, used to obtain motion sensor data from a target device, and determine the user state of a user using the Bluetooth headset based on the motion sensor data, wherein the target device is a device that establishes a communication connection with the Bluetooth headset; a generation subsystem, used to generate a content type vector based on the audio stream played by the Bluetooth headset, wherein the content type vector is used to characterize the speaking speed characteristics, pitch characteristics and frequency distribution characteristics of the audio stream; a third determination subsystem, used to determine the recommended volume of the Bluetooth headset based on the environmental state, the user state and the content type vector.
[0016] According to another aspect of an embodiment of the present application, a device for determining the volume of a Bluetooth headset is also provided, including: a first determination module, used to obtain ambient audio data of the environment in which the Bluetooth headset is located, and determine the environmental state based on the ambient audio data, wherein the environmental state is used to reflect the noise level of the environment in which the Bluetooth headset is located; a second determination module, used to obtain motion sensor data from a target device, and determine the user state of a user using the Bluetooth headset based on the motion sensor data, wherein the target device is a device that establishes a communication connection with the Bluetooth headset; a generation module, used to generate a content type vector based on the audio stream played by the Bluetooth headset, wherein the content type vector is used to characterize the speech speed characteristics, pitch characteristics and frequency distribution characteristics of the audio stream; a third determination module, used to determine the recommended volume of the Bluetooth headset based on the environmental state, the user state and the content type vector.
[0017] According to another aspect of the embodiment of the present application, a Bluetooth headset is also provided, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, wherein the processor implements the steps of the above-mentioned method for determining the volume of the Bluetooth headset when executing the computer program.
[0018] According to another aspect of the embodiments of the present application, a computer-readable storage medium is provided, in which a computer program is stored. The computer program is configured to execute the above-mentioned method for determining the volume of a Bluetooth headset when running.
[0019] According to another aspect of the embodiments of the present application, a computer program product is provided, including a computer program. When the computer program is executed by a processor, the method for determining the volume of a Bluetooth headset is implemented.
[0020] This application solves the problem of Bluetooth headset volume adjustment relying on manual settings by automatically analyzing environmental, user, and content characteristics. First, by collecting ambient audio data, the environmental state (noise level) is determined, achieving intelligent perception of the environment. Second, motion sensor data is obtained from the target device to analyze the user's state (such as movement or stillness) to ensure that the volume is adapted to the user's current activity. At the same time, a content type vector (used to characterize speech rate, pitch, and frequency distribution characteristics) is generated based on the audio stream, and the volume is adjusted according to different content to optimize the listening experience. Finally, the recommended volume is determined by combining the environmental state, user state, and content type vector, achieving personalized and dynamic volume adjustment. This application can achieve the following technical effects: 1) Intelligent determination of recommended volume to meet the needs of diverse scenarios and improve user experience; 2) Dynamic adaptability, with real-time recommended volume adaptation to the environment and user behavior. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present application and, together with the description, serve to explain the principles of the present application.
[0022] In order to more clearly illustrate the embodiments of the present application or the technology in the prior art, the following is a brief introduction to the drawings required for use in the embodiments or the description of the prior art. Obviously, for ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0023] Figure 1 1 is a flow chart of a method for determining the volume of a Bluetooth headset according to an embodiment of the present application; Figure 2 This is a structural block diagram of a system for determining the volume of a Bluetooth headset according to an embodiment of the present application; Figure 3 This is a structural block diagram of a device for determining the volume of a Bluetooth headset according to an embodiment of the present application. DETAILED DESCRIPTION
[0024] In order to enable those skilled in the art to better understand this application, the following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technology in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts should fall within the scope of protection of this application.
[0025] It should be noted that the terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequential order. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments of the present application described herein can be implemented in a sequence other than those illustrated or described herein. In addition, the terms "including" and "having" and any of their variations are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device comprising a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.
[0026] In this embodiment, a method for determining the volume of a Bluetooth headset is provided. Figure 1 FIG. 1 is a flow chart of a method for determining the volume of a Bluetooth headset according to an embodiment of the present application. Figure 1 As shown, the process includes the following steps S102 to S108: Step S102: Acquire ambient audio data of the environment in which the Bluetooth headset is located, and determine an environmental state based on the ambient audio data, wherein the environmental state is used to reflect the noisiness of the environment in which the Bluetooth headset is located; Optionally, the environmental state includes but is not limited to: a noisy environment and a quiet environment.
[0027] In an exemplary embodiment, the above step S102 can be implemented by the following steps S11-S14: Step S11: digitally processing the ambient audio data according to a preset sampling rate to obtain an original audio sequence, and analyzing the original audio sequence using a fast Fourier transform to obtain frequency distribution data, wherein the frequency distribution data includes amplitude information and phase information within different frequency ranges; It should be noted that the microphone of the Bluetooth headset serves as an audio input device, which can capture ambient sounds in real time and is suitable for scenarios such as environmental monitoring.
[0028] For example, the microphone captures audio at a 44.1kHz sampling rate to ensure high-fidelity reproduction of ambient sound. Choosing a sampling rate requires a balance between data accuracy and processing efficiency. 44.1kHz covers the human perceptible frequency range of 20Hz to 20kHz, making it suitable for noise analysis. After acquisition, the analog audio signal is digitized by an analog-to-digital converter to produce the original audio sequence.
[0029] For example, in a noisy cafe environment, the original audio sequence contains human voices, the sound of clinking cups and plates, and background music, and the data is represented in the time domain. A fast Fourier transform (FFT) is applied to the original audio sequence to convert the time domain signal into the frequency domain, obtaining frequency components. The FFT decomposes the signal and extracts the amplitude and phase information of each frequency. It should be noted that a frequency component consists of a frequency value and the amplitude and phase information corresponding to that frequency value.
[0030] In one possible implementation, the transformed frequency distribution data shows that the main frequency of the coffee shop environment is concentrated between 100 Hz and 5 kHz, reflecting the characteristics of the collision between human voice and utensils.
[0031] It should be noted that the window size of the fast Fourier transform affects the frequency resolution. A 1024-point window is often used to balance time and frequency accuracy. The frequency distribution data provides the basis for subsequent classification.
[0032] In one embodiment, a data preprocessing tool can be used to clean and convert the format of the raw ambient audio data to improve data quality. Specifically, the cleaning process includes removing outliers, such as spike signals generated by microphone vibrations; the format conversion standardizes the raw ambient audio data into a unified time series format. For example, assuming that the raw ambient audio data contains a 10-second audio clip, the preprocessing tool will split it into multiple 1-second clips and convert it into a numerical matrix form. This standardized data set lays the foundation for subsequent feature extraction.
[0033] Step S12: using a two-means clustering algorithm to group the frequency components in the frequency distribution data to obtain noise frequency component clusters and signal frequency component clusters; It's important to note that the K-means (K=2) clustering algorithm can be used to group frequency components in frequency distribution data. Through iterative optimization, K-means clustering categorizes frequency components into signal and noise. Based on frequency amplitude and distribution characteristics, the K-means (K=2) clustering algorithm can classify low-frequency human voices and high-frequency instrument sounds as signals, and background hum as noise.
[0034] Step S13: constructing an environmental feature vector based on the noise frequency component cluster, wherein the environmental feature vector includes: a noise decibel value corresponding to the noise frequency component cluster, a frequency range corresponding to the noise frequency component cluster, and a noise proportion; the noise proportion is equal to the ratio of the number of frequency components in the noise frequency component cluster to the number of frequency components in the frequency distribution data; Alternatively, the noise clustering at the coffee shop shows a dominant frequency range of 50 Hz to 200 Hz with low amplitude, indicating air conditioning or electrical appliance noise.
[0035] It's understandable that an environment feature vector is a multidimensional array containing noise decibel values, frequency ranges, and noise percentages. For example, a feature vector for a coffee shop might be [60dB, 50-200Hz, 30%], where 60dB represents the noise level, 50-200Hz is the frequency range, and 30% is the noise percentage.
[0036] Optionally, the environmental feature vector may also include a timestamp to reflect changes in noise over time, such as the noise level rising to 65 dB during the lunch rush hour.
[0037] Optionally, the environmental feature vector can be used for environment classification or adaptive noise reduction. For example, a Bluetooth headset can identify a coffee shop environment based on the environmental feature vector and automatically adjust the noise reduction algorithm to prioritize filtering out background noise between 50Hz and 200Hz, improving call clarity.
[0038] It's important to note that the environmental feature vector reveals that the main noise source in the cafe is electrical appliances. This allows managers to replace low-noise equipment and improve the customer experience. This has significant technical benefits. First, real-time audio processing combined with fast Fourier transforms and K-means clustering enables precise noise separation, improving the efficiency of environmental analysis. Second, this structured data, the environmental feature vector, provides the foundation for intelligent applications, supporting adaptive noise reduction for headphones or environmental optimization decisions. Logically, audio acquisition, frequency domain analysis, noise classification, and feature vector construction form a complete chain, with each step supporting each other to ensure data accuracy and application reliability.
[0039] Step S14: determining the environmental state according to the environmental feature vector using a preset noise classification model.
[0040] Optionally, the preset noise classification model is a machine learning model trained by supervised learning, and the training data includes multiple groups of data, each group of data includes a sample environment feature vector and the environment state corresponding to the sample environment feature vector.
[0041] Alternatively, the environmental state can be determined based on the environmental feature vector in the following simple manner: if the noise decibel value in the environmental feature vector exceeds a preset decibel value (based on a typical environmental noise level), a threshold comparison tool is used to determine that the environment is noisy, resulting in a first classification result; if the noise decibel value in the environmental feature vector is lower than the preset decibel value, the environment is determined to be quiet, resulting in a first classification result. The first classification result is converted to a specified format using a result output tool to obtain a final environmental classification result.
[0042] It is understood that the result output tool converts the first classification result into a specified format for easy application. For example, the classification result can be formatted as a JSON object containing the environment type and timestamp, such as {"environment":"noisy","timestamp":"2025-05-09 10:00:00"}. This structured output facilitates integration with other systems, such as allowing smart devices to adjust volume or switch to noise reduction mode based on the results. Optionally, a noise level exceeding 70 decibels is considered a noisy environment. For example, if the noise level of a feature vector is 75 decibels, the tool classifies it as a noisy environment; if it is 50 decibels, it is classified as a quiet environment. This classification method is simple and efficient, allowing for rapid response to environmental changes.
[0043] Alternatively, in real-world scenarios, Bluetooth headphones can detect high noise levels based on the above process when a user is walking on a noisy street and automatically switch to noise reduction mode; or detect low noise levels in a quiet library and turn off noise reduction to save power. This adaptive adjustment significantly improves the user experience while optimizing device performance.
[0044] Step S104: acquiring motion sensor data from a target device, and determining a user status of a user using the Bluetooth headset based on the motion sensor data, wherein the target device is a device that establishes a communication connection with the Bluetooth headset; It should be noted that when obtaining motion sensor data from the target device, the corresponding current time information will also be obtained. The target device communicates with the Bluetooth headset via the Bluetooth low energy protocol.
[0045] In an exemplary embodiment, the above step S104 can be implemented by the following steps S21-S24: Step S21: parsing the motion sensor data through the Bluetooth communication protocol between the Bluetooth headset and the target device to obtain a target data set, wherein the target data set includes the three-axis acceleration of the target device and the heart rate of the user with different timestamps; For example, when acquiring motion sensor data and current time information from a target device connected to a Bluetooth headset, the sensor typically includes a three-axis accelerometer and a heart rate sensor.
[0046] For example, the built-in accelerometer in a smartphone connected to a Bluetooth headset collects triaxial data 100 times per second, generating raw data including acceleration on the X, Y, and Z axes. Simultaneously, the heart rate sensor records heart rate values once per second, with millisecond timestamps. These data packets are transmitted to the Bluetooth headset via the Bluetooth protocol, where they are parsed to form a structured data stream containing acceleration, heart rate, and timestamps. This approach ensures real-time and complete data collection.
[0047] Step S22: Processing the target data set using a time series analysis algorithm and a Fourier transform algorithm to obtain a target feature set, wherein the target feature set includes a time period identifier, and acceleration features and cadence features within the time period corresponding to the time period identifier; Optionally, the acceleration feature can be generated by calculating the mean and variance of the motion acceleration using a sliding window, and the cadence value per unit time can be calculated using a step counting algorithm to generate a cadence feature. For example, the acceleration feature is 1.2 m / s², and the cadence feature is 1.8 steps / second.
[0048] In one possible implementation, when processing the raw data stream for time series analysis, a Fourier transform is used to extract the frequency components of motion acceleration. For example, a fast Fourier transform is applied to the acceleration data to analyze peaks in the 0.5-3 Hz frequency range to determine the user's cadence, such as a walking pace of 120 steps per minute. Timestamps are used to segment the data into minutes, generating time period identifiers such as "10:00-10:01." This processing method clearly distinguishes motion states within different time periods.
[0049] Step S23: If heart rate fluctuations exist within the time period corresponding to the time period identifier, determining the heart rate fluctuation rate within the time period corresponding to the time period identifier, and generating a user state vector based on the heart rate fluctuation rate and the target feature set, wherein the user state vector includes the time period identifier, the acceleration feature, the cadence feature, and the heart rate fluctuation rate within the time period corresponding to the time period identifier; Optionally, after obtaining the target feature set, if heart rate fluctuations exist within the time period corresponding to the time period identifier, for example, if the heart rate fluctuates between 60 and 100 beats per minute, the heart rate fluctuation amplitude per unit time can be calculated using a sliding window analysis to generate the heart rate fluctuation rate. Preferably, the window size is set to 10 seconds, and the standard deviation of the heart rate within the window is calculated to obtain the heart rate fluctuation rate.
[0050] Alternatively, a heart rate fluctuation rate exceeding 5 beats / minute indicates that the user may be engaging in high-intensity exercise. Acceleration and cadence features are combined to generate a user state vector. This approach provides a more comprehensive description of the user's state.
[0051] Optionally, after the user state vector is generated, the user state vector may be normalized to ensure comparability of different features.
[0052] For example, normalize acceleration values to the range of 0-1, cadence to 0-100 steps / minute, and heart rate fluctuation to 0-10 beats / minute. The concatenated user state vector contains normalized acceleration, cadence, time period identifiers, and heart rate fluctuation characteristics. This approach makes the vector suitable for subsequent analysis and maintains data consistency.
[0053] It's important to note that the user state vector can be represented as an array containing four-dimensional data: normalized acceleration 0.8, cadence 80, time period "10:00," and heart rate fluctuation 6. This vector can intuitively reflect the user's exercise state and physiological characteristics during a specific time period, providing a basis for subsequent state analysis.
[0054] For example, the fluctuation of heart rate can be used to infer whether a user is walking or running. High fluctuations and a fast cadence typically indicate running, while low fluctuations and a slow cadence suggest walking. This combined analysis of multi-dimensional features can more accurately characterize user activity and improve the reliability of data analysis.
[0055] It's important to note that the introduction of time period identifiers makes it easier to track a user's exercise patterns at different times of the day. For example, cadence and heart rate fluctuation data from 10 a.m. can be compared with data from 8 p.m. to analyze changes in a user's exercise habits. This granular time analysis helps generate more personalized user status reports.
[0056] For example, acquiring real-time data from wearable device sensors typically involves the coordinated operation of multiple sensors. Wearable devices such as smart bracelets collect user data through accelerometers, heart rate sensors, and pedometers. For example, the accelerometer collects triaxial acceleration data at a 50Hz frequency, recording the user's arm swing; the heart rate sensor records the heart rate value once per second, capturing the user's physiological state; the pedometer generates cadence data by detecting footstep vibrations; and the timestamp records each acquisition moment with millisecond-level accuracy. This multi-sensor collaboration ensures comprehensive data coverage, covering both movement and physiological information.
[0057] In one possible implementation, feature extraction is the core of data processing. For acceleration data, a sliding window can be set to 2 seconds, and the mean and variance of the acceleration within the window are calculated. For example, the mean of the three-axis acceleration within a window is 0.5 m / s², and the variance is 0.2, reflecting the smoothness of the user's movement. The cadence feature is calculated using a step counting algorithm. Assuming that the user takes 120 steps in 1 minute, the cadence is 2 steps / second. The heart rate fluctuation feature is calculated by the difference between the maximum and minimum heart rate within the window. For example, if the heart rate fluctuates between 70 and 80 beats / minute, the fluctuation amplitude is 10 beats / minute. The time period is identified based on the timestamp, such as marking the period from 22:00 to 6:00 as nighttime. This feature extraction method transforms raw data into analyzable indicators.
[0058] Step S24: Determine the user state according to the user state vector using a preset state classification model.
[0059] Optionally, the preset state classification model is a machine learning model trained by supervised learning, and the training data includes multiple groups of data, each group of data includes a user state vector and a user state corresponding to the user state vector.
[0060] Optionally, the user state can also be determined based on the user state vector in the following manner: a logical judgment is performed on the feature data set using a preset state classification rule to obtain a user state classification result. For example, if the acceleration feature or the cadence feature is higher than a corresponding preset threshold, the user state is determined to be in motion; if the time period is identified as nighttime and the heart rate fluctuation feature is lower than a corresponding preset threshold, the user state is determined to be in a resting state. Based on the user state classification result, the time series data of the motion state or resting state is stored in a database, and a time series analysis tool is used to smooth the time series data to generate a continuous user state sequence to determine the final user state.
[0061] For example, if the average acceleration exceeds 1m / s² or the cadence is higher than 1.5 steps / second, the user is judged to be in motion. For example, when the user is walking quickly, the average acceleration is 1.2m / s² and the cadence is 1.8 steps / second, which meets the conditions for motion. If it is nighttime and the heart rate fluctuation is less than 5 beats / minute, the user is judged to be in a resting state. For example, if the user's heart rate fluctuates by 3 beats / minute at 2:00 a.m., it indicates sleep. This rule clearly distinguishes user states.
[0062] It should be noted that the storage and smoothing of time series data further optimizes the state sequence. Activity and rest states are stored in the database using timestamps as indexes, for example, recording a user's activity status from 8:00 to 8:30. Smoothing can use a moving average method. For example, if a user's activity status is incorrectly classified as resting for 5 consecutive minutes, smoothing can correct the error to continuous activity. This process ensures the consistency of the state sequence and improves the reliability of analysis.
[0063] In one embodiment, the generation of a continuous user status sequence facilitates long-term monitoring. For example, a user's status sequence generated over a single day might show exercise from 8:00 AM to 9:00 AM and rest from 10:00 PM to 6:00 AM, reflecting their regular sleep schedule. This sequence can be used for health management, reminding users to maintain exercise or improve sleep.
[0064] Optionally, the above method forms a complete user status monitoring process through multi-sensor data fusion, feature extraction, status classification, and sequence generation. Each link supports each other, ensuring a logically rigorous data collection and analysis process, thereby improving the user experience.
[0065] Step S106: generating a content type vector according to the audio stream played by the Bluetooth headset, wherein the content type vector is used to characterize the speech rate characteristics, pitch characteristics, and frequency distribution characteristics of the audio stream; In an exemplary embodiment, the above step S106 may be performed through the following steps S31-S34: Step S31: using an audio decoder to parse the audio stream to obtain a metadata set, wherein the metadata set includes a timestamp, encoding format, and bit rate corresponding to the audio stream; Alternatively, metadata from the audio stream played by Bluetooth headsets is the basis for content classification. Bluetooth headset audio streams typically contain rich metadata, such as timestamps, encoding formats, and bit rates. This metadata is parsed by an audio decoder to form a structured metadata collection.
[0066] For example, a user listening to an MP3 song through Bluetooth headphones will decode the music and obtain metadata with a timestamp of 2025-05-09 10:00:00, an MP3 encoding format, and a bit rate of 320kbps. This parsing method ensures the accuracy of subsequent analysis because the metadata provides reliable raw information for feature extraction.
[0067] Step S32: Analyzing the audio stream based on the metadata set using a fast Fourier transform algorithm to obtain a reference feature set, wherein the reference feature set includes a speech rate feature and a pitch feature corresponding to the audio stream; Optionally, a content classification algorithm is used to analyze the audio stream based on the metadata set. The core is to calculate the audio frequency distribution through fast Fourier transform. Fast Fourier transform is a tool that converts time domain signals into frequency domain, which is used to extract the frequency characteristics of audio. For example, when analyzing a podcast audio, the fast Fourier transform can identify that the main frequencies are concentrated in the range of 200Hz to 3000Hz, indicating that the human voice is the main component. Subsequently, a feature set is formed by extracting speech rate features and pitch features. The speech rate feature can be calculated by detecting the interval time of the speech segments in the audio, such as a speech rate of 80 words per minute; the pitch feature is determined by the proportion of high and low notes in the frequency distribution, such as 30% of high notes. These features provide a quantitative basis for subsequent classification.
[0068] Step S33: determining the content type corresponding to the audio stream according to the reference feature set; Optionally, the reference feature set is input into a preset neural network model to determine the content type corresponding to the audio stream. Content types include but are not limited to: "podcast", "music", and "audiobook".
[0069] Step S34: Generate a content type vector based on the reference feature set, the frequency distribution features corresponding to the audio stream, and the first weight, second weight, and third weight corresponding to the content type, wherein the content type vector includes the evaluation values of the audio stream on the speaking rate feature, pitch feature, and frequency distribution feature, respectively, the first weight is the weight corresponding to the speaking rate feature, the second weight is the weight corresponding to the pitch feature, and the third weight is the weight corresponding to the frequency distribution feature.
[0070] The reference feature set and the frequency distribution features corresponding to the audio stream can be mapped into a three-dimensional vector. Assume that the content type corresponding to the audio stream is podcast, the corresponding speech rate feature weight is 0.4, the pitch feature weight is 0.3, the frequency distribution feature weight is 0.3, the speech rate feature initial score is 0.9, the pitch feature initial score is 0.7, and the frequency distribution feature initial score is 0.6. The resulting content type vector can be [0.36, 0.21, 0.18]. This vector representation simplifies complex audio features into a computable mathematical form, facilitating subsequent analysis and storage.
[0071] In one embodiment, the application scenarios of the present application may be expanded. For example, a Bluetooth headset may adjust the sound effect mode based on the classification results, such as enhancing the mid-frequency vocals in podcast audio and enhancing the low-frequency bass in music audio.
[0072] Step S108: Determine the recommended volume of the Bluetooth headset according to the environmental state, the user state, and the content type vector.
[0073] In an exemplary embodiment, the above-mentioned step S108 can be implemented in the following manner: based on preset rules, determining a first recommended volume according to the environmental state, determining a second recommended volume according to the user state, and determining a third recommended volume according to the content type vector; and performing a weighted summation of the first recommended volume, the second recommended volume, and the third recommended volume to obtain the recommended volume of the Bluetooth headset.
[0074] In an exemplary embodiment, before the above step S108, the method further includes the following steps S41-S44: Step S41: Acquire multiple historical environment feature vectors, historical user state vectors, and historical content type vectors, wherein the historical environment feature vector is an environment feature vector obtained by processing historical environment audio data, and the historical user state vector is a user state vector obtained by processing historical motion sensor data; Step S42: normalizing the multiple historical environment feature vectors, historical user state vectors, and historical content type vectors, and clustering the multiple vectors obtained after the normalization using a K-means clustering algorithm to obtain K clusters, wherein the K clusters correspond one-to-one to the K scenarios; It should be noted that the purpose of normalizing the multiple historical environment feature vectors, historical user state vectors, and historical content type vectors is to normalize these vectors to a unified scale and eliminate dimensional differences. It should be noted that the normalized historical environment feature vectors, historical user state vectors, and historical content type vectors are all three-dimensional feature vectors, where the values in the normalized historical environment feature vectors correspond to the noise decibel value, frequency range, and noise ratio, respectively; the values in the normalized historical user state vectors correspond to the acceleration characteristics, cadence characteristics, and heart rate fluctuation rate within the time period, respectively; and the values in the normalized historical content type vectors correspond to the speech rate characteristics, pitch characteristics, and frequency distribution characteristics, respectively.
[0075] It should be noted that when the Euclidean distance is used to calculate the distance between vectors, assuming that an environment feature vector is represented as [0.8, 0.3, 0.5] and the user state vector is [0.6, 0.7, 0.2], the Euclidean distance between the two vectors is calculated to obtain clustering groups.
[0076] Alternatively, the clustering result may group quiet environments, low heart rates, and slow speech speed audio into one group, indicating that the user is in a relaxing scene. The K scenes include but are not limited to relaxing scenes, active scenes, and focused scenes.
[0077] Step S43: generating a scene transition probability matrix based on the K clusters using a Markov chain method, wherein the scene transition probability matrix includes transition probabilities corresponding to any two scenes in the K scenes; Optionally, a dynamic scene transition probability model is constructed based on the clustering results. Scene transition refers to the user switching from one state to another, for example, from a quiet indoor environment to a noisy outdoor environment. Using a Markov chain approach, the transition frequencies between clustering categories are statistically analyzed to generate a scene transition probability matrix. Assuming three scenes: relaxed, active, and focused, statistics show that the transition probability from relaxed to active is 0.4, and from relaxed to focused is 0.3. A transition probability matrix is generated. This matrix reflects the probability of scene transitions and helps predict user behavior patterns.
[0078] Step S44: Calculating a scene similarity matrix based on the scene transition probability matrix and the eigenvectors corresponding to the K clusters, wherein the scene similarity matrix includes the scene similarity of any two scenes among the K scenes; when it is determined according to the scene transition probability matrix that the transition probabilities corresponding to two scenes are greater than or equal to a first preset threshold, calculating the scene similarity of the two scenes based on the eigenvectors corresponding to the clusters corresponding to the two scenes; when it is determined according to the scene transition probability matrix that the transition probabilities corresponding to two scenes are less than the first preset threshold, determining that the scene similarity of the two scenes is equal to zero; Optionally, if the transition probability for a particular category exceeds a preset threshold—for example, the probability of transitioning from relaxed to active exceeds 0.35—then the relationship between clustered categories is further analyzed using a spatiotemporal similarity metric. This spatiotemporal similarity metric is based on the cosine similarity method, comparing the angles between vectors. Assuming the feature vectors for the relaxed scenario are [0.7, 0.2, 0.4] and the active scenario are [0.6, 0.5, 0.3], the cosine similarity value is calculated to be 0.85, indicating that the two scenarios are relatively close in feature distribution. This similarity value reflects the strength of the association between the scenarios.
[0079] For example, a scene similarity matrix is generated based on spatiotemporal similarity values. Each entry in the matrix represents the similarity between two scenes. For example, the similarity between relaxation and activity is 0.85, while that between relaxation and focus is 0.78. Through matrix operations, these values are entered into the matrix to form a scene similarity matrix. This matrix can be used to analyze potential connections between scenes, for example, identifying which scenes are easily converted into each other, thereby optimizing content recommendation logic.
[0080] The advantage of this approach lies in its dynamic modeling of scenarios through clustering and similarity analysis of multi-dimensional features. For example, in an audio content recommendation scenario, the system can prioritize faster-paced audio content based on the user's transition from a relaxed to an active mood. This approach improves content adaptability and enhances the consistency of the user experience.
[0081] In one embodiment, the initial center point selection for K-means clustering may affect the stability of the results. This can be optimized through multiple random initializations. For example, K-means clustering can be run 10 times for the environmental feature vector, and the result with the smallest variance is selected as the final grouping. This method improves the reliability of the clustering results.
[0082] It's important to note that the combined use of the scene transition probability matrix and the similarity matrix provides multi-level support for dynamic scene analysis. For example, the probability matrix predicts the likelihood of scene switching, while the similarity matrix reveals the characteristic correlations between switching scenes. Together, the two can more accurately capture changes in user needs.
[0083] For example, suppose a user is listening to slow music in a quiet environment. The system detects an increase in ambient noise and, using the transition probability matrix, predicts that the user is likely entering an active scene. It also uses the similarity matrix to determine the relevance of the active scene to the current scene, and then recommends fast-paced music suitable for the active scene. This multi-faceted analysis ensures the rigor and consistency of the recommendation logic.
[0084] In an exemplary embodiment, the above step S108 may also be implemented in the following manner: determining the recommended volume of the Bluetooth headset according to the environmental state, the user state, the content type vector, and the scene similarity matrix.
[0085] In an exemplary embodiment, determining the recommended volume of the Bluetooth headset according to the environment state, the user state, the content type vector, and the scene similarity matrix may be performed through the following steps S51-S53: Step S51: determining a target scene according to the environment state, the user state and the content type vector; Optionally, the target scene may be determined according to the environmental state, the user state and the content type vector by a preset scene classification model, wherein the preset scene classification model is a neural network model trained using supervised learning.
[0086] Step S52: determining N reference scenes similar to the target scene according to the scene similarity matrix, where N is an integer greater than or equal to 0 and less than or equal to K-1; Step S53: Determine the recommended volume of the Bluetooth headset according to the target weight value corresponding to the target scene, the N weight values corresponding to the N reference scenes, the volume recommendation value corresponding to the target scene, and the N volume recommendation values corresponding to the N reference scenes.
[0087] Optionally, the target weight value corresponding to the target scene and the N weight values corresponding to the N reference scenes can be determined based on a similarity value between the target scene and the N reference scenes. If the similarity value between the target scene and the N reference scenes is greater, the difference between the target weight value corresponding to the target scene and the N weight values corresponding to the N reference scenes is smaller; if the similarity value between the target scene and the N reference scenes is smaller, the difference between the target weight value corresponding to the target scene and the N weight values corresponding to the N reference scenes is greater.
[0088] Optionally, assuming that the target weight value corresponding to the target scene is 0.5, N is equal to 2, the N weight values corresponding to the N reference scenes are 0.2 and 0.3 respectively, the recommended volume value corresponding to the target scene is 30 volume units, and the N recommended volume values corresponding to the N reference scenes are 35 volume units and 40 volume units respectively, then the recommended volume of the Bluetooth headset = 0.5*30+0.2*35+0.3*40=34.
[0089] In an exemplary embodiment, after the above step S108, the method further includes: when the difference between the current volume of the Bluetooth headset and the recommended volume is less than a second preset threshold, directly adjusting the volume of the Bluetooth headset to the recommended volume; when the difference between the current volume of the Bluetooth headset and the recommended volume is greater than the second preset threshold, determining the volume increase / deceleration rate according to the user's displacement trajectory and posture change rate, and dynamically adjusting the volume of the Bluetooth headset to the recommended volume according to the volume increase / deceleration rate. Optionally, when the difference between the current volume of the Bluetooth headset and the recommended volume is greater than a second preset threshold, a smooth interpolation method can also be used to process the recommended volume value to generate a continuous volume reference value to obtain the volume reference value. The current volume value of the device is obtained, and the deviation between the volume reference value and the current volume value is calculated. If the deviation exceeds the preset deviation threshold, the dynamic adjustment trigger signal is determined. According to the dynamic adjustment trigger signal, real-time displacement trajectory data and posture change rate data are obtained through the device sensor, and a weighted average method is used to calculate the comprehensive motion factor of the displacement trajectory data and the posture change rate data to obtain the volume increase / deceleration rate. The volume reference value is adjusted according to the volume increase / deceleration rate, and the volume of the Bluetooth headset is dynamically adjusted to the recommended volume based on the volume reference value.
[0090] In one possible implementation, a smooth interpolation method is used to process the recommended volume values, generating a continuous volume baseline. For example, if the device switches from a library to a cafe within 10 seconds, the volume needs to transition from 30dB to 50dB. Smooth interpolation breaks down this 20dB difference into a gradual 2dB per second increment, ensuring a natural volume change and avoiding abrupt jumps. This smooth transition is crucial to the user's listening experience.
[0091] Alternatively, when calculating the deviation between the volume reference value and the current volume, assume the current volume is 40 decibels, the reference value is 50 decibels, and the deviation is 10 decibels. If the preset deviation threshold reaches 5 decibels, a dynamic adjustment signal is triggered. This deviation detection mechanism can promptly detect any mismatch between the volume and the environment, initiating subsequent adjustments.
[0092] Optionally, obtaining real-time displacement trajectory data and attitude change rate data can be achieved through the accelerometer and gyroscope of the device.
[0093] For example, if a user moves from sitting to walking in a cafe, the displacement trajectory data will be reflected as 0.5 meters per second, and the posture change rate will be displayed as a tilt change of 10 degrees per second. The weighted average method assigns 60% weight to the displacement and 40% weight to the posture change rate. The combined calculation results in a motion factor that reflects the user's dynamic level. This motion factor accurately represents the potential impact of user behavior on volume requirements.
[0094] In one embodiment, when adjusting the volume baseline based on the volume increase / deceleration rate, if the motion factor indicates a highly dynamic user, the volume increase / deceleration rate might be 3 decibels per second, adjusting the baseline value of 50 decibels to 53 decibels. Conversely, if the user is stationary, the rate might be -2 decibels per second, adjusting the baseline value to 48 decibels. This dynamic adjustment optimizes the volume in real time based on the user's state.
[0095] Through the description of the above embodiments, those skilled in the art can clearly understand that the methods according to the above embodiments can be implemented by means of software plus the necessary general hardware platform. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technology of this application, or the part that contributes to the prior art, can be embodied in the form of a software product. The computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), and includes a number of instructions for enabling a terminal device (which can be a mobile phone, computer, server, or network device, etc.) to execute the methods of each embodiment of the present application.
[0096] This embodiment also provides a system and device for determining the volume of a Bluetooth headset. The system and device are used to implement the above-mentioned embodiments and preferred embodiments, and details already described will not be repeated. As used below, the term "module" may refer to a combination of software and / or hardware that implements a predetermined function. Although the devices described in the following embodiments are preferably implemented in software, implementation using hardware, or a combination of software and hardware, is also possible and contemplated.
[0097] Figure 2 : is a structural block diagram of a system for determining the volume of a Bluetooth headset according to an embodiment of the present application, the system comprising: A first determining subsystem 202 is configured to obtain ambient audio data of an environment in which the Bluetooth headset is located, and determine an environmental state based on the ambient audio data, wherein the environmental state is used to reflect the noisiness of the environment in which the Bluetooth headset is located; a second determining subsystem 204 for acquiring motion sensor data from a target device, and determining a user status of a user using the Bluetooth headset based on the motion sensor data, wherein the target device is a device that establishes a communication connection with the Bluetooth headset; A generating subsystem 206 is configured to generate a content type vector based on the audio stream played by the Bluetooth headset, wherein the content type vector is used to characterize the speech rate characteristics, pitch characteristics, and frequency distribution characteristics of the audio stream; The third determining subsystem 208 is configured to determine a recommended volume for the Bluetooth headset according to the environmental state, the user state, and the content type vector.
[0098] The above system solves the problem of manual volume adjustment for Bluetooth headsets by automatically analyzing environmental, user, and content characteristics. First, the environmental state (noise level) is determined by collecting environmental audio data, thereby achieving intelligent perception of the environment. Second, motion sensor data is obtained from the target device to analyze the user's state (such as movement or stillness) to ensure that the volume is adapted to the user's current activity. At the same time, a content type vector (used to characterize speech rate, pitch, and frequency distribution characteristics) is generated based on the audio stream, and the volume is adjusted according to different content to optimize the listening experience. Finally, the recommended volume is determined by combining the environmental state, user state, and content type vector, achieving personalized and dynamic volume adjustment. This application can achieve the following technical effects: 1) Intelligent determination of recommended volume to meet the needs of diverse scenarios and improve user experience; 2) Dynamic adaptability, with real-time recommended volume adaptation to the environment and user behavior.
[0099] Figure 3 : is a structural block diagram of a device for determining the volume of a Bluetooth headset according to an embodiment of the present application, the device comprising: A first determining module 302 is configured to obtain ambient audio data of an environment in which the Bluetooth headset is located, and determine an environmental state based on the ambient audio data, wherein the environmental state is used to reflect the noisiness of the environment in which the Bluetooth headset is located; a second determining module 304, configured to obtain motion sensor data from a target device, and determine a user status of a user using the Bluetooth headset based on the motion sensor data, wherein the target device is a device that establishes a communication connection with the Bluetooth headset; A generating module 306 is configured to generate a content type vector according to the audio stream played by the Bluetooth headset, wherein the content type vector is used to characterize the speech rate characteristics, pitch characteristics, and frequency distribution characteristics of the audio stream; The third determining module 308 is configured to determine a recommended volume for the Bluetooth headset according to the environment state, the user state, and the content type vector.
[0100] The above-mentioned device solves the problem of manual volume adjustment for Bluetooth headsets by automatically analyzing environmental, user, and content characteristics. First, the environmental state (noise level) is determined by collecting environmental audio data, thereby achieving intelligent perception of the environment. Second, motion sensor data is obtained from the target device to analyze the user's state (such as movement or stillness) to ensure that the volume is adapted to the user's current activity. At the same time, a content type vector (used to represent speech rate, pitch, and frequency distribution characteristics) is generated based on the audio stream, and the volume is adjusted according to different content to optimize the listening experience. Finally, the recommended volume is determined by combining the environmental state, user state, and content type vector, achieving personalized and dynamic volume adjustment. This application can achieve the following technical effects: 1) Intelligent determination of recommended volume to meet the needs of diverse scenarios and improve user experience; 2) Dynamic adaptability, with real-time recommended volume adaptation to the environment and user behavior.
[0101] An embodiment of the present application further provides a computer-readable storage medium, in which a computer program is stored. The computer program is configured to execute the steps of any of the above method embodiments when run.
[0102] Optionally, in this embodiment, the above-mentioned storage medium may be configured to store a computer program for executing the following steps.
[0103] In an exemplary embodiment, the computer-readable storage medium may include, but is not limited to, various media that can store computer programs, such as a USB flash drive, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk, or an optical disk.
[0104] For specific examples in this embodiment, reference may be made to the examples described in the above embodiments and exemplary implementation modes, and this embodiment will not be described in detail here.
[0105] An embodiment of the present application further provides a computer program product, including a computer program, and the computer program performs the steps of any of the above method embodiments when executed by a processor.
[0106] An embodiment of the present application further provides another computer program product, comprising a non-volatile computer-readable storage medium, wherein the non-volatile computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of any of the above method embodiments are implemented.
[0107] An embodiment of the present application further provides an electronic device, which includes a memory and a processor, wherein a computer program is stored in the memory, and the processor is configured to execute the steps of any of the above method embodiments through the computer program.
[0108] For specific examples in this embodiment, reference may be made to the examples described in the above embodiments and exemplary implementation modes, and this embodiment will not be described in detail here.
[0109] Obviously, those skilled in the art should understand that the modules or steps of the present application described above can be implemented using a general-purpose computing device, they can be concentrated on a single computing device, or distributed across a network composed of multiple computing devices, they can be implemented using program code executable by the computing device, and thus, they can be stored in a storage device and executed by the computing device, and in some cases, the steps shown or described can be performed in a different order than herein, or they can be fabricated into separate integrated circuit modules, or multiple modules or steps can be fabricated into a single integrated circuit module for implementation. Thus, the present application is not limited to any specific combination of hardware and software.
[0110] The above is only a preferred embodiment of the present application. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present application. These improvements and modifications should also be regarded as the scope of protection of the present application.
Claims
1. A method for determining the volume of a Bluetooth headset, characterized in that: include: Acquire ambient audio data of an environment in which the Bluetooth headset is located, and determine an environmental state based on the ambient audio data, wherein the environmental state is used to reflect the noisiness of the environment in which the Bluetooth headset is located; Acquiring motion sensor data from a target device, and determining a user status of a user using the Bluetooth headset based on the motion sensor data, wherein the target device is a device that establishes a communication connection with the Bluetooth headset; Generating a content type vector according to the audio stream played by the Bluetooth headset, wherein the content type vector is used to characterize the speech rate characteristics, pitch characteristics, and frequency distribution characteristics of the audio stream; Determine a recommended volume for the Bluetooth headset according to the environmental state, the user state, and the content type vector.
2. The method according to claim 1, characterized in that Determining an environmental state according to the environmental audio data includes: Digitally processing the ambient audio data according to a preset sampling rate to obtain an original audio sequence, and analyzing the original audio sequence using a fast Fourier transform to obtain frequency distribution data, wherein the frequency distribution data includes amplitude information and phase information within different frequency ranges; Using a two-means clustering algorithm to group the frequency components in the frequency distribution data to obtain noise frequency component clusters and signal frequency component clusters; Constructing an environmental feature vector based on the noise frequency component cluster, wherein the environmental feature vector includes: a noise decibel value corresponding to the noise frequency component cluster, a frequency range corresponding to the noise frequency component cluster, and a noise proportion; the noise proportion is equal to the ratio of the number of frequency components in the noise frequency component cluster to the number of frequency components in the frequency distribution data; The environmental state is determined according to the environmental feature vector by using a preset noise classification model.
3. The method according to claim 1, characterized in that Determining a user status of a user using the Bluetooth headset according to the motion sensor data includes: parsing the motion sensor data through a Bluetooth communication protocol between the Bluetooth headset and the target device to obtain a target data set, wherein the target data set includes three-axis acceleration of the target device and the heart rate of the user with different timestamps; The target data set is processed using a time series analysis algorithm and a Fourier transform algorithm to obtain a target feature set, wherein the target feature set includes a time period identifier, and an acceleration feature and a cadence feature within the time period corresponding to the time period identifier; In a case where heart rate fluctuations exist within the time period corresponding to the time period identifier, determining the heart rate fluctuation rate within the time period corresponding to the time period identifier, and generating a user state vector based on the heart rate fluctuation rate and the target feature set, wherein the user state vector includes the time period identifier, the acceleration feature, the cadence feature, and the heart rate fluctuation rate within the time period corresponding to the time period identifier; The user state is determined according to the user state vector using a preset state classification model.
4. The method according to claim 1, wherein Generate a content type vector according to the audio stream played by the Bluetooth headset, including: Parsing the audio stream using an audio decoder to obtain a metadata set, wherein the metadata set includes a timestamp, an encoding format, and a bit rate corresponding to the audio stream; Analyzing the audio stream based on the metadata set using a fast Fourier transform algorithm to obtain a reference feature set, wherein the reference feature set includes a speech rate feature and a pitch feature corresponding to the audio stream; Determining a content type corresponding to the audio stream according to the reference feature set; A content type vector is generated based on the reference feature set, the frequency distribution features corresponding to the audio stream, and the first weight, second weight, and third weight corresponding to the content type, wherein the content type vector includes evaluation values of the audio stream on the speaking rate feature, pitch feature, and frequency distribution feature, respectively, the first weight is the weight corresponding to the speaking rate feature, the second weight is the weight corresponding to the pitch feature, and the third weight is the weight corresponding to the frequency distribution feature.
5. The method according to claim 1, wherein Before determining the recommended volume of the Bluetooth headset according to the environmental state, the user state, and the content type vector, the method further includes: Acquire multiple historical environment feature vectors, historical user state vectors, and historical content type vectors, wherein the historical environment feature vector is an environment feature vector obtained by processing historical environment audio data, and the historical user state vector is a user state vector obtained by processing historical motion sensor data; Normalizing the multiple historical environment feature vectors, historical user state vectors, and historical content type vectors, and clustering the multiple vectors obtained after the normalization using a K-means clustering algorithm to obtain K clusters, wherein the K clusters correspond one-to-one to the K scenarios; Generate a scene transition probability matrix based on the K clusters using a Markov chain method, wherein the scene transition probability matrix includes transition probabilities corresponding to any two scenes in the K scenes; A scene similarity matrix is calculated based on the scene transition probability matrix and the eigenvectors corresponding to the K clusters, wherein the scene similarity matrix includes the scene similarity of any two scenes among the K scenes; when it is determined according to the scene transition probability matrix that the transition probabilities corresponding to two scenes are greater than or equal to a first preset threshold, the scene similarity of the two scenes is calculated based on the eigenvectors corresponding to the clusters corresponding to the two scenes; when it is determined according to the scene transition probability matrix that the transition probabilities corresponding to two scenes are less than the first preset threshold, the scene similarity of the two scenes is determined to be zero; Determining the recommended volume of the Bluetooth headset according to the environmental state, the user state, and the content type vector includes: determining the recommended volume of the Bluetooth headset according to the environmental state, the user state, the content type vector, and the scene similarity matrix.
6. The method according to claim 5, characterized in that Determining a recommended volume for the Bluetooth headset according to the environment state, the user state, the content type vector, and the scene similarity matrix includes: determining a target scenario according to the environment state, the user state, and the content type vector; Determining N reference scenes similar to the target scene according to the scene similarity matrix, where N is an integer greater than or equal to 0 and less than or equal to K-1; The recommended volume of the Bluetooth headset is determined according to the target weight value corresponding to the target scene, the N weight values corresponding to the N reference scenes, the volume recommendation value corresponding to the target scene, and the N volume recommendation values corresponding to the N reference scenes.
7. The method according to claim 1, characterized in that After determining the recommended volume of the Bluetooth headset according to the environmental state, the user state, and the content type vector, the method further includes: When the difference between the current volume of the Bluetooth headset and the recommended volume is less than a second preset threshold, directly adjusting the volume of the Bluetooth headset to the recommended volume; When the difference between the current volume of the Bluetooth headset and the recommended volume is greater than a second preset threshold, the volume increase / deceleration rate is determined according to the displacement trajectory and posture change rate of the user, and the volume of the Bluetooth headset is dynamically adjusted to the recommended volume according to the volume increase / deceleration rate.
8. A system for determining the volume of a Bluetooth headset, characterized in that: include: a first determining subsystem, configured to obtain ambient audio data of an environment in which the Bluetooth headset is located, and determine an environmental state based on the ambient audio data, wherein the environmental state is used to reflect the noisiness of the environment in which the Bluetooth headset is located; a second determining subsystem, configured to obtain motion sensor data from a target device, and determine a user status of a user using the Bluetooth headset based on the motion sensor data, wherein the target device is a device that establishes a communication connection with the Bluetooth headset; A generating subsystem, configured to generate a content type vector according to the audio stream played by the Bluetooth headset, wherein the content type vector is used to characterize the speech rate characteristics, pitch characteristics, and frequency distribution characteristics of the audio stream; The third determination subsystem is configured to determine a recommended volume for the Bluetooth headset according to the environmental state, the user state, and the content type vector.
9. A device for determining the volume of a Bluetooth headset, characterized in that: include: a first determining module, configured to obtain ambient audio data of an environment in which the Bluetooth headset is located, and determine an environmental state based on the ambient audio data, wherein the environmental state is used to reflect the noisiness of the environment in which the Bluetooth headset is located; a second determining module, configured to obtain motion sensor data from a target device and determine a user status of a user using the Bluetooth headset based on the motion sensor data, wherein the target device is a device that establishes a communication connection with the Bluetooth headset; A generating module, configured to generate a content type vector according to the audio stream played by the Bluetooth headset, wherein the content type vector is used to characterize the speech rate characteristics, pitch characteristics, and frequency distribution characteristics of the audio stream; The third determining module is configured to determine a recommended volume for the Bluetooth headset according to the environmental state, the user state, and the content type vector.
10. A Bluetooth headset comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 7 are implemented.