A Multi-Mode Interaction Method between a Bluetooth Speaker and a Smart Phone
By setting an adjustable angle magnet ring and magnetic adsorption between the Bluetooth speaker and the smartphone, combined with machine learning algorithms, detecting the physical adsorption state and generating a personalized parameter solution, the problem of inflexible interaction mode in the existing technology is solved, and adaptive multi-mode audio and video experience optimization is achieved.
Patent Information
- Application Number
- CN202411548867.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-01
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2044-11-01
AI Technical Summary
The interaction mode between existing Bluetooth speakers and smartphones lacks flexibility, and it is impossible to achieve accurate support angle adjustments under different adsorption states. The interaction mode is single, and it is impossible to fully optimize the user's audio and video experience.
By setting up an adjustable angle magnet ring on the Bluetooth speaker, magnetic adsorption with the smartphone, different physical adsorption states are detected, and user usage data is analyzed using machine learning algorithms, personalized parameter schemes are generated, and sound effects and functions are automatically adjusted to optimize the user experience.
It realizes adaptive adjustments based on user habits and environmental changes, improves the user's audio and video experience in different scenarios, and provides flexible multi-mode interaction and personalized sound effect optimization.
Smart Images

Figure CN119485789B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of Bluetooth speakers, and particularly to a multi-mode interaction method between a Bluetooth speaker and a smart phone. Background Art
[0002] Currently, the interaction between a Bluetooth speaker and a smart phone mainly relies on wireless connection to achieve audio playback or call functions. Such speakers are usually equipped with different brackets or accessories for fixing the phone on the speaker, so that users can obtain a better visual experience when watching videos or making video calls. At the same time, some speakers improve the audio effect through preset sound effect modes to meet different usage scenarios.
[0003] However, the physical combination method between the existing speakers and smart phones often lacks flexibility and cannot achieve precise adjustment of the support angle in different adsorption states. In addition, the existing interaction modes are relatively single, lacking personalized adaptation to different user behaviors and environmental conditions, and cannot fully optimize the user's audio and video experience. This results in the need for users to manually adjust the sound effect when switching scenarios, and the usage experience is rather cumbersome.
[0004] Therefore, there is an urgent need to develop a new multi-mode interaction method between a Bluetooth speaker and a smart phone. Summary of the Invention
[0005] The present application provides a multi-mode interaction method between a Bluetooth speaker and a smart phone to improve the user's audio and video experience.
[0006] The present application provides a multi-mode interaction method between a Bluetooth speaker and a smart phone, including:
[0007] Setting a first magnet ring capable of adjusting the angle on the Bluetooth speaker, where the first magnet ring is used for magnetic adsorption with a second magnet ring on the smart phone, so as to achieve the physical connection between the Bluetooth speaker and the smart phone;
[0008] Establishing a wireless connection between the Bluetooth speaker and the smart phone through the Bluetooth protocol;
[0009] Detecting the physical adsorption state between the smart phone and the Bluetooth speaker, where the physical adsorption state includes a first adsorption state, a second adsorption state, and a third adsorption state; wherein, the first adsorption state means that the smart phone is adsorbed on the Bluetooth speaker, the first magnet ring and the second magnet ring are magnetically adsorbed, and the smart phone forms a video viewing angle of 50 - 70 degrees with the horizontal plane; the second adsorption state means that the smart phone is adsorbed on the Bluetooth speaker, the first magnet ring and the second magnet ring are magnetically adsorbed, and the smart phone forms a video call angle of 0 - 15 degrees with the horizontal plane; the third adsorption state means that the Bluetooth speaker is not adsorbed with the smart phone;
[0010] Enter the corresponding interaction mode according to the detected physical adsorption state;
[0011] Use machine learning algorithms to analyze the usage data of users collected in different interaction modes, and generate personalized parameter solutions for different physical adsorption states;
[0012] When the corresponding physical adsorption state is detected, apply the corresponding personalized parameter solution to optimize the user experience.
[0013] The beneficial effects of the technical solution provided by this application include:
[0014] (1) By analyzing the usage data of users in different interaction modes through machine learning algorithms, personalized parameter solutions are generated for each physical adsorption state, realizing intelligent adjustment of parameters such as sound effects, and improving the audio-visual experience of users in different scenarios. (2) The system can apply the corresponding personalized parameter solution in real time when different physical adsorption states are detected, enabling the device to adaptively adjust according to the user's usage habits and environmental changes, thus significantly improving the overall user experience. Brief Description of the Drawings
[0015] Figure 1 is a flowchart of a multi-mode interaction method between a Bluetooth speaker and a smartphone provided in the first embodiment of this application.
[0016] Figure 2 is a schematic diagram of the connection between a Bluetooth speaker and a smartphone involved in the first embodiment of this application. Detailed Description of the Embodiment
[0017] Many specific details are set forth in the following description in order to provide a thorough understanding of this application. However, this application can be implemented in many other ways different from those described herein, and those skilled in the art can make similar generalizations without departing from the connotation of this application. Therefore, this application is not limited by the specific embodiments disclosed below.
[0018] The first embodiment of this application provides a multi-mode interaction method between a Bluetooth speaker and a smartphone. Please refer to Figure 1 , this figure is a schematic diagram of the first embodiment of this application. The following combines Figure 1 to detail the multi-mode interaction method between a Bluetooth speaker and a smartphone provided in the first embodiment of this application.
[0019] Step S101: Set a first magnet ring capable of adjusting the angle on the Bluetooth speaker. The first magnet ring is used for magnetic adsorption with a second magnet ring on the smartphone, so as to realize the physical connection between the Bluetooth speaker and the smartphone.
[0020] In step S101, first, an adjustable-angle first magnet ring needs to be set at the bottom or side of the Bluetooth speaker. The material of this magnet ring should be a material with strong magnetism, such as neodymium magnets or similar high-magnetic alloys, to ensure that it can generate a stable magnetic adsorption with the second magnet ring on the smartphone at different angles.
[0021] During the specific installation process, the first magnet ring can be embedded in the housing structure of the Bluetooth speaker to ensure that it can rotate or tilt in different directions. This design can be achieved by setting an adjustable hinge mechanism or sliding guide rail inside the speaker, allowing the first magnet ring to align with the second magnet ring on the smartphone at an appropriate angle. The adjustment mechanism should have sufficient damping force to ensure that the adjusted angle can remain stable during the adsorption process and will not be displaced or slide due to the weight of the smartphone.
[0022] During implementation, the first magnet ring on the Bluetooth speaker should be aligned with the second magnet ring on the smartphone to ensure that the two magnet rings can generate sufficient magnetic attraction when approaching, to achieve a fast and firm physical connection. During the adsorption process between the smartphone and the Bluetooth speaker, considering the weight and size differences of different models of smartphones, the adsorption force of the magnet ring should be able to adapt to a wide range of device types and ensure stability in various physical adsorption states.
[0023] In practical applications, to enhance the user experience, a layer of anti-slip material (such as silicone or rubber coating) can be covered on the surface of the first magnet ring to prevent the smartphone from sliding in the adsorbed state. At the same time, to prevent potential interference of magnetism on the internal electronic components of the smartphone, a shielding layer (such as an aluminum foil layer or a conductive foam layer) can be added between the magnet ring and the speaker housing to effectively shield unnecessary magnetic field leakage.
[0024] The implementation of this step needs to consider the impact of magnetic adsorption on the physical connection between the smartphone and the Bluetooth speaker, ensure stable physical support during the adsorption process, and allow flexible adjustment at different angles. Through the above method, this step can provide a reliable physical connection basis for the multi-mode interaction between the Bluetooth speaker and the smartphone.
[0025] Figure 2 It is a schematic diagram of the connection between the Bluetooth speaker 2 and the smartphone 100 according to the first embodiment of the present application. Figure 2 In it, the smartphone 100 is physically connected to the Bluetooth speaker 2. The magnetic ring 1 of the Bluetooth speaker 2 is connected to the magnetic ring of the smartphone ( Figure 2Not shown in the figure, the magnetic ring is placed on the back of the mobile phone and is the same shape as the magnetic ring 1). The magnetic ring 1 can adjust the angle along the rotation axis inside the connecting part 11. The Bluetooth speaker 2 has a multi-functional switch button 224, and the multi-functional switch button 224 has functions such as turning on, turning off, Bluetooth connection, pausing, playing, etc. A through hole 223 penetrating up and down is provided in the upper shell of the Bluetooth speaker 2, the sound output direction of the speaker 23 is upward, and the sound outlet of the speaker 23 is located in the through hole 223 of the upper shell. A cushion ring 12 is provided on one side of the mobile phone magnetic ring 1 facing the Bluetooth speaker 2.
[0026] Step S102: Establish a wireless connection between the Bluetooth speaker and the smart phone through the Bluetooth protocol.
[0027] In step S102, establishing a wireless connection between the Bluetooth speaker and the smart phone through the Bluetooth protocol is a key link for realizing multi-mode interaction. First, it is necessary to start the Bluetooth module on the smart phone and the Bluetooth speaker respectively. On the smart phone side, the Bluetooth module is usually activated through the API call built into the operating system. The user can turn on the Bluetooth function in the settings menu or automatically start the Bluetooth search function through an application. At the same time, the Bluetooth module on the Bluetooth speaker also needs to enter the discoverable pairing mode, which is usually achieved by pressing a specific button or switch on the speaker.
[0028] When the Bluetooth speaker enters the pairing mode, it broadcasts its device name and MAC address. After the smart phone receives this information within the search range, it will display a list of available devices on the user interface. The user can select the speaker to send a pairing request. At this time, the smart phone and the Bluetooth speaker will communicate through the device discovery process of the Bluetooth protocol stack, including the exchange of device names, the identification of device types, and the confirmation of connection methods.
[0029] Once the speaker accepts the pairing request from the smart phone, the two sides start the connection process of the Bluetooth protocol. First, the speaker and the smart phone will negotiate the communication parameters of the device through the Link Management Protocol (LMP), such as the Maximum Transmission Unit (MTU), the flow control method, the error detection and correction method, etc. Next, through the Service Discovery Protocol (SDP), the smart phone will obtain the service types supported by the Bluetooth speaker, such as function modules for audio playback, microphone input, etc. This step ensures that the two devices can communicate using the same protocol standard in subsequent interactions.
[0030] After the above protocol negotiation is completed, a secure wireless connection is established between the smartphone and the Bluetooth speaker. To ensure communication security, encryption is usually used during the connection process, such as AES encryption of the Bluetooth channel, to protect the privacy of data transmission. This encryption process starts automatically when the pairing is successful, without the need for manual intervention by the user. After the connection is successfully established, the Bluetooth speaker emits a beep to notify the user that the device has been successfully connected and entered the normal working mode.
[0031] Step S103: Detect the physical adsorption state between the smartphone and the Bluetooth speaker. The physical adsorption state includes a first adsorption state, a second adsorption state, and a third adsorption state. Among them, the first adsorption state means that the smartphone is adsorbed on the Bluetooth speaker, the first magnet ring and the second magnet ring are magnetically adsorbed, and the smartphone forms a video viewing angle of 50 - 70 degrees with the horizontal plane. The second adsorption state means that the smartphone is adsorbed on the Bluetooth speaker, the first magnet ring and the second magnet ring are magnetically adsorbed, and the smartphone forms a video call angle of 0 - 15 degrees with the horizontal plane. The third adsorption state means that the Bluetooth speaker is not adsorbed to the smartphone.
[0032] In step S103, detecting the physical adsorption state between the smartphone and the Bluetooth speaker is a key step to achieve multi-mode interaction. First, by integrating an angle sensor, a pressure sensor, and a magnetic sensor inside the Bluetooth speaker, the adsorption state of the smartphone is monitored in real time. These sensors work together to determine the angle and adsorption situation between the smartphone and the speaker.
[0033] When the smartphone is adsorbed to the Bluetooth speaker through the magnet ring, the angle sensor immediately senses the tilt angle of the smartphone relative to the horizontal plane. If it is detected that the angle formed by the smartphone and the horizontal plane is between 50 - 70 degrees, it is determined as the first adsorption state, that is, the angle in the video viewing mode. In this state, the Bluetooth speaker will receive the feedback signal from the sensor and record the current adsorption state information.
[0034] If the angle sensor detects that the smartphone forms a tilt angle of 0 - 15 degrees with the horizontal plane, it can be determined as the second adsorption state, that is, the video call mode. At this time, the system of the Bluetooth speaker will mark this state as the corresponding physical adsorption state, and at the same time, combine the information of the magnetic sensor to further confirm the adsorption position of the smartphone. The magnetic sensor can accurately sense the adsorption position of the smartphone and ensure stable magnetic adsorption of the two magnet rings at different angles.
[0035] When the smartphone is not physically adsorbed to the Bluetooth speaker, the angle sensor will detect that the Bluetooth speaker is not within the preset angle range, and at the same time, the pressure sensor does not sense the pressure change from the smartphone. In this case, the system determines it as the third adsorption state, that is, the Bluetooth speaker is in an independent placement state.
[0036] To ensure the accuracy of detection, the system will sample and average the data of the sensor multiple times to exclude short-term jitters or misjudgments. During the above detection process, the processor of the Bluetooth speaker will transmit the data collected by the sensor to the central control module in real time. This module is responsible for judging the current physical adsorption state and transmitting this state information to the subsequent interaction mode control module to implement the next interaction operation.
[0037] Step S104: Enter the corresponding interaction mode according to the detected physical adsorption state.
[0038] In step S104, the system automatically switches to the corresponding interaction mode according to the detected physical adsorption state. Specifically, when the physical adsorption state is determined, the control module of the Bluetooth speaker will immediately identify this state and activate the audio and video processing modules corresponding to this state, thereby entering the interaction mode suitable for the current usage scenario.
[0039] When the first adsorption state is detected, that is, the smartphone and the Bluetooth speaker are adsorbed together at an angle of 50 - 70 degrees, the system will automatically enter the video playback mode. In this mode, the Bluetooth speaker will switch the sound effect mode to the theater sound effect mode and simultaneously turn on the stereo sound effect enhancement function to provide a better surround stereo effect. At this time, the system will also dynamically adjust the audio parameters according to the type of the currently played video content (such as movies, TV series, or music videos) to optimize the user's viewing experience.
[0040] When the second adsorption state is detected, that is, the smartphone and the Bluetooth speaker are adsorbed together at an angle of 0 - 15 degrees, the system will switch to the video call mode. In this mode, the microphone sensitivity of the Bluetooth speaker will be automatically increased to ensure that the user's voice signal can be clearly captured during voice input. At the same time, the system will enable the call noise reduction function to reduce the interference of ambient noise and optimize the voice clarity during the call. The system will also adjust the voice processing parameters in real time to cope with different voice environment changes and ensure stable voice output under various call conditions.
[0041] When the third adsorption state is detected, that is, the smartphone is not physically adsorbed to the Bluetooth speaker, the system will enter the music playback mode. In this mode, the sound effect settings of the Bluetooth speaker will be adjusted to the omnidirectional stereo sound effect mode to enhance the sense of space of the sound effect. In addition, the system will automatically optimize the balance of bass and treble to ensure a more rich sound quality performance during music playback. The speaker will also enable the spatial sound effect optimization function to enhance the surround effect of the music through multi-directional sound propagation and increase the user's auditory enjoyment.
[0042] Throughout the process, the control module of the Bluetooth speaker continuously monitors the changes in the current physical adsorption state and rapidly adjusts the audio and video parameters according to the switching of each state. In this way, the system can quickly and accurately switch to the corresponding interaction mode based on different adsorption states, achieving seamless conversion between multiple modes and the best audio-visual effects.
[0043] When the system enters the corresponding interaction mode based on the detected physical adsorption state, specific sound effect adjustments and function optimizations will be implemented according to different adsorption states. When the first adsorption state is detected, that is, when the smartphone and the Bluetooth speaker are adsorbed together at an angle of 50 - 70 degrees, the system will automatically switch to the video playback mode. In this mode, the system first adjusts the Bluetooth speaker to the theater sound effect mode. This sound effect mode is specifically designed to enhance the sound effect of videos, making the sound effect of the low-frequency part more shocking and improving the clarity of the high-frequency part at the same time. At this time, the system will turn on the stereo sound effect enhancement function, significantly enhancing the sense of surround and space of the sound, thereby enhancing the user's audio-visual experience. In addition, the system will dynamically adjust the audio parameters according to the type of video content, such as strengthening the low-frequency impact in action movies and enhancing the clarity of human voices in dialogue segments, so as to optimize the user's viewing experience during video playback.
[0044] When the system detects the second adsorption state, that is, when the smartphone and the Bluetooth speaker are adsorbed at an angle of 0 - 15 degrees, the system will enter the video call mode. In this mode, the system will automatically increase the microphone sensitivity of the Bluetooth speaker, enabling the microphone to capture the user's voice signal more clearly. This adjustment is especially suitable for the voice input requirements of users in a quiet environment or at a relatively large call distance. At the same time, the system will enable the call noise reduction function, reducing the interference of environmental noise on the call quality by analyzing the environmental noise in real time and performing dynamic noise reduction processing. In terms of voice clarity, the system will optimize the audio processing parameters, such as enhancing the voice signal in the middle frequency band and making real-time adjustments according to the user's speaking speed and intonation, to ensure the stability and clarity of the voice output, thereby providing users with a more natural and high-quality video call experience.
[0045] When the system detects the third adsorption state, that is, when the smartphone is not physically adsorbed to the Bluetooth speaker, the system will automatically switch to the music playback mode. In this mode, the system will adjust the Bluetooth speaker to the omnidirectional stereo sound effect mode. This mode can make the sound spread evenly in all directions, forming a stronger sense of space and surround. In addition, the system will optimize the balance of bass and treble according to the user's sound effect preference, making the bass more mellow and the treble clearer. During the playback process, the system will also enable the spatial sound effect optimization function, further enhancing the sense of distribution and layering of the music in space, thereby enhancing the overall auditory effect.
[0046] Through the above mechanism, different physical adsorption states will trigger corresponding interactive mode switches. Each mode includes a series of optimized settings for sound effects and functions to ensure that users can obtain the best audio experience in different scenarios. The system can not only automatically adjust according to the adsorption state, but also dynamically optimize according to environmental changes and user preferences, making the multi-mode interaction more intelligent and personalized.
[0047] When the first adsorption state is detected, that is, when the smartphone is adsorbed to the Bluetooth speaker at an angle of 50 - 70 degrees, the system will enter the video playback mode. To further optimize the user's viewing experience in this mode, the system will dynamically select the most suitable cinema sound effect mode according to the audio characteristics of the video content. First, the system will perform real-time analysis on the audio signal of the video content to extract the characteristic information in the audio. These characteristic information may include human voice conversations, background music, and sound effect details, etc. By analyzing these characteristics, the system can determine the type of the current video, such as a drama mainly featuring human voices, an action movie with rich sound effects, or a musical with a large proportion of music, etc.
[0048] According to the analysis results, the system will automatically select the most suitable cinema sound effect mode. If it is detected that the video content mainly features human voice conversations, the system will switch to the clear dialogue mode. In this mode, the Bluetooth speaker will enhance the output of the human voice frequency, making the dialogue part clearer and more prominent. If the proportion of music in the video content is large, the system will switch to the music enhancement mode. In this mode, the Bluetooth speaker will improve the performance of the low-frequency and high-frequency parts, making the music effect more plump and three-dimensional. For content with a large proportion of sound effects and background sounds (such as action movies or science fiction movies), the system will select the stereo surround mode to enhance the effect of the surround sound, enabling users to obtain a more immersive audio-visual experience.
[0049] In terms of volume adjustment, the system will use the user's historical volume adjustment habit data to set the initial volume value. This initial volume value is set according to the user's volume preference in the past video playback mode, ensuring that users can obtain a suitable volume level without manual adjustment every time they enter this mode. In addition, the system will automatically fine-tune the volume according to the real-time change of the ambient noise. The detection of the ambient noise is usually achieved through the microphone of the smartphone or the ambient noise sensor on the Bluetooth speaker. When the ambient noise increases, the system will automatically increase the volume output to ensure that the sound effect of the video content is not interfered by the background noise; when the ambient noise decreases, the system will appropriately reduce the volume to avoid the impact of excessive volume on the user's auditory comfort.
[0050] Through the above steps, the system can not only intelligently adjust the sound effects and volume according to the video content and user habits, but also respond to the changes in the environment in real time, thus providing users with a continuously optimized viewing experience. This dynamic sound effect and volume adjustment mechanism not only improves the intelligence level of the system, but also significantly enhances the audio-visual enjoyment of users in the video playback mode.
[0051] When the second adsorption state is detected, that is, when the smartphone and the Bluetooth speaker are adsorbed at an angle of 0 - 15 degrees, the system will enter the video call mode. To optimize the audio quality in this mode, the system will first use machine learning algorithms to analyze the frequency of the user's adjustment of the microphone sensitivity during past video calls. This analysis is based on the user's historical audio settings and adjustment behaviors, and extracts the user's preference patterns in different environments and call scenarios. For example, if the user often reduces the microphone sensitivity in a relatively quiet environment, the system will automatically set the initial sensitivity lower in future similar environments. On the contrary, if the user repeatedly increases the microphone sensitivity in a relatively noisy environment, the system will correspondingly automatically increase the initial sensitivity to better adapt to the current call environment.
[0052] During a real-time call, the system will also dynamically adjust the noise reduction intensity according to the changes in the ambient noise. Specifically, the noise sensors on the Bluetooth speaker and the smartphone will continuously monitor the ambient noise level around. If the noise increases, the system will correspondingly enhance the noise reduction function to ensure that the clarity of the voice is not interfered by the external noise. This adjustment is calculated and implemented in real time by the algorithm, and can respond to the changes in the ambient noise within milliseconds to ensure the stability of the call quality. At the same time, when the ambient noise decreases, the system will automatically weaken the noise reduction intensity to avoid over-processing of the voice signal by the excessive noise reduction effect, thus retaining the natural sound quality of the voice.
[0053] In addition, to further improve the clarity of the voice and the call quality, the system will use the multi-microphone array technology of the smartphone and the Bluetooth speaker. This multi-microphone array can not only identify and eliminate background noise through the phase difference and signal strength difference, but also enhance the directivity of the user's voice through the sound source localization technology. In practical applications, the multi-microphone array will first identify the direction of the user's voice, and then filter the background noise in other directions, so that the user's voice becomes clearer and more prominent during the call. This sound source localization technology can be adjusted in real time in a dynamic environment, enabling the system to still maintain high-quality voice output when the user moves or the environment changes.
[0054] Through the above series of optimization steps, the system can not only automatically adjust the microphone sensitivity according to the user's habits in the video call mode, but also respond to changes in ambient noise in real time, and improve the directivity and clarity of the voice through multiple microphone arrays and sound source localization technology. In this way, users can obtain a clearer and more natural voice effect during video calls, and at the same time, the system can also stably provide high-quality audio performance in different environments.
[0055] When the third adsorption state is detected, that is, when the smartphone is not physically adsorbed to the Bluetooth speaker, the system will switch to the music playback mode. To optimize the sound effect performance in this mode, the system will first select a suitable equalizer preset mode by analyzing the user's sound effect preferences under different music types. Specifically, the system will analyze the user's historical playback records to identify the sound effect settings selected by the user in different types of music such as pop music, classical music, jazz, and electronic music. This analysis involves not only the historical data of volume and frequency adjustment, but also the preferences for equalizer settings, such as bass boost, clear vocal mode, or treble boost. Based on the analysis results, the system will automatically select the equalizer preset mode that best matches the user's current music type preference to ensure that the played music matches the user's auditory needs.
[0056] During music playback, the system will also dynamically adjust the spatial sound effect parameters according to the rhythm changes of the song. The spatial sound effect parameters are adjusted by real-time analysis of the music signal, paying special attention to features such as the rhythm, beats, and intensity changes of the song. For example, when the system detects an increase in music rhythm or an enhancement of low-frequency components, it will enhance the surround effect of the spatial sound effect, making the sound have a greater sense of extension in space, thereby enhancing the three-dimensional and immersive feeling of hearing; while when the music rhythm slows down or turns soft, the system will appropriately contract the spatial sound effect, making the sound concentrate in the direct front of the listener, creating a more delicate and private listening experience. This dynamic adjustment mechanism can respond to the changes in the song in real time, synchronize the sound effect with the rhythm of the music, and thus improve the sense of presence and music expressiveness of hearing.
[0057] In addition, the system will also set the initial sound effect parameters by analyzing the user's historical bass and treble adjustment data. For example, if the user frequently boosts the bass during past playbacks, then the system will automatically set a higher bass parameter at the start of a new round of music playback, and vice versa. As the music plays, the system will monitor the current audio output in real time and make further fine-tuning according to the user's real-time adjustment behavior. If the user adjusts the bass or treble parameter again during playback, the system will quickly respond to this change and update the sound effect parameters through an algorithm to ensure the continuous optimization of the auditory experience.
[0058] Through the above steps, the system can not only automatically select the appropriate equalizer mode according to the user's sound effect preferences in the music playback mode, but also adjust the sound effect parameters in real time to match the rhythm of the music and the user's immediate needs. This highly flexible and personalized sound effect optimization mechanism not only enhances the user's auditory experience, but also improves the overall sound quality performance of music playback, enabling users to enjoy the best sound effect experience in different types of music.
[0059] Step S105: Use machine learning algorithms to analyze the usage data collected from the user in different interaction modes, and generate personalized parameter solutions for different physical adsorption states.
[0060] In step S105, the system analyzes the usage data of the user in different interaction modes through machine learning algorithms to generate personalized parameter solutions for different physical adsorption states. First, the system obtains relevant usage data from the sensors, application interfaces, and user inputs of the Bluetooth speaker and the smartphone. This data may include the user's volume adjustment habits, the selection frequency of sound effect modes, the adjustment records of microphone sensitivity, the adjustment history of the support angle, as well as environmental characteristics such as ambient noise level and ambient light intensity. This step can be executed in the processing unit of the Bluetooth speaker.
[0061] Once these data are collected, the system will preprocess them, including data standardization, denoising, and normalization. The standardization process ensures that the numerical ranges of different data are the same, the denoising process is used to filter out outliers and invalid data, and normalization scales the data to the same range for subsequent processing by machine learning algorithms.
[0062] Next, the preprocessed data will be input into a pre-trained neural network model for analysis.
[0063] After completing the above feature analysis, the output layer of the neural network will generate personalized parameter solutions for different physical adsorption states according to the analysis results. Specifically, for the first adsorption state, the personalized parameter solution may include a specific enhancement level of the video sound effect mode and the optimal support angle; for the second adsorption state, the solution may focus on optimizing voice clarity and adjusting noise reduction parameters; while for the third adsorption state, it may focus on adjusting the equalizer settings and enhancing spatial sound effects. Each personalized parameter solution contains a combination of multiple parameters, ensuring the optimal adjustment of the user experience in different interaction modes.
[0064] Through the adaptive characteristics of the machine learning algorithm, the entire process can continuously optimize according to the user's historical usage data and environmental changes. This means that during the user's usage, the system will continuously update the personalized parameter solution based on new data, making it closer to the user's actual needs and preferences each time it is used. In this way, through in-depth analysis and feature extraction of the user's behavior data, the system can not only provide highly personalized settings for different physical adsorption states, but also achieve smooth transitions during multi-mode switching, ensuring continuous optimization of the user experience.
[0065] Furthermore, the use of the machine learning algorithm to analyze the usage data of the user collected in different interaction modes to generate a personalized parameter solution for different physical adsorption states includes:
[0066] Using a pre-trained neural network model to process the usage data of the user collected in different interaction modes to generate a personalized parameter solution for different physical adsorption states; wherein, the neural network model includes an input encoding module, a multi-dimensional attention feature extraction module, an adaptive multi-state feature fusion module, and a multi-task optimization output module;
[0067] Among them, the input encoding module is used to receive various types of input features, including user behavior features, physical adsorption state features, and environmental features; the input encoding module uses a convolutional encoding network to perform preliminary extraction and non-linear transformation on the input features, mapping the low-dimensional input features to a high-dimensional space, thereby obtaining an encoded feature vector;
[0068] The multi-dimensional attention feature extraction module is used to receive the encoded feature vector provided by the input encoding module; the multi-dimensional attention feature extraction module is implemented by a recurrent neural network based on a multi-dimensional attention mechanism, and is used to dynamically adjust the weights of features for different adsorption states and interaction modes to obtain a feature vector weighted by multi-dimensional attention;
[0069] The adaptive multi-state feature fusion module is used to receive the feature vector provided by the multi-dimensional attention feature extraction module; the adaptive multi-state feature fusion module uses a cross-feature combination strategy to cross-multiply and weighted-sum the features from different modes to generate a richer multi-mode feature representation; through a dynamic weighting function, adjust the fusion weight of the features according to the current physical adsorption state and user preferences to obtain a fused feature vector;
[0070] The multi-task optimization output module is used to receive the fused feature vector provided by the adaptive multi-state feature fusion module; the multi-task optimization output module is designed based on a multi-task learning framework, and decomposes the final personalized parameter solution into multiple task outputs, and the multiple task outputs include:
[0071] Task 1: Parameter output for video playback mode, including sound enhancement and stereo enhancement;
[0072] Task 2: Parameter output for video call mode, including microphone sensitivity, noise reduction intensity, and speech clarity parameters;
[0073] Task 3: Parameter output for music playback mode, including equalizer adjustment, spatial sound effect, and bass enhancement parameters.
[0074] When implementing this multi - mode interaction method, using machine learning algorithms to analyze the usage data of users in different interaction modes is the core step. First, the system collects various feature data of users in different physical adsorption states and interaction modes. These data include user behavior characteristics (such as volume adjustment, sound effect selection, microphone adjustment, etc.), physical adsorption state characteristics (such as angle change, adsorption position), and environmental characteristics (such as environmental noise, light intensity, etc.). The diversity of these data provides a sufficient basis for the subsequent generation of personalized parameter solutions.
[0075] After the data collection is completed, these multiple types of input features are received by the encoding module. This module uses a convolutional encoding network to perform preliminary extraction and non - linear transformation on the input features, aiming to map the low - dimensional input features to a high - dimensional space. The convolutional encoding network extracts features of different dimensions layer by layer, ensuring that the non - linear relationships between features are fully captured and represented. In this way, the system can extract key features from complex user behavior patterns, thus obtaining an encoded feature vector. This encoded feature vector has higher expressiveness and generalization ability, laying a foundation for subsequent in - depth analysis and processing.
[0076] Next, the multi - dimensional attention feature extraction module further processes the encoded feature vector. This module is implemented using a recurrent neural network based on a multi - dimensional attention mechanism. The purpose of this process is to dynamically adjust the weights of the input features according to different physical adsorption states and interaction modes. The multi - dimensional attention mechanism can finely adjust the weights of features, thereby enhancing the key features in a specific interaction mode. For example, in the video playback mode, the attention mechanism will highlight the sound - related features, while in the video call mode, it will assign higher weights to the features related to microphone sensitivity and speech clarity. Through this dynamic weighting, the system can generate a feature vector weighted by multi - dimensional attention, which can better adapt to the current physical adsorption state and interaction mode.
[0077] Subsequently, the adaptive multi-state feature fusion module fuses the feature vectors provided by the multi-dimensional attention feature extraction module. This module adopts a cross-feature combination strategy, which cross-multiplies and weighted-sums the features from different modes to generate a richer multi-mode feature representation. Through this cross-feature combination, the system can capture the feature interaction relationships in different modes, and through a dynamic weighting function, adjust the fusion weights of the features according to the current physical adsorption state and user preferences, thereby obtaining the fused feature vector. This feature fusion strategy ensures the adaptability and personalized performance of the system under different adsorption states.
[0078] Finally, the multi-task optimization output module decomposes the fused feature vector into multiple task outputs. These outputs are designed as personalized parameter schemes for different interaction modes. For the video playback mode, the outputs include parameter settings for sound enhancement and stereo enhancement; for the video call mode, the outputs contain optimized parameters for microphone sensitivity, noise reduction intensity, and speech clarity; and in the music playback mode, the outputs include equalizer adjustment, spatial sound effects, and bass enhancement settings. Each task output is optimized for a specific physical adsorption state and user needs to ensure the best audio-visual experience in different interaction modes.
[0079] Through the above implementation steps, the present invention can generate highly personalized parameter schemes based on the usage data of users in different modes. This not only improves the adaptability and intelligence level of the system, but also significantly optimizes the overall experience of users in multi-mode interactions.
[0080] In this embodiment, the design and implementation of the multi-dimensional attention feature extraction module are the key parts, and its purpose is to calculate the attention weights of the input features at different time steps and use these weights to extract and adjust the temporal features. The whole process depends on a series of calculation steps, involving the specific application of multiple mathematical formulas and parameters.
[0081] First, the system calculates the multi-dimensional attention weights of each input feature at time step using Formula 1. The specific form of Formula 1 is as follows:
[0082] ;
[0083] In Formula 1, represents the attention weight at time step and is used to reflect the importance of the current input feature in the multi-dimensional space. The attention weight is achieved by performing a non-linear transformation on each input feature and normalizing the result. Specifically:
[0084] is a multi-dimensional attention weight matrix, which is a trainable matrix used to perform a linear transformation on the input feature vector. The elements in this matrix are automatically adjusted through the neural network training process to optimally extract different dimensions of features.
[0085] is the encoded feature vector generated by the input encoding module. It represents the feature result after being processed by the convolutional encoding network and is a high-dimensional representation of user behavior features, physical adsorption state features, and environmental features. This feature vector is extracted through multiple layers of the convolutional network, retaining the non-linear relationships between different dimensions.
[0086] is the standard Sigmoid activation function, which is defined as . In this formula, the Sigmoid function is used to limit the input result between , thereby converting the feature value into the basis of the attention weight and ensuring that the numerical range of the attention weight is suitable for subsequent normalization processing.
[0087] is the bias term, which is a parameter used to adjust the initial value of the input feature. The main role of the bias term is to balance the initial weights of different features, thereby ensuring that the model has better robustness and adaptability during the feature processing process.
[0088] is the encoded feature vector generated by the input encoding module at time step . It represents the input feature at the -th time step in the time series, similar to the feature at time step . Its role is to provide a comparison reference during the attention weight normalization process to ensure that the sum of the attention scores for all time steps is 1. Similar. Its role is to provide a comparison reference during the attention weight normalization process to ensure that the sum of the attention scores for all time steps is 1.
[0089] 1 is the total number of time steps, that is, the length of the time series considered when calculating the attention weight. It represents the range of the time dimension of the input feature and is used to represent the end point of the time series in the formula. This parameter determines the range of normalization, thereby ensuring that the calculation of the attention weight covers the entire time series.
[0090] Once the attention weights for all time steps are calculated, they are used to extract temporal features. For this purpose, the system extracts temporal features through Equation 2:
[0091] ;
[0092] In Equation 2, Represents the time step of the temporal feature vector, which contains the evolution of multi-dimensional features in the time dimension. The specific explanations are as follows:
[0093] is the element-wise Hadamard product, which is used to multiply the attention weights with the input features element by element, thereby adjusting the weight of each feature. This operation ensures that the attention weights play a role in temporal feature extraction.
[0094] is the temporal feature state of the previous time step.
[0095] is the encoded feature vector generated by the input encoding module at time step .
[0096] is the input weight matrix of the recurrent neural network, which is used to perform a linear transformation on the features adjusted by the attention weights. The weight values in this matrix are optimized through training and are used to extract features at the time step.
[0097] is the hidden layer weight matrix of the recurrent unit, which combines the feature state of the previous time step with the current feature to capture the long-term and short-term dependencies in the temporal features. This process enables the recurrent network to model the feature changes in the time dimension.
[0098] is the hyperbolic tangent activation function, which is used to limit the feature values within the range of , thereby ensuring the non-linear expression ability of the model.
[0099] is the bias term of the recurrent neural network, which is used to balance the initial values of the input features during the recurrent calculation process.
[0100] After completing the temporal feature extraction, the system calculates the weighted feature vector through Equation 3:
[0101] ;
[0102] In Equation 3, is the feature vector weighted by the multi-dimensional attention mechanism and is used for the final feature fusion and output. The specific explanations are as follows:
[0103] Represents the accumulation of weighted features over all time steps. This accumulation process ensures the application of attention weights in the time dimension, enabling the weighted feature vectors to reflect the importance of multi-dimensional features. Represents the total number of time steps.
[0104] Is a periodic adjustment term used to introduce a modulating factor with periodic variations. This term remains fixed during the calculation and is mainly used to adjust the periodic fluctuations of features in the time dimension.
[0105] Is a periodic function that is used to simulate the natural periodic variations of features in the time dimension. By introducing this term, the system can capture the periodic characteristics of features in the time dimension, thereby improving the response ability to periodic signals.
[0106] In this embodiment, the role of the adaptive multi-state feature fusion module is to achieve the fusion of multi-mode features through dynamic cross-feature combination and adaptive weight adjustment. This process uses Formula 4 for complex combination of features, aiming to generate a more accurate and adaptable final fused feature vector. The specific implementation steps are as follows.
[0107] During the feature fusion process, the system first obtains a series of weighted feature vectors from the multi-dimensional attention feature extraction module. Each feature vector represents the weighted result of features in the current mode and contains important information extracted from multi-mode interactions. To enable these features to be optimally combined under different physical adsorption states, the adaptive multi-state feature fusion module adopts a strategy of dynamic cross-feature combination to further fuse and optimize these features.
[0108] Formula 4 is used to implement this fusion strategy, and its specific form is:
[0109] ;
[0110] In this formula:
[0111] Represents the finally fused feature vector, which is the result of comprehensively processing multi-dimensional features in different modes. This feature vector will be used to generate personalized parameter schemes for different physical adsorption states. Its calculation is achieved through cross-combination and adaptive adjustment of all weighted features.
[0112] Is the th weighted feature vector output by the multi-dimensional attention feature extraction module. Each Represents the feature weight distribution in a specific mode, which is generated by weighting features through an attention mechanism. These feature vectors serve as the basis of the input data during the feature fusion process.
[0113] Is the element-wise Hadamard product, used to cross-combine different features element by element. Through this element-wise multiplication operation, the interaction relationships between different features can be retained and enhanced during the fusion process, strengthening the combination effect of features in the non-linear space.
[0114] Is the part that performs non-linear transformation on the feature vector, where Is the hyperbolic tangent activation function, which limits the input value within The range. This activation function is used to improve the non-linear expression ability of features, while Is the adjustment coefficient, used to control the non-linear amplitude of features. Specifically, Is a trainable parameter, whose initial value is determined by the pre-training of the model and is dynamically adjusted according to the performance of features during the training process to ensure adaptability to features.
[0115] Is the -th element of the fusion weight matrix. Each weight element Is used to control the proportion of different features in the final fused feature. The initial weight values are determined by the model through learning on a large-scale dataset during the pre-training stage. During the operation of the model, these weights are adaptively updated according to the current physical adsorption state and user behavior to achieve the optimal feature combination in different interaction modes.
[0116] Is a periodic adjustment term, used to introduce periodic changes in the time dimension. Here, Is the periodic adjustment coefficient, which is an adjustable parameter used to control the amplitude of the periodic changes. Specifically, Is automatically optimized during the model training stage to ensure that the periodic adjustment can match the natural fluctuations of features. The periodic function Further enhances the dynamic adaptability of features by simulating the periodic characteristics of feature changes over time, where Is the current timestamp, Is a specific adjustment period. The adjustment period Is set according to the time characteristics of a specific interaction mode to ensure stable periodic changes of features in different time periods.
[0117] In this embodiment, the function and implementation of the input encoding module are the key parts of the entire multi-modal interaction system. Its main purpose is to receive various input data from user behaviors, physical adsorption states, and environmental characteristics, and perform necessary preprocessing and feature extraction on this data, so as to perform more accurate feature weighting and temporal processing in subsequent steps.
[0118] First, the input encoding module receives three types of input features: user behavior features, physical adsorption state features, and environmental features. User behavior features include operation records of the user in different interaction modes, such as volume adjustment, sound effect selection, and microphone sensitivity setting. The physical adsorption state features describe the physical connection method between the smartphone and the Bluetooth speaker, such as the adsorption angle and position. Environmental features cover external conditions such as environmental noise, light intensity, and temperature. These input features come from a variety of sensors and user input interfaces, so they may have different value ranges and scales in their original state.
[0119] To ensure the consistency of these input data during model processing, the input encoding module first performs standardization and normalization on the received feature data. The purpose of standardization is to make different features consistent in distribution by subtracting the mean and dividing by the standard deviation, usually adjusting their distribution to a standard normal distribution with a mean of 0 and a standard deviation of 1. Normalization processing is to scale the values of all features to the same range (such as 0 to 1) to avoid biases caused by features with overly large or small values to the model.
[0120] After completing the preprocessing, the standardized and normalized features will be input into the convolutional encoding network for feature extraction. The design of the convolutional encoding network aims to map the input low-dimensional features to a high-dimensional feature space, thereby enhancing the ability to capture the non-linear relationships between multi-dimensional features. The convolutional encoding network extracts the complex relationships and patterns of the input features layer by layer through the stacking of multiple convolutional layers and non-linear activations.
[0121] Specifically, each layer of the convolutional encoding network weights the input features through a convolutional kernel and passes the result to the next layer for further non-linear mapping. Through this layer-by-layer convolution and activation, the input features are gradually transformed into high-dimensional feature vectors. The high-dimensional feature vectors not only contain the basic information of the original input features, but also introduce the mutual relationships and complex interactions between the features, which helps to perform more accurate feature weighting and temporal processing in the subsequent multi-dimensional attention feature extraction module.
[0122] Finally, the extracted high-dimensional feature vectors will be used as the output and passed to the multi-dimensional attention feature extraction module. This output not only provides richer feature information for the multi-dimensional attention mechanism, but also lays the foundation for subsequent feature weighting and temporal processing.
[0123] In this embodiment, the main function of the multi-task optimization output module is to decompose multi-mode interaction tasks into multiple subtasks, and dynamically adjust the priorities and related settings of output parameters according to personalized requirements in different modes. Through the decomposition of interaction tasks and parameter optimization, this module provides more accurate and personalized support for the multi-mode interaction between the Bluetooth speaker and the smartphone.
[0124] First of all, during the multi-mode interaction process, the system divides the overall task into multiple subtasks according to the user's operation behavior and physical adsorption state. These subtasks include sound enhancement, noise reduction intensity adjustment, equalizer settings, etc. This decomposition method enables each subtask to independently optimize parameters, ensuring the best user experience in different interaction modes. For example, when the system is in the video playback mode, the priority of the sound enhancement subtask may be higher than that of the noise reduction function, while in the video call mode, the noise reduction intensity and speech clarity will be optimized first. This dynamic adjustment of priorities ensures that the system can quickly adapt to changes in user needs in different modes.
[0125] When generating personalized parameter solutions for each subtask, the system will refer to the user's historical preferences and current environmental conditions. For example, the output threshold of sound enhancement is set according to the user's sound preferences when using the video playback mode in the past. If the user has adjusted the bass enhancement multiple times in the historical record, the system will automatically set a higher bass output threshold; if the user has reduced the volume multiple times in a relatively quiet environment, the system will correspondingly lower the sound enhancement threshold when the current environmental noise is low. Similarly, the output parameters of the noise reduction intensity are set according to the real-time changes in environmental noise and the user's historical preferences in similar environments. For example, in a noisy environment, the threshold of the noise reduction intensity will be set higher to ensure that the speech clarity is not disturbed. The output of the equalizer adjustment is personalized according to the user's adjustment preferences for different music types, so that the sound performance in the music playback mode can be consistent with the user's preferences.
[0126] Through the multi-task learning framework, the system can update the output weights of the model in real time during operation. Specifically, the multi-task learning framework ensures that the output parameters of each subtask can achieve the optimal configuration when switching between multiple modes by simultaneously optimizing the loss functions of multiple subtasks. This real-time update mechanism enables the model to maintain a high degree of consistency in different modes and at the same time have the ability to quickly respond to changes in user preferences. In actual operation, when the user's behavior characteristics and environmental conditions change, the model will automatically adjust the output weights, so that parameters such as sound enhancement, noise reduction intensity, and equalizer adjustment can quickly adapt to the new usage scenario. This real-time update not only improves the system's adaptability but also ensures that the personalized performance in different modes always meets the user's needs.
[0127] Through the above steps, the multi-task optimization output module can decompose complex tasks in multi-modal interactions into multiple sub-tasks and dynamically adjust the output parameters of each sub-task according to the user's historical preferences and environmental conditions. At the same time, through the real-time update mechanism of the multi-task learning framework, the system can continuously optimize various parameter settings in different modes, thereby achieving a highly personalized and adaptive user experience.
[0128] The following is a reference implementation of the neural network model:
[0129] import torch
[0130] import torch.nn as nn
[0131] import torch.nn.functional as F
[0132] import numpy as np
[0133] # Define the input encoding module: receive various types of input features, including user behavior, physical adsorption state, and environmental features
[0134] class InputEncodingModule(nn.Module):
[0135] def __init__(self, input_dim, conv_out_dim):
[0136] super(InputEncodingModule, self).__init__()
[0137] # Convolutional encoding network: perform preliminary extraction and non-linear transformation on input features
[0138] self.conv1 = nn.Conv1d(in_channels=input_dim, out_channels=conv_out_dim, kernel_size=3, padding=1)
[0139] self.conv2 = nn.Conv1d(in_channels=conv_out_dim, out_channels=conv_out_dim, kernel_size=3, padding=1)
[0140] self.batch_norm = nn.BatchNorm1d(conv_out_dim)
[0141] def forward(self, x):
[0142] # Data preprocessing: standardization and normalization
[0143] x = (x - x.mean(dim=0)) / (x.std(dim=0) + 1e-5)
[0144] # Convolution layer 1: Feature extraction from input features
[0145] x = F.relu(self.conv1(x))
[0146] # Convolution layer 2: Further feature extraction
[0147] x = F.relu(self.conv2(x))
[0148] # Batch normalization: Ensure that the output features are in the same numerical range
[0149] x = self.batch_norm(x)
[0150] return x
[0151] # Define the multi-dimensional attention feature extraction module
[0152] class MultiDimAttentionModule(nn.Module):
[0153] def __init__(self, att_dim):
[0154] super(MultiDimAttentionModule, self).__init__()
[0155] self.W_a = nn.Parameter(torch.randn(att_dim, att_dim)) # Multi-dimensional attention weight matrix
[0156] self.b_a = nn.Parameter(torch.zeros(att_dim)) # Bias term
[0157] def forward(self, X_enc):
[0158] # Calculate multi-dimensional attention weights
[0159] T1 = X_enc.size(0) # Total number of time steps
[0160] A_t = []
[0161] for t in range(T1):
[0162] score = torch.sigmoid(X_enc[t] @ self.W_a + self.b_a)
[0163] exp_score = torch.exp(score)
[0164] A_t.append(exp_score)
[0165] A_t = torch.stack(A_t)
[0166] A_t = A_t / A_t.sum(dim=0, keepdim=True) # Normalization
[0167] H_t = []
[0168] # Extract temporal features
[0169] H_prev = torch.zeros_like(X_enc[0]) # Initialize H_(t-1)
[0170] W_h = nn.Parameter(torch.randn(X_enc.size(1), X_enc.size(1)))
[0171] U_h = nn.Parameter(torch.randn(X_enc.size(1), X_enc.size(1)))
[0172] b_h = nn.Parameter(torch.zeros(X_enc.size(1)))
[0173] for t in range(T1):
[0174] H_t_curr = torch.tanh(W_h @ (A_t[t] X_enc[t]) + U_h @H_prev + b_h)
[0175] H_t.append(H_t_curr)
[0176] H_prev = H_t_curr
[0177] H_t = torch.stack(H_t)
[0178] # Weighted feature vector
[0179] T = T1 # Set the adjustment period
[0180] X_att = sum(A_t[t] H_t[t] for t in range(T)) + 0.1 torch.cos(2 np.pi torch.arange(T) / T).sum()
[0181] return X_att
[0182] # Define the adaptive multi-state feature fusion module
[0183] class AdaptiveFeatureFusionModule(nn.Module):
[0184] def __init__(self, feature_dim, N):
[0185] super(AdaptiveFeatureFusionModule, self).__init__()
[0186] self.W_fuse = nn.Parameter(torch.randn(N, feature_dim)) # Initial fusion weight matrix
[0187] self.gamma = nn.Parameter(torch.ones(N)) # Nonlinear adjustment coefficient
[0188] self.beta = nn.Parameter(torch.randn(N)) # Periodic adjustment coefficient
[0189] def forward(self, X_att):
[0190] X_fuse = 0 # Initialize the final fused feature vector
[0191] t = torch.arange(0, X_att.size(0), dtype=torch.float) # Current timestamp
[0192] T = X_att.size(0) # Adjustment period
[0193] for i in range(len(self.W_fuse)):
[0194] X_fuse += (X_att[i] torch.tanh(self.gamma[i] X_att[i])) (
[0195] self.W_fuse[i] + self.beta[i] torch.sin(t / T))
[0196] return X_fuse
[0197] # Define the multi-task optimization output module
[0198] class MultiTaskOptimizationModule(nn.Module):
[0199] def __init__(self, feature_dim):
[0200] super(MultiTaskOptimizationModule, self).__init__()
[0201] self.fc_video = nn.Linear(feature_dim, 2) # Video playback parameters: sound enhancement, stereo enhancement
[0202] self.fc_call = nn.Linear(feature_dim, 3) # Video call parameters: microphone sensitivity, noise reduction intensity, speech clarity
[0203] self.fc_music = nn.Linear(feature_dim, 3) # Music playback parameters: equalizer, spatial sound effect, bass enhancement
[0204] def forward(self, X_fuse):
[0205] # Video playback mode parameter output
[0206] video_params = self.fc_video(X_fuse)
[0207] # Video call mode parameter output
[0208] call_params = self.fc_call(X_fuse)
[0209] # Music playback mode parameter output
[0210] music_params = self.fc_music(X_fuse)
[0211] return video_params, call_params, music_params
[0212] # Define the complete neural network model
[0213] class MultiModeInteractionModel(nn.Module):
[0214] def __init__(self, input_dim, conv_out_dim, att_dim, N):
[0215] super(MultiModeInteractionModel, self).__init__()
[0216] self.input_encoder = InputEncodingModule(input_dim, conv_out_dim)
[0217] self.attention_module = MultiDimAttentionModule(att_dim)
[0218] self.fusion_module = AdaptiveFeatureFusionModule(att_dim, N)
[0219] self.output_module = MultiTaskOptimizationModule(att_dim)
[0220] def forward(self, x):
[0221] # Input encoding
[0222] X_enc = self.input_encoder(x)
[0223] # Multi-dimensional attention feature extraction
[0224] X_att = self.attention_module(X_enc)
[0225] # Adaptive multi-state feature fusion
[0226] X_fuse = self.fusion_module(X_att)
[0227] # Multi-task optimization output
[0228] video_params, call_params, music_params = self.output_module(X_fuse)
[0229] return video_params, call_params, music_params
[0230] To train the above neural network model, an effective dataset needs to be constructed and multiple steps need to be executed to optimize the weights of the model. The following are the brief training steps:
[0231] First, a multi-modal dataset containing user behavior features, physical adsorption state features, and environmental features needs to be prepared. This data should include historical behavior records under various interaction modes and mark the corresponding target parameters (such as sound effect enhancement, noise reduction intensity, equalizer settings, etc.). The data should be preprocessed, including standardization and normalization, to ensure that the numerical ranges of the input features are consistent.
[0232] When initializing the model, appropriate hyperparameters are selected, such as learning rate, batch size, number of training epochs, etc. The model is instantiated, and an optimizer (such as Adam or SGD) is used to adjust the weights of the model, while an appropriate loss function (such as mean squared error) is selected to evaluate the prediction effect of the model. The calculation of the loss function can set weights separately for different task outputs to ensure the effect of multi-task learning.
[0233] During the training process, the input data is fed into the model batch by batch for forward propagation, and the loss of each batch is calculated. The loss value is based on the error between the parameter outputs of different modes such as video playback, video call, and music playback and the true labels. According to the calculated total loss, the weights of the model are updated through the backpropagation algorithm, enabling the model to gradually learn to generate accurate personalized parameters in different modes.
[0234] After training is completed, the performance of the model is tested on an unseen dataset to ensure that the model can still output high-quality personalized parameter solutions under various physical adsorption states and environmental conditions. The weights of the model can be further fine-tuned according to the test results to achieve the best effect.
[0235] Step S106: When the corresponding physical adsorption state is detected, apply the corresponding personalized parameter solution to optimize the user experience.
[0236] In step S106, after the system detects the corresponding physical adsorption state, it will immediately call and apply the previously generated personalized parameter solution to optimize the user experience. The key to this step lies in how to quickly and accurately apply the generated personalized parameter solution to the control systems of the Bluetooth speaker and the smartphone.
[0237] When the physical adsorption state is determined, the central processing module of the system will automatically retrieve the personalized parameter solution corresponding to this state. This solution contains a variety of parameter settings, which involve sound effect modes, microphone sensitivity, support angle, noise reduction intensity, equalizer adjustment, etc. When entering different interaction modes, these parameters will be assigned to the corresponding functional modules to ensure that the device is adjusted according to the current physical state. For example, in the first adsorption state (video viewing mode), the system will adjust the sound effect mode of the Bluetooth speaker to the theater sound effect mode and turn on the stereo sound enhancement function according to the personalized solution. At the same time, the screen support angle of the smartphone will also be fine-tuned according to the recommended value in the solution to make it more compatible with the user's viewing angle.
[0238] In the second adsorption state (video call mode), the system will immediately increase the microphone sensitivity of the Bluetooth speaker and enable the noise reduction function by calling the parameter solution. At this time, the noise reduction intensity of the speaker will be adjusted in real time according to the user's preference and the ambient noise level. To ensure that the voice clarity reaches the best state, the system will also optimize the input voice through the real-time voice processing module to further improve the voice quality of the call.
[0239] In the third adsorption state (music playback mode), the system adjusts the equalizer settings of the Bluetooth speaker according to the personalized scheme to enhance the performance of bass and treble. At the same time, the spatial sound effect function is activated to make the sound effect have a better sense of surround. The parameter adjustment in this mode focuses on the overall auditory effect of the music. Therefore, the omnidirectional stereo sound effect of the speaker will be preferentially activated to increase the sense of space and layering of the sound.
[0240] The second embodiment of the present application provides an electronic device, and the electronic device includes:
[0241] A processor;
[0242] A memory for storing a program, which when read and executed by the processor, executes the multi-mode interaction method between the Bluetooth speaker and the smart phone provided in the first embodiment of the present application.
[0243] The third embodiment of the present application provides a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, it executes the multi-mode interaction method between the Bluetooth speaker and the smart phone provided in the first embodiment of the present application.
[0244] Although the present application is disclosed above with preferred embodiments, it is not used to limit the present application. Any person skilled in the art can make possible changes and modifications without departing from the spirit and scope of the present application. Therefore, the protection scope of the present application should be subject to the scope defined by the claims of the present application.
Claims
1. A multi-mode interaction method between a Bluetooth speaker and a smart phone, characterized in that, Including: A first magnet ring capable of adjusting the angle is provided on the Bluetooth speaker, and the first magnet ring is used for magnetic adsorption with a second magnet ring on the smartphone, so as to realize the physical connection between the Bluetooth speaker and the smartphone; Establish a wireless connection between the Bluetooth speaker and the smartphone through the Bluetooth protocol; Detect the physical adsorption state between the smartphone and the Bluetooth speaker, and the physical adsorption state includes a first adsorption state, a second adsorption state, and a third adsorption state; wherein, the first adsorption state means that the smartphone is adsorbed on the Bluetooth speaker, the first magnet ring and the second magnet ring are magnetically adsorbed, and the smartphone forms a video viewing angle of 50-70 degrees with the horizontal plane; the second adsorption state means that the smartphone is adsorbed on the Bluetooth speaker, the first magnet ring and the second magnet ring are magnetically adsorbed, and the smartphone forms a video call angle of 0-15 degrees with the horizontal plane; the third adsorption state means that the Bluetooth speaker is not adsorbed with the smartphone; Enter the corresponding interaction mode according to the detected physical adsorption state; Use machine learning algorithms to analyze the usage data collected by users in different interaction modes, and generate personalized parameter solutions for different physical adsorption states; When the corresponding physical adsorption state is detected, apply the corresponding personalized parameter solution to optimize the user experience.
2. The multi-mode interaction method between the Bluetooth speaker and the smart phone according to claim 1, characterized in that The entering the corresponding interaction mode according to the detected physical adsorption state includes: When the first adsorption state is detected, enter the video playback mode, adjust the Bluetooth speaker to the cinema sound effect mode; turn on the stereo sound effect enhancement function; dynamically adjust the audio parameters according to the type of video content to optimize the viewing experience; When the second adsorption state is detected, enter the video call mode, increase the microphone sensitivity of the Bluetooth speaker; enable the call noise reduction function to reduce environmental noise interference; optimize the voice clarity parameters to enhance the call quality; When the third adsorption state is detected, enter the music playback mode, adjust the Bluetooth speaker to the omnidirectional stereo sound effect mode; optimize the balance of bass and treble to improve the sound quality; enable the spatial sound effect optimization function to enhance the music playback effect.
3. The multi-mode interaction method between a Bluetooth speaker and a smart phone according to claim 2, wherein, When the first adsorption state is detected, the user experience of the video playback mode is further optimized through the following steps: According to the audio characteristics of the video content, select different cinema sound effect modes, including the dialogue clarity mode, the music enhancement mode, and the stereo surround mode; Use the user's volume adjustment habit data to automatically set the volume startup value and perform real-time fine-tuning according to the ambient noise level.
4. The multi-mode interaction method between a Bluetooth speaker and a smart phone according to claim 2, wherein When the second adsorption state is detected, optimize the audio quality of the video call mode, specifically including: Analyze the adjustment frequency of the microphone sensitivity by the user during the video call through machine learning algorithms, and automatically optimize the initial sensitivity setting of the microphone; During a real-time call, dynamically adjust the noise reduction intensity according to the change of ambient noise to ensure that the voice clarity is not interfered by ambient noise; Use the multiple microphone arrays of the smartphone and the Bluetooth speaker to identify and eliminate background noise, and enhance the directivity of the user's voice through sound source localization technology.
5. The multi-mode interaction method between a Bluetooth speaker and a smart phone according to claim 2, characterized in that, When the third adsorption state is detected, optimize the sound effect performance of the music playback mode, specifically including: Select the preset mode of the equalizer by analyzing the user's sound effect preferences under different music genres; During music playback, dynamically adjust the spatial sound effect parameters according to the changes in the song rhythm, so that the spatial sense of the sound effect matches the rhythm changes of the music; Automatically set the initial sound effect parameters based on the user's historical bass and treble adjustment data, and adjust the balance of bass and treble in real time during playback to enhance the user's auditory experience.
6. The multi-mode interaction method between the Bluetooth speaker and the smart phone according to claim 1, characterized in that, The use of machine learning algorithms to analyze the collected usage data of users in different interaction modes to generate personalized parameter solutions for different physical adsorption states, including: Use a pre-trained neural network model to process the collected usage data of users in different interaction modes to generate personalized parameter solutions for different physical adsorption states; wherein, the neural network model includes an input encoding module, a multi-dimensional attention feature extraction module, an adaptive multi-state feature fusion module, and a multi-task optimization output module; Among them, the input encoding module is used to receive various types of input features, including user behavior features, physical adsorption state features, and environmental features; the input encoding module uses a convolutional encoding network to perform preliminary extraction and non-linear transformation on the input features, mapping the low-dimensional input features to a high-dimensional space, thereby obtaining an encoded feature vector; The multi-dimensional attention feature extraction module is used to receive the encoded feature vector provided by the input encoding module; the multi-dimensional attention feature extraction module is implemented by a recurrent neural network based on a multi-dimensional attention mechanism, and is used to dynamically adjust the weights of features for different adsorption states and interaction modes to obtain a multi-dimensionally attention-weighted feature vector; The adaptive multi-state feature fusion module is used to receive the feature vector provided by the multi-dimensional attention feature extraction module; the adaptive multi-state feature fusion module uses a cross-feature combination strategy to cross-multiply and weighted-sum the features from different modes to generate a richer multi-mode feature representation; through a dynamic weighting function, adjust the fusion weights of the features according to the current physical adsorption state and user preferences to obtain a fused feature vector; The multi-task optimization output module is used to receive the fused feature vector provided by the adaptive multi-state feature fusion module; the multi-task optimization output module is designed based on a multi-task learning framework, and decomposes the final personalized parameter solution into multiple task outputs, and the multiple task outputs include: Task 1: Parameter output for video playback mode, including sound effect enhancement, stereo enhancement; Task 2: Parameter output for video call mode, including microphone sensitivity, noise reduction intensity, and speech clarity parameters; Task 3: Parameter output for music playback mode, including equalizer adjustment, spatial sound effect, and bass enhancement parameters.
7. The multi-mode interaction method between a Bluetooth speaker and a smart phone according to claim 6, wherein, The multi-dimensional attention feature extraction module is specifically used for: Calculate the multi-dimensional attention weights of each input feature through the following formula 1: ; Among them, is the attention weight at time step ; is the multi-dimensional attention weight matrix; is the encoded feature vector generated by the input encoding module; is the standard Sigmoid activation function; is the bias term used to balance the initial weights of different features; is the encoded feature vector generated by the input encoding module at time step ; 1 is the total number of time steps; Extract the time series features using the following formula 2: ; Among them, is the temporal feature vector at time step ; is the element-wise Hadamard product, which is used to multiply the attention weights with the features; is the input weight matrix of the recurrent neural network; is the hidden layer weight matrix of the recurrent unit; is the temporal feature state at the previous time step; is the bias term of the recurrent neural network; is the encoded feature vector generated by the input encoding module at time step ; Calculate the weighted feature vector according to the following formula 3: ; Among them, is the feature vector weighted by the multi-dimensional attention mechanism; is the periodic adjustment term; is the specific adjustment period.
8. The multi-mode interaction method between the Bluetooth speaker and the smart phone according to claim 6, characterized in that, The adaptive multi-state feature fusion module adopts a fusion strategy based on dynamic cross-feature combination and adaptive weight adjustment, and realizes the fusion through the following formula 4: ; Among them, is the finally fused feature vector; is the th weighted feature vector output by the multi-dimensional attention feature extraction module; is the element-wise Hadamard product, which is used to perform non-linear cross-combination of features in different modes; is the adjustment coefficient, which is used to control the non-linear amplitude of the features; is the th element of the fusion weight matrix, and its initial value is determined by the pre-training of the model and is dynamically updated through adaptive adjustment during the running process; is the periodic adjustment coefficient; is the current timestamp; is the specific adjustment period.
9. The multi-mode interaction method between a Bluetooth speaker and a smart phone according to claim 6, characterized in that The input encoding module is specifically used for: Receiving user behavior features, physical adsorption state features, and environmental features, and performing data preprocessing on these features, including standardization and normalization processing, to ensure that the input data is within the same numerical range; Performing feature extraction on the standardized features through a convolutional encoding network, mapping the input low-dimensional features to a high-dimensional feature space, so as to enhance the ability to capture the non-linear relationship between multi-dimensional features; Outputting the extracted high-dimensional feature vectors to the multi-dimensional attention feature extraction module for subsequent feature weighting and temporal processing.
10. The multi-mode interaction method between the Bluetooth speaker and the smart phone according to claim 6, characterized in that, The multi-task optimization output module is specifically used for: Decomposing the multi-mode interaction task into multiple sub-tasks, and adjusting the priority of the output parameters according to the personalized requirements in different modes; When generating the personalized parameter scheme for each sub-task, setting the output thresholds for sound enhancement, noise reduction intensity, and equalizer adjustment respectively according to the user's historical preferences and current environmental conditions; Updating the output weights of the model in real time through a multi-task learning framework.
Citation Information
Patent Citations
High-sensitivity multifunctional intelligent sound equipment
CN110278498A
Sound effect adjusting method and device of intelligent sound box
CN115374305A