Sound effect adjustment method and device, equipment and storage medium
By obtaining audio equipment information and analyzing audio content scenes, and dynamically adjusting the sound effect settings in combination with the personalized sound effect recommendation model, the problem of inflexible and accurate sound effect settings in the existing technology is solved, and a more accurate and personalized sound effect experience is achieved.
Patent Information
- Application Number
- CN202510422814.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-07
- Publication Date
- 2025-05-06
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
In the prior art, due to limited to the type information of the audio device, the sound effect settings are not flexible and accurate enough, and cannot provide the user with the best auditory experience.
By obtaining device information of the currently connected audio device, determine the audio device type, and select a preset sound effect configuration as the initial setting based on this type. Then, the audio signal processing algorithm is used to analyze the currently played audio content, identify the target audio scene, and adjust the initial sound effect settings according to the personalized sound effect recommendation model, and dynamically adjust to adapt to the current audio content and user preferences.
It realizes more accurate adaptation of sound effect settings, improves the accuracy, flexibility, personalization and dynamic sound effect settings, and provides users with a better listening experience.
Smart Images

Figure CN119937975A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of audio processing technology, and in particular to a sound effect adjustment method, device, equipment and storage medium. Background Art
[0002] By accurately setting the sound effects, the auditory experience of the audio content can be significantly improved, making the audience feel as if they are in a more real, delicate and layered sound world, thereby enjoying richer, three-dimensional and immersive sound effects, greatly enhancing the appeal and attraction of the audio content.
[0003] In the prior art, the type of connected audio device is generally obtained through device identification technology, and the sound effect configuration matching the device type is selected from the preset sound effect library as the target sound effect setting. However, the data obtained by this technical means is relatively single and less relevant, and is mainly limited to the type information of the audio device. This limitation makes the sound effect setting often not flexible and accurate, and cannot provide users with the best listening experience. Summary of the invention
[0004] The present invention provides a sound effect adjustment method, device, equipment and storage medium to solve the problem in the prior art that the sound effect setting is not flexible and accurate due to being limited to the type information of the audio device.
[0005] The first aspect of the present invention provides a sound effect adjustment method, comprising: obtaining device information of a currently connected target audio device; determining the type of the audio device according to the device information, and selecting a suitable preset sound effect configuration based on the audio device type to obtain an initial sound effect setting; analyzing the currently played target audio, and identifying the corresponding target audio scene using an audio signal processing algorithm; adjusting the initial sound effect setting according to the target audio scene and a preset personalized sound effect recommendation model to obtain a target sound effect setting that adapts to the current audio content and meets the user's preferences; continuously monitoring the playback status of the target audio content, and re-identifying the corresponding audio scene when the playback content changes; and dynamically adjusting the sound effect setting according to the new audio scene.
[0006] In a feasible implementation, before obtaining the device information of the currently connected target audio device, it also includes: collecting and analyzing the user's historical sound effect adjustment records, and extracting the user's sound effect preference characteristics in different scenarios; obtaining the user's usage habits, the user's usage habits including the frequency of use of sound effects, usage duration, and commonly used devices; using the user's sound effect preference characteristics in different scenarios and the user's usage habits as training samples to perform model training, and obtain a personalized sound effect recommendation model, which can recommend corresponding sound effect adjustment plans in real time according to the user's sound effect preferences in the current scenario.
[0007] In a feasible implementation, the obtaining of device information of the currently connected target audio device includes: detecting surrounding connectable audio devices and obtaining signal strength information of each connectable audio device; obtaining user historical connection records and device priority settings to evaluate each connectable audio device; based on the signal strength information of each connectable audio device and the evaluation result, selecting a target audio device for connection, and obtaining the device information of the target audio device.
[0008] In a feasible implementation, the obtaining of user historical connection records and device priority settings to evaluate each connectable audio device includes: retrieving user historical connection records to count the connection frequency and usage time of each connectable audio device, and comprehensively scoring each connectable audio device in combination with the device priority setting.
[0009] In a feasible implementation, the analysis of the currently playing audio content and the use of an audio signal processing algorithm to identify the corresponding target audio scene include: performing time-frequency analysis on the currently playing audio content to obtain multi-dimensional audio features; using a deep learning architecture to perform feature encoding and context association learning on the multi-dimensional audio features to obtain a high-level audio representation; based on the high-level audio representation, performing scene classification using a preset classifier to obtain a target audio scene.
[0010] In a feasible implementation, the time-frequency analysis of the currently playing audio content to obtain multi-dimensional audio features includes: extracting an audio signal corresponding to the currently playing audio content; applying a wavelet transform to decompose the audio signal into multiple frequency components and multiple time segments; and analyzing the multiple frequency components and the multiple time segments to form multi-dimensional audio features.
[0011] In a feasible implementation, the analyzing the multiple frequency components and the multiple time segments to form multi-dimensional audio features includes: using signal processing technology to process the multiple frequency components, analyzing the energy intensity, spectrum shape and changes of each frequency component in different time periods, and obtaining a first processing result; extracting audio events, rhythm patterns and sound textures in each time segment to obtain a second processing result; analyzing based on the first processing result and the second processing result to obtain a third processing result, the third processing result including the evolution law of frequency components with time segments, the interaction mechanism between frequency components and the correlation between audio events and frequency components in time segments; and combining the first processing result, the second processing result and the third processing result to obtain multi-dimensional audio features.
[0012] In a feasible implementation, the initial sound effect setting is adjusted according to the target audio scene and a preset personalized sound effect recommendation model to obtain a target sound effect setting that adapts to the current audio content and meets the user's preference, including: inputting the target audio scene into a preset personalized sound effect recommendation model to obtain a sound effect adjustment scheme that matches the target audio scene; determining adjustment content in combination with the sound effect adjustment scheme and the initial sound effect setting, the adjustment content including the sound effect parameters and adjustment amplitude that need to be adjusted; and adjusting the initial sound effect setting accordingly according to the adjustment content to obtain a target sound effect setting that adapts to the current audio content and meets the user's preference.
[0013] In a feasible implementation, the continuously monitoring the playback status of the target audio content and re-identifying the corresponding audio scene when the playback content changes includes: continuously monitoring the playback status of the target audio content, identifying key events, and determining whether the playback content has changed based on the key events; if it is determined that the playback content has changed based on the key events, analyzing the new audio content using audio feature extraction technology and scene classification algorithm to re-identify and determine its corresponding audio scene.
[0014] In a feasible implementation manner, the identifying of key events and determining whether the playback content has changed based on the key events include: identifying key events by monitoring metadata changes, chapter markers, advertisement insertion points, and audio feature analysis of the audio stream; setting a time window, and if a key event of audio stream switching, chapter change, advertisement insertion, or the start of a new track is detected within the time window, it is determined that the playback content has changed.
[0015] In a feasible implementation, after adjusting the initial sound effect setting according to the target audio scene to obtain the target sound effect setting adapted to the current audio content, it also includes: continuously monitoring the user's position change, and when the position change exceeds a preset change value, re-detecting the connectable audio device and reconnecting to the most suitable audio device.
[0016] In a feasible implementation, after adjusting the initial sound effect setting according to the target audio scene to obtain the target sound effect setting adapted to the current audio content, it also includes: real-time monitoring of user feedback on the sound effect setting, the feedback including user manual adjustment of sound effect parameters, user satisfaction score of the sound effect setting, and user voice evaluation of the sound effect; inputting the feedback data into the personalized sound effect recommendation model for model optimization to update the recommendation strategy of the personalized sound effect recommendation model.
[0017] The second aspect of the present invention provides a sound effect adjustment device, including: an acquisition module, used to obtain device information of a currently connected target audio device; a setting module, used to determine the type of audio device according to the device information, and select a suitable preset sound effect configuration based on the audio device type to obtain an initial sound effect setting; an identification module, used to analyze the currently played target audio, and use an audio signal processing algorithm to identify the corresponding target audio scene; a first adjustment module, used to adjust the initial sound effect setting according to the target audio scene and a preset personalized sound effect recommendation model, to obtain a target sound effect setting that adapts to the current audio content and meets the user's preference; a re-identification module, used to continuously monitor the playback status of the target audio content, and re-identify the corresponding audio scene when the playback content changes; a second adjustment module, used to dynamically adjust the sound effect setting according to the new audio scene.
[0018] In a feasible implementation, the sound effect adjustment device also includes: a model building module, which is used to collect and analyze the user's historical sound effect adjustment records and extract the user's sound effect preference characteristics in different scenarios; obtain the user's usage habits, which include the frequency of use of sound effects, usage duration and commonly used devices; use the user's sound effect preference characteristics in different scenarios and the user's usage habits as training samples to perform model training to obtain a personalized sound effect recommendation model, which can recommend corresponding sound effect adjustment plans in real time according to the user's sound effect preferences in the current scenario.
[0019] In a feasible implementation, the acquisition module includes: an acquisition unit, used to detect the connectable audio devices in the surroundings and obtain the signal strength information of each connectable audio device; an evaluation unit, used to obtain the user's historical connection records and device priority settings to evaluate each connectable audio device; a selection unit, used to select a target audio device for connection based on the signal strength information of each connectable audio device and the evaluation results, and obtain the device information of the target audio device.
[0020] In a feasible implementation, the evaluation unit is specifically used to retrieve the user's historical connection records to count the connection frequency and usage time of each connectable audio device, and comprehensively score each connectable audio device in combination with the device priority setting. In a feasible implementation, the recognition module includes: an analysis unit, used to perform time-frequency analysis on the currently played audio content to obtain multi-dimensional audio features; a processing unit, used to use a deep learning architecture to perform feature encoding and context association learning on the multi-dimensional audio features to obtain a high-level audio representation; a classification unit, used to perform scene classification based on the high-level audio representation through a preset classifier to obtain a target audio scene.
[0021] In a feasible implementation, the analysis unit includes: an extraction subunit, used to extract the audio signal corresponding to the currently played audio content; a decomposition subunit, used to apply wavelet transform to decompose the audio signal into multiple frequency components and multiple time segments; and an analysis subunit, used to analyze the multiple frequency components and the multiple time segments to form multi-dimensional audio features.
[0022] In a feasible implementation manner, the analysis subunit is specifically used to: use signal processing technology to process the multiple frequency components, analyze the energy intensity, spectral shape and changes of each frequency component in different time periods, and obtain a first processing result; extract the audio events, rhythm patterns and sound textures in each time segment to obtain a second processing result; analyze based on the first processing result and the second processing result to obtain a third processing result, and the third processing result includes the evolution law of the frequency components with the time segment, the interaction mechanism between the frequency components, and the correlation between the audio events and the frequency components in the time segment; and combine the first processing result, the second processing result and the third processing result to obtain multi-dimensional audio features.
[0023] In a feasible implementation, the first adjustment module is specifically used to: input the target audio scene into a preset personalized sound effect recommendation model to obtain a sound effect adjustment scheme that matches the target audio scene; determine the adjustment content in combination with the sound effect adjustment scheme and the initial sound effect setting, the adjustment content including the sound effect parameters that need to be adjusted and the adjustment range; and make corresponding adjustments to the initial sound effect setting according to the adjustment content to obtain a target sound effect setting that adapts to the current audio content and meets the user's preferences.
[0024] In a feasible implementation, the re-identification module includes: a judgment unit, which is used to continuously monitor the playback status of the target audio content, identify key events, and determine whether the playback content has changed based on the key events; an identification unit, which is used to analyze the new audio content using audio feature extraction technology and scene classification algorithm to re-identify and determine its corresponding audio scene if it is determined that the playback content has changed based on the key events.
[0025] In a feasible implementation, the judgment unit is specifically used to: identify key events by monitoring metadata changes, chapter markers, advertisement insertion points and audio feature analysis of the audio stream; set a time window, and if a key event of audio stream switching, chapter change, advertisement insertion or new track start is detected within the time window, it is determined that the playback content has changed.
[0026] In a feasible implementation, the sound effect adjustment device further includes: a reconnection module, which is used to continuously monitor the position change of the user, and when the position change exceeds a preset change value, re-detect the connectable audio devices and reconnect to the most suitable audio device.
[0027] In a feasible implementation, the sound effect adjustment device also includes: a feedback and optimization module, which is used to monitor in real time the user's feedback on the sound effect settings, the feedback including the user's manual adjustment of the sound effect parameters, the user's satisfaction score for the sound effect settings, and the user's voice evaluation of the sound effect; the feedback data is input into the personalized sound effect recommendation model for model optimization to update the recommendation strategy of the personalized sound effect recommendation model.
[0028] The third aspect of the present invention provides a sound effect adjustment device, comprising: a memory and at least one processor, wherein the memory stores instructions; the at least one processor calls the instructions in the memory so that the sound effect adjustment device executes the above-mentioned sound effect adjustment method.
[0029] A fourth aspect of the present invention provides a computer-readable storage medium, wherein the computer-readable storage medium stores instructions, which, when executed on a computer, enable the computer to execute the above-mentioned sound effect adjustment method.
[0030] In the technical solution provided by the present invention, the device information of the currently connected target audio device is obtained; the type of the audio device is determined according to the device information, and a suitable preset sound effect configuration is selected based on the type of the audio device to obtain the initial sound effect setting; the target audio currently being played is analyzed, and the corresponding target audio scene is identified by using an audio signal processing algorithm; the initial sound effect setting is adjusted according to the target audio scene and a preset personalized sound effect recommendation model to obtain a target sound effect setting that is adapted to the current audio content and meets the user's preference; the playback state of the target audio content is continuously monitored, and when the playback content changes, the corresponding audio scene is re-identified; and the sound effect setting is dynamically adjusted according to the new audio scene. In the embodiment of the present invention, the initial sound effect setting is determined according to the device information, and the current playback content is analyzed in combination with the audio signal processing algorithm to identify the target audio scene, so that the sound effect setting is more accurately adapted to the characteristics of the playback content. At the same time, a personalized sound effect recommendation model is introduced to meet the personalized needs of users. In addition, by continuously monitoring the playback state, content changes can be identified in real time and the sound effect setting can be dynamically adjusted to ensure that the sound effect always matches the current playback content, thereby improving the accuracy, flexibility, personalization and dynamism of the sound effect setting, and providing users with a better listening experience. BRIEF DESCRIPTION OF THE DRAWINGS
[0031] Figure 1 A schematic diagram of an embodiment of a sound effect adjustment method in an embodiment of the present invention; Figure 2 A schematic diagram of another embodiment of the sound effect adjustment method in an embodiment of the present invention; Figure 3 Schematic diagram of another embodiment of the sound effect adjustment method in the embodiment of the present invention; Figure 4 A schematic diagram of an embodiment of a sound effect adjustment device in an embodiment of the present invention; Figure 5 A schematic diagram of another embodiment of the sound effect adjustment device in an embodiment of the present invention; Figure 6 Schematic diagram of an embodiment of a sound effect adjustment device in an embodiment of the present invention. DETAILED DESCRIPTION
[0032] The embodiments of the present invention provide a sound effect adjustment method, apparatus, device and storage medium, which dynamically adjust the sound effect settings by combining device information, audio scene analysis and a personalized recommendation model, so that the sound effects are adapted to the device characteristics and audio content and meet the user preferences, thereby significantly improving the accuracy and flexibility of the sound effect settings.
[0033] The terms "first", "second", "third", "fourth", etc. (if any) in the specification and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged where appropriate, so that the embodiments described herein can be implemented in an order other than that illustrated or described herein. In addition, the terms "including" or "having" and any variations thereof are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units that are clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.
[0034] It is understandable that the execution subject of the present invention may be a sound effect adjustment device, or a terminal or a server, which is not limited here. The embodiment of the present invention is described by taking a server as the execution subject as an example.
[0035] For ease of understanding, the specific process of the embodiment of the present invention is described below. Figure 1 , an embodiment of the sound effect adjustment method in the embodiment of the present invention includes: 101. Get the device information of the currently connected target audio device; It should be stated that all user data involved in this application strictly complies with relevant laws and regulations, is obtained through legal channels with the user's explicit authorization, and is processed using anonymization or encryption technology to ensure that the data collection, storage and use process complies with the Personal Information Protection Law and GDPR and other privacy protection requirements.
[0036] The server communicates with the connected target audio device through a device driver or a dedicated audio interface protocol, such as Bluetooth audio protocols such as A2DP, AVRCP, or USB audio class protocols. During this process, the server sends a query request to the target audio device. The request contains the device information fields that need to be obtained, such as the device model, manufacturer identification, supported audio format list, and sound effect processing capabilities. After receiving the request, the target audio device returns the corresponding device information data packet. The server can parse these data packets to obtain detailed device information.
[0037] 102. Determine the audio device type according to the device information, and select a suitable preset sound effect configuration based on the audio device type to obtain an initial sound effect setting; The server parses the key fields in the device information obtained, such as the device model, manufacturer code, and hardware feature description, and compares and matches them with the preset device type database, so as to accurately identify the type of the currently connected audio device, such as headphones, home theater audio system, or car audio. Based on the determined audio device type, a set of the most suitable preset sound effect configurations is automatically selected as the initial sound effect settings from the preset sound effect configuration library according to the sound effect processing capabilities of the device type and the general preferences of users. General user preferences refer to the common or mainstream needs and preferences for sound effects of a large number of user groups for a specific type of audio device. These preferences can be obtained through user surveys, online evaluation analysis, historical usage data mining, etc., and reflect the user's emphasis on and preference for different sound effect features such as bass enhancement, treble clarity, sound balance, surround effects, etc.
[0038] 103. Analyze the currently played target audio and identify the corresponding target audio scene using an audio signal processing algorithm; The server obtains the target audio signal being played, and identifies the key features of the target audio signal through audio signal processing algorithms such as spectrum analysis, Mel-frequency cepstral coefficient extraction, and sound event detection. The key features can be sound frequency distribution, rhythm pattern, instrument type, vocal characteristics, and background noise, etc. The key features are compared and matched with the known scene features in the preset deep learning model to determine the target audio scene to which the target audio is most likely to belong. Audio scenes include concert scenes, natural environments, movie dialogues, or intense sports scenes. The deep learning model has been trained with a large amount of labeled data and can learn the mapping relationship between different audio scenes and key features.
[0039] 104. According to the target audio scene and the preset personalized sound effect recommendation model, the initial sound effect setting is adjusted to obtain the target sound effect setting that is adapted to the current audio content and meets the user's preference; After the server identifies the target audio scene to which the target audio belongs, it will input the relevant information of the target audio scene into a trained personalized sound effect recommendation model, which comprehensively considers the characteristics of the audio scene (such as the richness of instruments in music scenes and the clarity of dialogues in movie scenes), the user's historical sound effect preference data (such as the preference for bass and sensitivity to sound balance), and the device's own sound effect processing capabilities. The model uses the corresponding algorithm logic to calculate the sound effect adjustment parameters that can highlight the scene characteristics and fit the user's preferences in the current audio scene, such as equalizer settings, dynamic range compression ratio, surround sound effect intensity, etc.; then, the initial sound effect settings are adjusted according to these parameters to generate a set of target sound effect settings that are both adapted to the current audio content characteristics and meet the user's personalized hearing needs, thereby bringing users a more immersive and enjoyable audio experience.
[0040] For example, the personalized sound effect recommendation algorithm in the personalized sound effect recommendation model can generate customized sound effect parameter adjustment solutions by combining user historical preferences, current audio scenes and device characteristics. First, the algorithm maps the user ID, scene type and device type to low-dimensional embedding vectors to capture their potential characteristics; then, based on the collaborative filtering method, the sound effect preferences of similar users in the same scene are analyzed, and the deep neural network is used to fuse user embedding, scene embedding and device embedding to predict the optimal sound effect parameters; in addition, the algorithm also introduces rule-based basic recommendations as a supplement; finally, through the weighted fusion of collaborative filtering scores, deep network prediction results and basic recommendation values, personalized sound effect settings such as equalizer parameters, bass enhancement and treble adjustment are output. This technology can adapt to different devices and scenes, continuously optimize the recommendation effect, and significantly improve the user's auditory experience.
[0041] The expression formula for personalized sound effect recommendation is:
[0042] in, is the recommendation score of sound effect e when user u uses device d in scenario s, , , is a weight parameter used to balance the contributions of collaborative filtering, deep neural networks, and basic recommendations. represents the collaborative filtering recommendation score of the sound effect parameter configuration e by user u in the audio scene s. DNN() is the deep neural network prediction. is the user preference embedding vector, is the scene feature embedding vector, is the device feature embedding vector, base(s,d,e) is the rule-based recommendation; The calculation formula is:
[0043] in, are the k most similar users to user u, is the similarity between user u and user v, is the rating of sound effect e by user v in scene s; in, The calculation formula is:
[0044] in, is the user embedding matrix, is the one-hot encoding of the user ID; in, The calculation formula is:
[0045] in, is the scene embedding matrix, is the one-hot encoding of the scene ID; in, The calculation formula is:
[0046] in, is the device embedding matrix, is the one-hot encoding of the device ID.
[0047] 105. Continuously monitor the playback status of the target audio content, and when the playback content changes, re-identify the corresponding audio scene; The server performs spectral analysis on the audio signal through fast Fourier transform or short-time Fourier transform, extracts time-frequency features such as spectral energy distribution, frequency peaks and harmonic structure, and classifies the audio scenes in combination with machine learning models such as convolutional neural networks or support vector machines. At the same time, the audio stream is segmented using sliding window technology, and the feature vector of each audio segment is calculated in real time. The pre-trained audio scene recognition model is used to dynamically determine whether the current scene has changed. If a significant feature change is detected, such as a sudden change in the spectral energy distribution or a switch in the scene classification result, the process of re-identifying the audio scene is triggered.
[0048] 106. Dynamically adjust sound effect settings based on new audio scenarios.
[0049] When the server detects a change in the audio scene, it dynamically adjusts the sound effect settings based on the new scene. For example, when switching from music to a game scene, the spatial sense and positioning effects may be enhanced to improve the gaming experience. This dynamic adjustment ensures that the sound effect settings always match the current playback content, providing users with the best listening experience.
[0050] In an embodiment of the present invention, the initial sound effect setting is determined based on the device information, and the current playback content is analyzed in combination with the audio signal processing algorithm to identify the target audio scene, so that the sound effect setting is more accurately adapted to the characteristics of the playback content. At the same time, a personalized sound effect recommendation model is introduced to meet the personalized needs of users. In addition, by continuously monitoring the playback status, content changes can be identified in real time and the sound effect settings can be dynamically adjusted to ensure that the sound effects always match the current playback content, thereby improving the accuracy, flexibility, personalization and dynamism of the sound effect setting, and providing users with a better auditory experience.
[0051] See also Figure 2 Another embodiment of the sound effect adjustment method in the embodiment of the present invention includes: 201. Build a personalized sound effect recommendation model; The server collects and analyzes the user's historical sound effect adjustment records, extracts the user's sound effect preference characteristics in different scenarios; obtains the user's usage habits, which include the frequency of use of sound effects, usage duration, and commonly used devices; uses the user's sound effect preference characteristics in different scenarios and the user's usage habits as training samples for model training to obtain a personalized sound effect recommendation model. The personalized sound effect recommendation model can recommend corresponding sound effect adjustment plans in real time based on the user's sound effect preferences in the current scenario.
[0052] The server comprehensively collects users' past sound adjustment behavior data, including but not limited to users' adjustment records of volume, equalizer settings, surround sound effects and special sound effects in different scenarios such as watching movies, listening to music, and playing games; then, using data analysis techniques, such as cluster analysis or association rule mining, these records are deeply analyzed to identify the sound parameters and their trends that users repeatedly adjust in specific scenarios, so as to accurately extract the user's personalized sound preference characteristics in different audio application scenarios.
[0053] With user authorization, the server will regularly record the user's activity logs on various audio applications or devices, including the number of times the user activates the sound effect function, the duration of each use, and the specific device model used. In order to enhance the accuracy and completeness of the data, the server will also integrate usage data from third-party applications or devices with the user's explicit consent, and then use data analysis tools to summarize and classify these log data to calculate the user's frequency of use of the sound effect function in different time periods, the average duration of each use, and the type of audio device most commonly used by the user.
[0054] The server selects appropriate machine learning algorithms, such as collaborative filtering, content-based recommendations, or deep learning algorithms, which can accurately capture and predict users' sound preferences. In the model training phase, the training samples are divided into training sets and validation sets. The training set is used to teach the model to learn the intrinsic relationship between user sound preference characteristics, usage habits, and sound recommendation, while the validation set is used to verify the generalization ability of the model. A supervised learning strategy is adopted to guide model training with the help of data with labels such as user satisfaction ratings and click-through rates, and the model parameters are continuously optimized to improve its performance on the training set. After training, the validation set is used to conduct a comprehensive evaluation of the model, including key indicators such as recommendation accuracy and recall rate, and the effectiveness of the model in practical applications is verified through experimental methods such as A / B testing. Finally, based on the evaluation feedback, the model is carefully tuned, such as adjusting parameters, adding data, optimizing feature engineering, etc., to further improve recommendation performance and user experience.
[0055] 202. Detecting surrounding connectable audio devices and obtaining signal strength information of each connectable audio device; The server starts the device scanning function through the device's wireless communication module (such as Bluetooth, Wi-Fi, etc.), actively searches for and lists all connectable audio devices within the communication range, and uses specific communication protocols or API interfaces to establish temporary communication links with these devices to obtain the signal strength indicator (RSSI) value or other parameters indicating signal quality of each device. In this process, the security and privacy of communication must be ensured to avoid unauthorized device access and data leakage.
[0056] 203. Obtaining user historical connection records and device priority settings to evaluate each connectable audio device; The server retrieves the user's historical connection records to count the connection frequency and usage time of each connectable audio device, and gives a comprehensive score to each connectable audio device based on the device priority setting.
[0057] The server retrieves the user's historical connection records and counts the connection frequency and usage time of each audio device. These data can reflect the user's preferences and usage habits for the device. Then, based on the device type, brand, model and user-defined priority settings, combined with the device's connection frequency, usage time and weight, a comprehensive score is calculated through weighted summation to measure the priority and importance of each device. The user-defined priority settings can be: whether to prefer to use a specific brand of headphones or speakers, whether to tend to connect to high-quality devices, etc.
[0058] 204. Based on the signal strength information and evaluation results of each connectable audio device, select a target audio device for connection, and obtain device information of the target audio device; The server selects devices with stable signal strength and high comprehensive scores as candidates. These scores are based on factors such as the device's connection frequency, usage time, and user-defined priority. Then, from these candidate devices, the device with the strongest signal and the best overall performance is selected as the target audio device for connection to ensure the stability of audio transmission and the satisfaction of user experience. Once the connection is successfully established, the necessary device information is obtained from the target device through a specific communication protocol or API interface.
[0059] 205. Determine the audio device type according to the device information, and select a suitable preset sound effect configuration based on the audio device type to obtain an initial sound effect setting; The execution process of step 205 is similar to that of step 102, and will not be described again here.
[0060] 206. Analyze the currently played target audio and identify the corresponding target audio scene using an audio signal processing algorithm; The execution process of step 206 is similar to that of step 103, and will not be described again here.
[0061] 207. According to the target audio scene and the preset personalized sound effect recommendation model, the initial sound effect setting is adjusted to obtain the target sound effect setting that is adapted to the current audio content and meets the user's preference; The server inputs the target audio scene into a preset personalized sound effect recommendation model to obtain a sound effect adjustment plan that matches the target audio scene; determines the adjustment content based on the sound effect adjustment plan and the initial sound effect settings, and the adjustment content includes the sound effect parameters that need to be adjusted and the adjustment range; and makes corresponding adjustments to the initial sound effect settings based on the adjustment content to obtain a target sound effect setting that adapts to the current audio content and meets the user's preferences.
[0062] The server inputs the target audio scene into the preset personalized sound effect recommendation model. The model will analyze and match the sound effect parameters most relevant to the target audio scene through algorithms according to the user's sound effect preference characteristics and usage habits in different scenes, and automatically generate a sound effect adjustment plan that matches the scene. Check each sound effect parameter in the sound effect adjustment plan one by one, including but not limited to volume, equalizer band gain, surround sound mode selection, bass enhancement effect, etc., compare the difference between each parameter in the sound effect adjustment plan and the initial setting, and calculate the specific amplitude of each parameter that needs to be adjusted. Through this comparative analysis, the required adjustment sound effect parameters and the specific direction and amplitude of the adjustment can be clearly defined. You can use the control interface of the audio device or the dedicated sound effect adjustment software to modify the sound effect parameters one by one according to the adjustment instructions to ensure that the adjustment of each parameter is accurate. Through such an adjustment method, it can be ensured that the sound effect settings of the audio device can fully comply with the recommendation scheme of the personalized sound effect recommendation model, thereby providing users with a more accurate and personalized audio experience.
[0063] 208. Continuously monitor the playback status of the target audio content, and when the playback content changes, re-identify the corresponding audio scene; Continuously monitor the playback status of the target audio content, identify key events, and determine whether the playback content has changed based on the key events; if the playback content is determined to have changed based on the key events, use audio feature extraction technology and scene classification algorithms to analyze the new audio content to re-identify and determine its corresponding audio scene. Among them: the specific execution steps of identifying key events and determining whether the playback content has changed based on key events are: identifying key events by monitoring metadata changes, chapter markers, advertising insertion points and audio feature analysis of the audio stream; setting a time window, if a key event of audio stream switching, chapter change, advertising insertion or the start of a new track is detected within the time window, it is determined that the playback content has changed.
[0064] The server pays attention to the metadata changes of the audio stream. These metadata usually contain rich information about the audio content, such as track name, artist, album details, etc. Any change in metadata may indicate a switch in the playback content; at the same time, it also monitors the chapter markers in the audio stream. These markers are usually used to identify logical segments in the audio content, such as chapters in a book, parts of a podcast, etc. Changes in chapter markers also mean that the playback content may have changed; in addition, for audio streams containing advertisements, special attention is paid to the detection of advertisement insertion points. By analyzing the audio features or specific identifiers before and after the advertisement insertion, the start and end of the advertisement can be accurately captured, and then it can be determined whether the playback content has changed due to the advertisement insertion; in addition to the above methods, audio feature analysis technology is also used to monitor the rhythm, melody, timbre and other features of the audio stream in real time. Any significant changes in audio features may be a signal of a switch in playback content.
[0065] 209. Dynamically adjust sound effect settings according to new audio scenarios; The execution process of step 209 is similar to that of step 106, and will not be described again here.
[0066] 210. Continuously monitor the user's location changes. When the location changes exceed a preset change value, re-detect the connectable audio devices and reconnect to the most suitable audio device.
[0067] The user's location information is obtained in real time through positioning technology, and the change in distance between the user's location and the location when the audio device was last connected is calculated; when the distance change exceeds the preset threshold, the process of re-detecting connectable audio devices is triggered; based on the user's current location, available audio devices nearby are retrieved, and the most suitable audio device is selected for connection based on the device signal strength, device type, user's historical connection records, and device priority settings.
[0068] The specific steps for retrieving available audio devices nearby are: using geo-fencing technology, setting a preset radius with the user's current location as the center, and retrieving all connectable audio devices within the preset radius; obtaining real-time status information of connectable audio devices, including whether the device is online, the current load status of the device, and the sound configuration capabilities supported by the device; based on the real-time status information, filtering out audio devices that meet the connection conditions as available nearby audio devices.
[0069] In an embodiment of the present invention, the most suitable target audio device is selected by detecting surrounding connectable audio devices and combining signal strength, user history and device priority to ensure the stability and adaptability of the device connection. According to the device type and audio scene analysis, the initial settings are selected from the preset sound effect configuration, and the sound effect parameters are dynamically adjusted using a personalized sound effect recommendation model, so that the sound effect settings are consistent with both device characteristics and user preferences. In addition, the audio content playback status and user position changes are continuously monitored, content changes or position changes are identified in real time, and the sound effect settings are dynamically adjusted or the most suitable audio device is reconnected to ensure that the sound effects always match the current playback content and environment, providing users with a more accurate, flexible and personalized sound effect experience, and significantly improving the auditory effect and user satisfaction.
[0070] See also Figure 3 Another embodiment of the sound effect adjustment method in the embodiment of the present invention includes: 301. Obtain device information of the currently connected target audio device; 302. Determine the audio device type according to the device information, and select a suitable preset sound effect configuration based on the audio device type to obtain an initial sound effect setting; The execution process of steps 301-302 is similar to the above steps 101-102, and will not be repeated here.
[0071] 303. Perform time-frequency analysis on the currently played audio content to obtain multi-dimensional audio features; Extract the audio signal corresponding to the currently playing audio content; apply wavelet transform to decompose the audio signal into multiple frequency components and multiple time segments; analyze multiple frequency components and multiple time segments to form multi-dimensional audio features. Among them, the specific execution steps of analyzing multiple frequency components and multiple time segments to form multi-dimensional audio features are: use signal processing technology to process multiple frequency components, analyze the energy intensity, spectrum shape and changes of each frequency component in different time periods, and obtain the first processing result; extract the audio events, rhythm patterns and sound textures in each time segment to obtain the second processing result; analyze based on the first processing result and the second processing result to obtain the third processing result, which includes the evolution law of frequency components with time segments, the interaction mechanism between frequency components and the correlation between audio events and frequency components in time segments; combine the first processing result, the second processing result and the third processing result to obtain multi-dimensional audio features.
[0072] Wavelet transform technology is applied to perform multi-scale decomposition of audio signals, and the audio signals are decomposed into multiple frequency components and multiple time segments. Wavelet transform can capture the time domain and frequency domain characteristics of the signal at the same time, and is particularly suitable for the analysis of non-stationary signals. In the decomposition process, the discrete wavelet transform algorithm is used to decompose the audio signal into detail coefficients and approximate coefficients of different scales, thereby obtaining multiple frequency components and corresponding time segments. Signal processing technology is used to perform detailed analysis on each frequency component, and its energy intensity, spectrum shape and changes in different time periods are calculated to form the first processing result. At the same time, the audio events, rhythm patterns and sound textures in each time segment are extracted by the time domain analysis method to generate the second processing result. On this basis, the evolution law of frequency components with time segments, the interaction mechanism between frequency components and the correlation between audio events and frequency components in time segments are further analyzed to obtain the third processing result. Finally, the first processing result, the second processing result and the third processing result are integrated to form multi-dimensional audio features, which include energy distribution, spectrum characteristics, rhythm patterns, audio events and their correlation with frequency components in the time and frequency domains.
[0073] 304. Using a deep learning architecture to perform feature encoding and contextual learning on multi-dimensional audio features to obtain advanced audio representation; The server inputs the extracted multi-dimensional audio features into the deep learning architecture for feature encoding and contextual association learning. The deep learning architecture adopts a hybrid model of convolutional neural networks and recurrent neural networks. The convolutional neural network is used to capture local patterns and spectral characteristics in audio features, while the recurrent neural network is used to learn the contextual association of audio features in time series. In the convolutional neural network part, multi-layer convolutional layers and pooling layers are used to abstract the multi-dimensional audio features layer by layer to extract high-level spectral features and time-frequency patterns. In the recurrent neural network part, long short-term memory networks or gated recurrent units are used to perform sequence modeling on audio features within time segments to capture the dynamic changes and contextual dependencies of audio content in time. The outputs of the convolutional neural network and the recurrent neural network are fused, and high-level audio representations are further extracted through the fully connected layer. These high-level audio representations not only contain the spectral features of the audio, but also integrate the contextual information on the time series, which can more comprehensively describe the characteristics of the audio content.
[0074] 305. Based on the advanced audio representation, scene classification is performed by a preset classifier to obtain a target audio scene; The server inputs the high-level audio representation into a preset classifier for scene classification. The classifier uses a support vector machine or a deep neural network model to classify the high-level audio representation through a trained classification model to identify the target audio scene to which the current audio content belongs. The classifier's training data includes labeled samples of multiple audio scenes to ensure that the classifier can accurately distinguish the characteristics of different scenes.
[0075] The calculation formula of the audio scene probability distribution y is:
[0076] in, and is a learnable parameter, c is a context vector, which represents the weighted feature, where the calculation formula of c is:
[0077] in, is the attention weight, indicating the importance of the t-th time step, is the final output of the bidirectional LSTM, where The calculation formula is:
[0078] Among them, softmax normalizes the scores into probability distribution, is a learnable weight vector used to calculate The importance score of the output, is a linear transformation of the hidden state (W is the weight matrix, b is the bias), is a nonlinear activation function that compresses the value to the interval [−1,1] to enhance the expressive power of the model. The calculation formula is:
[0079] in, is the output of the forward LSTM (the hidden state of the forward LSTM, encoding the past context), is the output of the reverse LSTM (the hidden state of the reverse LSTM, encoding the future context), which is used to capture the temporal dependency of audio features, where and The calculation formula is
[0080]
[0081] in, is the feature vector of the tth time step in the input sequence, refers to the hidden state of the previous time step, refers to the hidden state at the next time step.
[0082] 306. According to the target audio scene and the preset personalized sound effect recommendation model, the initial sound effect setting is adjusted to obtain the target sound effect setting that is adapted to the current audio content and meets the user's preference; After the server adjusts the initial sound effect settings according to the target audio scene and the preset personalized sound effect recommendation model to obtain the target sound effect settings that adapt to the current audio content and meet the user's preferences, it also performs the following steps: real-time monitoring of the power status of the target audio device, and when the power status is lower than a preset threshold, automatically adjusting the sound effect settings to reduce the device's power consumption; re-evaluating the sound effect based on the sound effect settings after the device's power consumption is reduced, and ensuring that the audio content after the sound effect adjustment still meets the user's preferences.
[0083] The server identifies the sound effect parameters with higher power consumption in the current sound effect settings, including bass enhancement, surround sound effect, and high-frequency gain; based on the preset power consumption optimization strategy, the server gradually reduces the intensity of the sound effect parameters with higher power consumption; while reducing the intensity of the sound effect parameters, the server monitors the changes in the sound effect in real time to ensure that the audio content after the sound effect adjustment still meets the user's basic auditory needs.
[0084] The server monitors the power status of the target audio device in real time through the power management module of the device, obtains key information such as the current power percentage and remaining usage time, and automatically triggers the power consumption optimization process when it detects that the power status is lower than the preset threshold. The server first identifies the sound effect parameters with high power consumption in the current sound effect settings, including bass enhancement, surround sound effect, and high-frequency gain. These parameters usually require more computing resources and hardware support, resulting in higher power consumption. Then, according to the preset power consumption optimization strategy, the intensity of these high-power sound effect parameters is gradually reduced, for example, the gain value of bass enhancement is gradually reduced, the intensity of surround sound effect is reduced, and the amplitude of high-frequency gain is reduced. During the adjustment process, a progressive adjustment algorithm is used to ensure smooth transition of changes in sound effect parameters to avoid users from perceiving obvious sudden changes in sound quality. At the same time, the changes in sound effects are monitored in real time, and the fit between the adjusted sound effects and the audio content is analyzed through audio signal processing technology. The time-frequency analysis algorithm is used to evaluate key indicators such as spectral distribution, dynamic range, and sound clarity after the sound effect is adjusted to ensure that the sound effect can still meet the basic auditory needs of users. In addition, the adjusted sound effect is personalized and optimized in combination with the user's historical preference data to ensure that the audio content after the sound effect adjustment still meets the user's preferences. Finally, the adjusted sound effect settings are applied to the target audio device, and the device's power status and sound effect are continuously monitored to ensure that a good user experience is maintained while reducing power consumption.
[0085] 307. Continuously monitor the playback status of the target audio content, and when the playback content changes, re-identify the corresponding audio scene; 308. Dynamically adjust sound effect settings according to new audio scenarios.
[0086] The execution process of steps 307-308 is similar to the above steps 105-106, and will not be repeated here.
[0087] 309. Monitor user feedback on sound effect settings in real time and input into a personalized sound effect recommendation model for model optimization.
[0088] The server monitors user feedback on sound effect settings in real time, including user manual adjustment of sound effect parameters, user satisfaction rating of sound effect settings, and user voice evaluation of sound effect; the feedback data is input into the personalized sound effect recommendation model for model optimization to update the recommendation strategy of the personalized sound effect recommendation model.
[0089] The server captures the user's manual adjustment of sound effect parameters through the user interface, and records the changes in sound effect parameters before and after the adjustment; analyzes the user's voice evaluation of the sound effect through voice recognition technology, and extracts keywords and emotional tendencies; obtains the user's rating of the current sound effect settings through a preset satisfaction rating interface; and uses the sound effect parameter changes, voice evaluation keywords, emotional tendencies, and satisfaction ratings as feedback data, which are input into the personalized sound effect recommendation model for model optimization.
[0090] The server captures the user's manual adjustment of sound effect parameters in real time through the user interface, and records the changes in sound effect parameters before and after the adjustment, including the numerical changes of specific parameters such as bass gain, treble gain, and surround sound intensity; at the same time, the server collects the user's voice evaluation of the sound effect through the microphone, analyzes the voice content using natural language processing technology, extracts keywords and emotional tendencies, and keyword extraction uses an algorithm based on word frequency and semantic importance. Emotional tendency analysis is achieved through a pre-trained sentiment classification model to determine whether the user's evaluation is positive, negative, or neutral; in addition, through the preset satisfaction rating interface, guide users to rate the current sound effect settings. The rating range is set according to actual conditions, and the user's rating data is recorded ; The changes in sound effect parameters manually adjusted by users, keywords and emotional tendencies in voice evaluations, and satisfaction scores are used as multi-dimensional feedback data and input into the personalized sound effect recommendation model for model optimization; the model optimization uses an incremental learning algorithm to fuse new feedback data with historical data and update the model's weights and parameters; during the optimization process, the gradient descent algorithm is used to adjust the model's learning rate to ensure that the model can quickly adapt to the user's latest preferences; at the same time, the performance of the optimized model is evaluated through cross-validation technology to ensure the accuracy and stability of the recommendation strategy; the updated personalized sound effect recommendation model can generate a sound effect adjustment plan that is more in line with the user's preferences in real time based on the user's latest feedback, thereby improving the user experience.
[0091] In an embodiment of the present invention, the device type is determined and the initial sound effect settings are selected based on the device information of the target audio device to ensure that the sound effects match the device characteristics. The target audio scene is identified by performing time-frequency analysis and deep learning processing on the currently playing audio content, and the sound effect parameters are dynamically adjusted in combination with a personalized sound effect recommendation model to ensure that the sound effect settings are consistent with the characteristics of the audio content and meet user preferences. In addition, the audio content playback status is continuously monitored, content changes are identified in real time, and the sound effect settings are dynamically adjusted to ensure that the sound effects are always adapted to the currently playing content. At the same time, by real-time monitoring of user feedback and optimizing the personalized sound effect recommendation model, the sound effect settings are constantly adapted to the user's latest preferences, providing users with a more accurate, flexible and personalized sound effect experience, significantly improving the auditory effect and user satisfaction.
[0092] The above describes the sound effect adjustment method in the embodiment of the present invention. The following describes the sound effect adjustment device in the embodiment of the present invention. Figure 4 In one embodiment of the present invention, a sound effect adjustment device includes: The acquisition module 401 is used to acquire device information of the currently connected target audio device; The setting module 402 is used to determine the audio device type according to the device information, and select a suitable preset sound effect configuration based on the audio device type to obtain an initial sound effect setting; The recognition module 403 is used to analyze the currently played target audio and identify the corresponding target audio scene using an audio signal processing algorithm; The first adjustment module 404 is used to adjust the initial sound effect settings according to the target audio scene and the preset personalized sound effect recommendation model to obtain the target sound effect settings that are suitable for the current audio content and meet the user's preferences; A re-identification module 405 is used to continuously monitor the playback status of the target audio content and re-identify the corresponding audio scene when the playback content changes; The second adjustment module 406 is used to dynamically adjust the sound effect settings according to the new audio scene.
[0093] In an embodiment of the present invention, the initial sound effect setting is determined based on the device information, and the current playback content is analyzed in combination with the audio signal processing algorithm to identify the target audio scene, so that the sound effect setting is more accurately adapted to the characteristics of the playback content. At the same time, a personalized sound effect recommendation model is introduced to meet the personalized needs of users. In addition, by continuously monitoring the playback status, content changes can be identified in real time and the sound effect settings can be dynamically adjusted to ensure that the sound effects always match the current playback content, thereby improving the accuracy, flexibility, personalization and dynamism of the sound effect setting, and providing users with a better auditory experience.
[0094] See also Figure 5 Another embodiment of the sound effect adjustment device in the embodiment of the present invention includes: The acquisition module 401 is used to acquire device information of the currently connected target audio device; The setting module 402 is used to determine the audio device type according to the device information, and select a suitable preset sound effect configuration based on the audio device type to obtain an initial sound effect setting; The recognition module 403 is used to analyze the currently played target audio and identify the corresponding target audio scene using an audio signal processing algorithm; The first adjustment module 404 is used to adjust the initial sound effect settings according to the target audio scene and the preset personalized sound effect recommendation model to obtain the target sound effect settings that are suitable for the current audio content and meet the user's preferences; A re-identification module 405 is used to continuously monitor the playback status of the target audio content and re-identify the corresponding audio scene when the playback content changes; The second adjustment module 406 is used to dynamically adjust the sound effect settings according to the new audio scene.
[0095] Optionally, the sound effect adjustment device further includes: Model building module 407 is used to collect and analyze the user's historical sound effect adjustment records, extract the user's sound effect preference characteristics in different scenarios; obtain the user's usage habits, which include the frequency of use of sound effects, usage duration and commonly used devices; use the user's sound effect preference characteristics in different scenarios and the user's usage habits as training samples to train the model, and obtain a personalized sound effect recommendation model. The personalized sound effect recommendation model can recommend corresponding sound effect adjustment plans in real time based on the user's sound effect preferences in the current scenario.
[0096] Optionally, the acquisition module 401 includes: The acquisition unit 4011 is used to detect the connectable audio devices in the surroundings and acquire the signal strength information of each connectable audio device; An evaluation unit 4012 obtains user historical connection records and device priority settings to evaluate each connectable audio device; The selection unit 4013 is used to select a target audio device for connection based on the signal strength information and the evaluation result of each connectable audio device, and obtain device information of the target audio device.
[0097] Optionally, the evaluation unit 4012 may be specifically configured to: Retrieve the user's historical connection records to count the connection frequency and usage time of each connectable audio device, and make a comprehensive score for each connectable audio device based on the device priority setting. Optionally, the identification module 403 includes: The analysis unit 4031 is used to perform time-frequency analysis on the currently played audio content to obtain multi-dimensional audio features; Processing unit 4032, for performing feature encoding and contextual learning on multi-dimensional audio features using a deep learning architecture to obtain a high-level audio representation; The classification unit 4033 is used to perform scene classification based on the high-level audio representation through a preset classifier to obtain a target audio scene.
[0098] Optionally, the analysis unit 4031 includes: The extraction subunit 40311 is used to extract the audio signal corresponding to the currently played audio content; A decomposition subunit 40312, for decomposing the audio signal into a plurality of frequency components and a plurality of time segments by applying a wavelet transform; The analysis subunit 40313 is used to analyze multiple frequency components and multiple time segments to form multi-dimensional audio features.
[0099] Optionally, the analysis subunit 40311 may be specifically used for: Signal processing technology is used to process multiple frequency components, and the energy intensity, spectrum shape and changes of each frequency component in different time periods are analyzed to obtain a first processing result; audio events, rhythm patterns and sound textures in each time segment are extracted to obtain a second processing result; based on the first processing result and the second processing result, analysis is performed to obtain a third processing result, which includes the evolution law of frequency components with time segments, the interaction mechanism between frequency components and the correlation between audio events and frequency components in time segments; the first processing result, the second processing result and the third processing result are combined to obtain multi-dimensional audio features.
[0100] Optionally, the first adjustment module 404 may be specifically configured to: The target audio scene is input into a preset personalized sound effect recommendation model to obtain a sound effect adjustment plan that matches the target audio scene; the adjustment content is determined by combining the sound effect adjustment plan and the initial sound effect settings, and the adjustment content includes the sound effect parameters that need to be adjusted and the adjustment range; the initial sound effect settings are adjusted accordingly according to the adjustment content to obtain a target sound effect setting that adapts to the current audio content and meets the user's preferences.
[0101] Optionally, the re-identification module 405 includes: The judgment unit 4051 is used to continuously monitor the playback status of the target audio content, identify key events, and judge whether the playback content has changed according to the key events; The identification unit 4052 is used to analyze the new audio content by using audio feature extraction technology and scene classification algorithm to re-identify and determine its corresponding audio scene if it is determined that the playback content has changed according to the key event.
[0102] Optionally, the determining unit 4051 may be specifically configured to: Key events are identified by monitoring metadata changes, chapter markers, ad insertion points, and audio feature analysis of the audio stream. A time window is set, and if key events such as audio stream switching, chapter changes, ad insertion, or the start of a new track are detected within the time window, it is determined that the playback content has changed.
[0103] Optionally, the sound effect adjustment device further includes: The reconnection module 408 is used to continuously monitor the position change of the user, and when the position change exceeds a preset change value, re-detect the connectable audio devices and reconnect to the most suitable audio device.
[0104] Optionally, the sound effect adjustment device further includes: The feedback and optimization module 409 is used to monitor the user's feedback on the sound effect settings in real time. The feedback includes the user's manual adjustment of the sound effect parameters, the user's satisfaction score on the sound effect settings, and the user's voice evaluation of the sound effect; the feedback data is input into the personalized sound effect recommendation model for model optimization to update the recommendation strategy of the personalized sound effect recommendation model.
[0105] In an embodiment of the present invention, by detecting the surrounding connectable audio devices and combining the signal strength, user history and device priority, the most suitable target audio device is selected to ensure the stability and adaptability of the device connection, the initial sound effect settings are determined according to the device type, and the target audio scene is identified by performing time-frequency analysis and deep learning processing on the currently playing audio content, and the sound effect parameters are dynamically adjusted in combination with the personalized sound effect recommendation model, so that the sound effect settings are consistent with the device characteristics and audio content characteristics, and meet the user preferences. In addition, the audio content playback status and user position changes are continuously monitored, content changes or position changes are identified in real time, and the sound effect settings are dynamically adjusted or the most suitable audio device is reconnected to ensure that the sound effects always match the currently playing content and environment. At the same time, by real-time monitoring of user feedback and optimizing the personalized sound effect recommendation model, the sound effect settings are constantly adapted to the user's latest preferences, providing users with a more accurate, flexible and personalized sound effect experience, and significantly improving the auditory effect and user satisfaction.
[0106] above Figure 4 and Figure 5 The sound effect adjustment device in the embodiment of the present invention is described in detail from the perspective of modular functional entities, and the sound effect adjustment device in the embodiment of the present invention is described in detail from the perspective of hardware processing.
[0107] See also Figure 6 As shown, the sound effect adjustment device includes a processor 600 and a memory 601. The memory 601 stores machine executable instructions that can be executed by the processor 600. The processor 600 executes the machine executable instructions to implement the above-mentioned sound effect adjustment method.
[0108] Further, Figure 6 The sound effect adjustment device shown also includes a bus 602 and a communication interface 603 , and the processor 600 , the communication interface 603 and the memory 601 are connected via the bus 602 .
[0109] Among them, the memory 601 may include a high-speed random access memory (Random Access Memory, RAM), and may also include a non-volatile memory (non-volatile memory), for example, at least one disk storage. The communication connection between the system network element and at least one other network element is realized through at least one communication interface 603 (which can be wired or wireless), and the Internet, wide area network, local area network, metropolitan area network, etc. can be used. The bus 602 can be an ISA bus, a PCI bus or an EISA bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 6 Only one bidirectional arrow is used in the diagram, but this does not mean that there is only one bus or only one type of bus.
[0110] The processor 600 may be an integrated circuit chip with signal processing capabilities. In the implementation process, each step of the above method can be completed by the hardware integrated logic circuit or software instructions in the processor 600. The above processor 600 can be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; it can also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA) or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components. The disclosed methods, steps and logic block diagrams in the embodiments of the present disclosure can be implemented or executed. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor. The steps of the method disclosed in conjunction with the embodiments of the present disclosure can be directly embodied as a hardware decoding processor to be executed, or a combination of hardware and software modules in the decoding processor can be executed. The software module can be located in a mature storage medium in the field such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory or an electrically erasable programmable memory, a register, etc. The storage medium is located in the memory 601 , and the processor 600 reads the information in the memory 601 and completes the method steps of the above-mentioned embodiment in combination with its hardware.
[0111] The present invention also provides a sound effect adjustment device, wherein the computer device includes a memory and a processor, wherein the memory stores computer-readable instructions, and when the computer-readable instructions are executed by the processor, the processor executes the steps of the sound effect adjustment method in the above-mentioned embodiments. The present invention also provides a computer-readable storage medium, which can be a non-volatile computer-readable storage medium or a volatile computer-readable storage medium, wherein the computer-readable storage medium stores instructions, and when the instructions are executed on a computer, the computer executes the steps of the sound effect adjustment method.
[0112] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.
[0113] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention is essentially or the part that contributes to the prior art or the whole or part of the technical solution can be embodied in the form of a software product. The computer software product is stored in a storage medium, including several instructions to enable a computer device (which can be a personal computer, server, or network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM), random access memory (RAM), disk or optical disk and other media that can store program code.
[0114] As described above, the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit the same. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that the technical solutions described in the aforementioned embodiments may still be modified, or some of the technical features thereof may be replaced by equivalents. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A sound effect adjustment method, characterized in that: The sound effect adjustment method comprises: Get the device information of the currently connected target audio device; Determining the audio device type according to the device information, and selecting a suitable preset sound effect configuration based on the audio device type to obtain an initial sound effect setting; Analyze the currently playing target audio and identify the corresponding target audio scene using an audio signal processing algorithm; According to the target audio scene and the preset personalized sound effect recommendation model, the initial sound effect setting is adjusted to obtain a target sound effect setting that is adapted to the current audio content and meets the user's preference; Continuously monitoring the playback status of the target audio content, and re-identifying the corresponding audio scene when the playback content changes; Dynamically adjust sound settings based on new audio scenarios.
2. The sound effect adjustment method according to claim 1, characterized in that: Before obtaining the device information of the currently connected target audio device, it also includes: Collect and analyze the user's historical sound effect adjustment records to extract the user's sound effect preference characteristics in different scenarios; Acquire user usage habits, including the frequency of use of sound effects, duration of use, and commonly used devices; The user's sound effect preference characteristics in different scenarios and the user's usage habits are used as training samples for model training to obtain a personalized sound effect recommendation model. The personalized sound effect recommendation model can recommend corresponding sound effect adjustment plans in real time according to the user's sound effect preferences in the current scenario.
3. The sound effect adjustment method according to claim 1, characterized in that: The obtaining of device information of the currently connected target audio device includes: Detect the surrounding connectable audio devices and obtain the signal strength information of each connectable audio device; Obtain user connection history and device priority settings to evaluate each connectable audio device; Based on the signal strength information and the evaluation result of each connectable audio device, a target audio device is selected for connection, and device information of the target audio device is obtained.
4. The sound effect adjustment method according to claim 3, characterized in that: The obtaining of the user's historical connection records and device priority settings to evaluate each connectable audio device includes: Retrieve the user's historical connection records to count the connection frequency and usage time of each connectable audio device, and make a comprehensive score for each connectable audio device based on the device priority setting.
5. The sound effect adjustment method according to claim 1, characterized in that: The analyzing the currently played audio content and identifying the corresponding target audio scene using an audio signal processing algorithm includes: Perform time-frequency analysis on the currently playing audio content to obtain multi-dimensional audio features; Using a deep learning architecture to perform feature encoding and contextual learning on the multi-dimensional audio features to obtain a high-level audio representation; Based on the high-level audio representation, scene classification is performed through a preset classifier to obtain a target audio scene.
6. The sound effect adjustment method according to claim 5, characterized in that: The time-frequency analysis of the currently played audio content is performed to obtain multi-dimensional audio features, including: Extracting the audio signal corresponding to the currently playing audio content; Applying wavelet transform to decompose the audio signal into a plurality of frequency components and a plurality of time segments; The multiple frequency components and the multiple time segments are analyzed to form multi-dimensional audio features.
7. The sound effect adjustment method according to claim 6, characterized in that: The analyzing the multiple frequency components and the multiple time segments to form a multi-dimensional audio feature includes: Processing the multiple frequency components using a signal processing technique, analyzing the energy intensity, spectrum shape, and changes in different time periods of each frequency component, and obtaining a first processing result; extracting audio events, rhythmic patterns, and sound textures within each time segment to obtain a second processing result; Analyze the first processing result and the second processing result to obtain a third processing result, wherein the third processing result includes the evolution law of the frequency components with the time segment, the interaction mechanism between the frequency components, and the correlation between the audio event and the frequency components in the time segment; The first processing result, the second processing result and the third processing result are integrated to obtain multi-dimensional audio features.
8. The sound effect adjustment method according to claim 1, characterized in that: The step of adjusting the initial sound effect setting according to the target audio scene and the preset personalized sound effect recommendation model to obtain a target sound effect setting adapted to the current audio content and in accordance with the user's preference includes: Inputting the target audio scene into a preset personalized sound effect recommendation model to obtain a sound effect adjustment solution matching the target audio scene; Determine the adjustment content by combining the sound effect adjustment scheme and the initial sound effect setting, wherein the adjustment content includes the sound effect parameters to be adjusted and the adjustment range; The initial sound effect settings are adjusted accordingly according to the adjustment content to obtain target sound effect settings that are adapted to the current audio content and meet the user's preferences.
9. The sound effect adjustment method according to claim 1, characterized in that: The continuously monitoring the playing state of the target audio content and re-identifying the corresponding audio scene when the playing content changes includes: Continuously monitoring the playback status of the target audio content, identifying key events, and determining whether the playback content has changed based on the key events; If it is determined according to the key event that the playback content has changed, the new audio content is analyzed using audio feature extraction technology and scene classification algorithm to re-identify and determine its corresponding audio scene.
10. The sound effect adjustment method according to claim 9, characterized in that: The identifying of the key event and judging whether the playback content has changed according to the key event includes: Identify key events by monitoring audio stream metadata changes, chapter markers, ad insertion points, and audio feature analysis; A time window is set, and if a key event such as audio stream switching, chapter change, advertisement insertion, or start of a new track is detected within the time window, it is determined that the playback content has changed.
11. The sound effect adjustment method according to claim 1, characterized in that: After the initial sound effect setting is adjusted according to the target audio scene to obtain the target sound effect setting adapted to the current audio content, the method further includes: The user's position change is continuously monitored, and when the position change exceeds a preset change value, the connectable audio device is re-detected and the most suitable audio device is re-connected.
12. The sound effect adjustment method according to claim 1, characterized in that: After the initial sound effect setting is adjusted according to the target audio scene to obtain the target sound effect setting adapted to the current audio content, the method further includes: Real-time monitoring of user feedback on sound effect settings, including user manual adjustment of sound effect parameters, user satisfaction rating of sound effect settings, and user voice evaluation of sound effect; The feedback data is input into the personalized sound effect recommendation model for model optimization to update the recommendation strategy of the personalized sound effect recommendation model.
13. A sound effect adjustment device, characterized in that: The sound effect adjustment device comprises: The acquisition module is used to obtain the device information of the currently connected target audio device; A setting module, used to determine the type of audio device according to the device information, and select a suitable preset sound effect configuration based on the audio device type to obtain an initial sound effect setting; The recognition module is used to analyze the currently playing target audio and identify the corresponding target audio scene using an audio signal processing algorithm; A first adjustment module, configured to adjust the initial sound effect setting according to the target audio scene and a preset personalized sound effect recommendation model to obtain a target sound effect setting that is adapted to the current audio content and meets the user's preference; A re-identification module, used to continuously monitor the playback status of the target audio content, and re-identify the corresponding audio scene when the playback content changes; The second adjustment module is used to dynamically adjust the sound effect settings according to the new audio scene.
14. A sound effect adjustment device, characterized in that: The sound effect adjustment device comprises: a memory and at least one processor, wherein instructions are stored in the memory; The at least one processor calls the instructions in the memory so that the sound effect adjustment device executes the sound effect adjustment method as described in any one of claims 1-12.
15. A computer-readable storage medium having instructions stored thereon, characterized in that: When the instructions are executed by the processor, the sound effect adjustment method as described in any one of claims 1-12 is implemented.
Citation Information
Patent Citations
Display equipment and sound effect setting method of audio equipment
CN116939262A
Method for realizing personalized sound effect adjustment of loudspeaker through collaborative filtering algorithm
CN119201031A
Sound effect intelligent automatic adjustment method and device based on television scene and intelligent television
CN119545076A
Sound channel correction method and device, equipment and storage medium
CN119629543A
Cited By
Sound effect generation method and device, television and computer readable storage medium
CN120980314A
Intelligent radio content dynamic adjustment method based on scene perception
CN121117254A
Audio enhancement method and system for Bluetooth earphone
CN121126190A
Sound effect adjustment method and system, electronic equipment and storage medium
CN121635837A
A fine sound effect intelligent adjustment method and system
CN122844788A