Audio playing control method and device, equipment and storage medium

By constructing a three-dimensional audio spatial morphological model and deep audio semantic analysis, combined with adaptive audio parameter regulation, intelligent audio playback control is realized, solving the problems of low efficiency and poor flexibility of traditional methods, and improving user experience and audio quality.

CN120215870AInactive Publication Date: 2025-06-27QILIXING TECHNOLOGY (SHENZHEN) CO LTD
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510358026.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-25
Publication Date
2025-06-27
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Traditional audio playback control methods are inefficient and have poor flexibility, and cannot achieve smooth and convenient control effects. They lack perception and adaptation to the user environment and operating habits, resulting in poor user experience.

Method used

By obtaining real-time playback audio signals and spatial multi-angle reflected audio signals, a three-dimensional audio spatial morphology model is constructed, spatial audio attenuation correction and deep audio semantic analysis, speculate user audio emotions and predict needs, and adaptive audio parameter regulation is carried out to realize intelligent playback control.

Benefits of technology

It realizes accurate propagation and environmental adaptation of audio in space, improves audio quality and user immersion experience, and provides personalized and intelligent audio playback services that can respond to user needs and environmental changes in real time.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120215870A_ABST
    Figure CN120215870A_ABST
Patent Text Reader

Abstract

The invention relates to the field of audio playing control, in particular to an audio playing control method and device, equipment and a storage medium. The method comprises the following steps: acquiring a real-time playing audio signal and a spatial multi-angle reflection audio signal; performing one-by-one angle signal reflection time calculation and audio spatial three-dimensional form modeling on the spatial multi-angle reflection audio signals to construct a three-dimensional audio spatial form model; performing spatial angle reflection attenuation evolution on the spatial multi-angle reflection audio signals one by one, and performing spatial audio attenuation correction based on the three-dimensional audio spatial form model so as to obtain spatial attenuation compensation audio signals; and carrying out audio rhythm fluctuation analysis on the spatial attenuation compensation audio signal, and carrying out deep audio semantic analysis so as to obtain an audio semantic feature of the audio signal. According to the invention, personalized audio playing is realized, the audio parameters are dynamically adjusted based on the environmental noise, and the audio experience of the user is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of audio playback control, and particularly to an audio playback control method, apparatus, device, and storage medium. Background Art

[0002] As one of the core functions in multimedia applications, audio playback is widely used in fields such as entertainment, education, and communication. With the popularization of smart devices and the continuous progress of technology, audio playback systems have evolved from traditional single-playback functions to diversified and intelligent directions. In this process, users have higher and higher requirements for the audio playback experience, not only limited to the improvement of playback quality, but also including intelligent control, personalized needs, and multi-task collaboration during the audio playback process.

[0003] Traditional audio playback control methods mostly rely on physical buttons or simple touch operations. Although they can meet basic playback control requirements, with the continuous change of user needs, traditional methods have gradually exposed problems of low efficiency and poor flexibility. When quickly adjusting the volume, selecting a playlist, or switching audio, users often need to repeatedly operate the control interface, resulting in cumbersome operations and unable to achieve smooth and convenient control effects. At the same time, most existing audio playback control methods lack the perception and adaptation to the user environment, and cannot intelligently identify the user's operation habits or situational needs, resulting in poor user experience. Therefore, with the continuous evolution of technology and the change of user needs, there is an urgent need for a new audio playback control method. Summary of the Invention

[0004] To solve the above technical problems, the present invention provides an audio playback control method, apparatus, device, and storage medium to solve at least one of the above technical problems.

[0005] To achieve the above object, the present invention provides an audio playback control method, including the following steps: Step S1: Obtain a real-time played audio signal and a spatially multi-angled reflected audio signal; calculate the signal reflection time for each angle of the spatially multi-angled reflected audio signal and perform three-dimensional audio space form modeling to construct a three-dimensional audio space form model; Step S2: Perform attenuation evolution for each spatial angle reflection of the spatially multi-angled reflected audio signal, and perform spatial audio attenuation correction based on the three-dimensional audio space form model to obtain a spatially attenuated compensated audio signal; Step S3: Perform audio rhythm fluctuation analysis on the spatially attenuated compensated audio signal and perform deep audio semantic parsing to obtain the audio semantic features of the audio signal; Step S4: Infer the user's audio emotion based on the audio semantic features of the audio signal, and predict the evolution of the audio demand, so as to generate the user's real-time audio demand prediction data; Step S5: Perform all-round environmental noise perception on the spatially attenuated compensated audio signal, and then perform adaptive audio parameter regulation according to the user's real-time audio demand prediction data, so as to generate a dynamically regulated audio signal; Step S6: Perform real-time playback processing based on the dynamically regulated audio signal, and make an intelligent playback control decision to construct an intelligent playback control model.

[0006] In this specification, an audio playback control device is provided for executing the audio playback control method as described above, including: An audio space module, configured to obtain a real-time playback audio signal and a spatially multi-angle reflected audio signal; calculate the signal reflection time for each angle of the spatially multi-angle reflected audio signal and perform three-dimensional audio space morphology modeling, so as to construct a three-dimensional audio space morphology model; An audio attenuation correction module, configured to perform the evolution of the reflection attenuation for each spatial angle of the spatially multi-angle reflected audio signal, and perform spatial audio attenuation correction based on the three-dimensional audio space morphology model, so as to obtain a spatially attenuated compensated audio signal; An audio semantics module, configured to perform audio rhythm fluctuation analysis on the spatially attenuated compensated audio signal and perform deep audio semantics parsing, so as to obtain the audio semantic features of the audio signal; An audio demand prediction module, configured to infer the user's audio emotion based on the audio semantic features of the audio signal, and predict the evolution of the audio demand, so as to generate the user's real-time audio demand prediction data; An audio parameter regulation module, configured to perform all-round environmental noise perception on the spatially attenuated compensated audio signal, and then perform adaptive audio parameter regulation according to the user's real-time audio demand prediction data, so as to generate a dynamically regulated audio signal; An intelligent playback control module, configured to perform real-time playback processing based on the dynamically regulated audio signal, and make an intelligent playback control decision to construct an intelligent playback control model.

[0007] The present invention also provides a computer device, including a memory and a processor, where a computer program is stored in the memory, and when the processor executes the computer program, the steps of the audio playback control method described in any one of the above are implemented.

[0008] The present invention also provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps of the audio playback control method described in any one of the above are implemented.

[0009] The beneficial effects of the present invention are specifically as follows: By obtaining audio signals and multi-angle reflection signals in real time, the propagation path and reflection characteristics of audio in space can be comprehensively captured, ensuring the accuracy and sense of space of audio playback. By constructing a three-dimensional audio space form model, the propagation characteristics of sound in different physical environments can be accurately presented, forming a comprehensive understanding of the audio propagation environment. This spatial modeling method can provide a more realistic and three-dimensional audio performance for audio players, especially suitable for scenarios such as virtual reality (VR) and augmented reality (AR), greatly enhancing the user's immersive experience. The attenuation evolution processing of the reflection signal eliminates the sound quality distortion caused by spatial factors (such as reflections from objects like walls, ceilings, floors, etc.), enabling the audio to maintain good quality in different environments. Through precise spatial audio attenuation correction, the attenuation of audio signals in different directions and angles can be compensated, ensuring that users can hear balanced sounds at any position without being affected by the environmental structure. Through audio rhythm fluctuation analysis and deep semantic parsing, the emotional tone, rhythm changes, speech content, and tone of the audio can be identified, deeply understanding the potential meaning of the audio signal, identifying the climax part of a piece of music, or identifying the emotional turning point in speech. By extracting the semantic features of the audio, a more personalized audio playback service can be provided for users, such as making appropriate adjustments or recommendations according to the emotional characteristics and content of the audio, which better meets the needs of users. This stage enables the player to have content recognition and semantic processing capabilities, no longer relying solely on simple audio signal processing, but being able to perform deeper intelligent analysis on the audio. By inferring the user's emotions based on the semantic features of the audio, the audio content and playback mode can be automatically adjusted. If the system recognizes that the user is in a low mood, it will play more relaxing or inspiring audio to achieve an emotional adjustment effect. By predicting the user's audio needs, the player can react before the user's needs change, providing a more proactive and intelligent service. Predicting the music preference or audio content according to the user's usage habits at different time periods, as the user's needs evolve, the system can be adjusted according to the user's long-term behavior, making each user's audio experience more personalized and unique. Through the perception of environmental noise, the system can automatically adjust the audio playback parameters according to changes in the external environment, such as enhancing speech clarity in a noisy environment and reducing the volume in a quiet environment, ensuring the consistency of the audio experience. Dynamically adjusting audio parameters (such as volume, frequency, equalizer settings, etc.) to ensure the best sound quality and listening experience in different environments. This adaptive function enables the audio playback to respond to environmental changes in real time without the need for manual intervention by the user. By optimizing the audio parameters, excessive volume or overly harsh audio can be avoided, enhancing the comfort and experience of users in different environments. Through intelligent playback control decisions, the audio player can optimize and process in real time according to factors such as the environment, user emotions, and audio content, improving the playback effect, whether it is automatically adjusting the audio effect or changing the playback mode,Both can ensure that the audio output always matches the user's needs. Through the intelligent playback control model, the player can interact with the user in a more natural and intelligent way, continuously adjust the audio playback method according to real-time feedback, avoid rigid operation processes, and through the intelligent decision-making model, the audio player makes personalized adjustments according to specific situations, thus greatly improving the user's satisfaction and forming a smooth and seamless audio experience. BRIEF DESCRIPTION OF THE DRAWINGS

[0010] Figure 1 is a schematic flow chart of the steps of a method for controlling audio playback according to the present invention; Figure 2 is a schematic detailed implementation step flow chart of step S1; Figure 3 is a schematic detailed implementation step flow chart of step S2; Figure 4 is a schematic detailed implementation step flow chart of step S3. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0011] It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0012] The embodiments of the present application provide a method, device, equipment and storage medium for controlling audio playback. The execution subjects of the method, device, equipment and storage medium for controlling audio playback include but are not limited to: mechanical equipment, data processing platforms, cloud server nodes, network upload devices, etc. that are equipped with this system and can be regarded as general computing nodes of the present application. The data processing platform includes but is not limited to: at least one of an audio image management system, an information management system, and a cloud data management system.

[0013] Please refer to Figures 1 to 4 , the present invention provides a method for controlling audio playback, and the method for controlling audio playback includes the following steps: Step S1: Obtain real-time played audio signals and spatially multi-angular reflected audio signals; calculate the signal reflection time for each angle of the spatially multi-angular reflected audio signals and perform three-dimensional audio spatial form modeling to construct a three-dimensional audio spatial form model; Step S2: Perform attenuation evolution of the spatially multi-angular reflected audio signals for each spatial angle and perform spatial audio attenuation correction based on the three-dimensional audio spatial form model to obtain spatially attenuated compensated audio signals; Step S3: Perform audio rhythm fluctuation analysis on the spatially attenuated compensated audio signals and perform in-depth audio semantic parsing to obtain the audio semantic features of the audio signals; Step S4: Infer the user's audio emotion based on the audio semantic features of the audio signals and predict the evolution of audio requirements to generate user real-time audio requirement prediction data; Step S5: Perform omni-directional ambient noise perception on the spatially attenuated compensated audio signal, and then adaptively adjust audio parameters according to the predicted data of the user's real-time audio requirements, so as to generate a dynamically adjusted audio signal; Step S6: Perform real-time playback processing based on the dynamically adjusted audio signal, and make an intelligent playback control decision to construct an intelligent playback control model.

[0014] In the embodiment of the present invention, refer to Figure 1 , which is a schematic diagram of the step flow of an audio playback control method of the present invention. In this example, the steps of the audio playback control method include: Step S1: Obtain a real-time playback audio signal and a spatially multi-angle reflected audio signal; calculate the signal reflection time for each angle of the spatially multi-angle reflected audio signal and perform audio spatial three-dimensional form modeling to construct a three-dimensional audio space form model; In this embodiment, a high-quality digital audio workstation (DAW) or audio recording software (such as Ableton Live or Pro Tools) is used to capture real-time audio signals. Ensure that the sampling rate is set to 44.1 kHz to capture high-fidelity audio details. Record multiple audio tracks of the audio signal, such as musical instruments, vocals, and ambient sounds, for subsequent analysis. Each track should be collected separately and ensure that its signal strength is moderate to avoid distortion. Store the real-time captured audio signal in a lossless format, such as WAV or FLAC, to maintain audio quality. Use a structured data format (such as JSON or CSV) to record the metadata of the audio signal, including information such as timestamps, track types, and audio lengths. Arrange multiple microphones (such as supercardioid or omnidirectional microphones) at specific spatial positions to ensure that reflected signals can be captured from different angles. It is recommended to use at least 8 microphones to form a circular or stereo configuration to cover a 360-degree space. Set the sampling rate of the microphones to be the same as the real-time played audio signal (44.1 kHz), for subsequent signal processing and analysis. Connect all microphone signals to a computer using an audio interface to ensure the synchronization of the data stream. Use a digital audio interface (such as Focusrite Scarlett) for signal recording to prevent latency and phase issues. Collect the signals of all microphones simultaneously through audio recording software and use clock synchronization technology to ensure the time alignment of all audio signals. Conduct time-domain analysis on the audio signals captured by each microphone. Use the Autocorrelation Function to detect the reflected components in the signals. Use libraries in Python (such as NumPy or SciPy) for signal processing. Set a threshold to determine the starting and ending points of the reflected signals. Calculate the reflection time based on the time difference between the reflected signal and the original signal. Extract the time delay of the reflected signal. Record the reflection time of each microphone in structured data, ensuring that the reflection time of each signal corresponds to the position of the microphone. Select a suitable 3D modeling software (such as Blender or MATLAB) to create a spatial model. Define a 3D coordinate system based on the data of the reflection time and microphone position. Convert the position of each microphone and its corresponding reflection time into 3D coordinate points, (x, y, z) = (r ⋅ cos(θ), r ⋅ sin(θ), h), where r is the distance between the microphone and the sound source, 𝜃 is the angle, and h is the height. Represent the propagation path and spatial form of the sound using surfaces or cubes based on the reflection time and spatial coordinates. Map the intensity of the reflected signal into the 3D model and use different colors or transparencies to represent different reflection intensities. After completing the model, generate a visualization image of the audio space. Generate high-quality 3D images through a rendering tool (such as the Cycles engine of Blender) to display the propagation characteristics of the audio in space. Use visualization techniques to analyze the behavior of the sound and observe the influence of different reflection angles on the audio signal to help understand the propagation characteristics of the sound in a specific environment.

[0015] Step S2: Perform one-by-one spatial angle reflection attenuation evolution on the spatially multi-angle reflected audio signals and conduct spatial audio attenuation correction based on the 3D audio spatial form model, so as to obtain spatially attenuation-compensated audio signals; In this embodiment, based on the spatially multi - angular reflected audio signals obtained in the previous step, the signals of each microphone are selected for analysis. According to different directions and distances, the signals are divided into multiple samples for analysis from each angle. A time window (such as 512 samples) is set to frame each signal to ensure capturing the transient characteristics of the signal. For each microphone signal, the least - squares method is used to fit the attenuation curve of the signal to estimate the attenuation coefficient 𝛼. The SciPy library in Python is used for curve fitting. The attenuation coefficient of each direction and the curve of signal intensity varying with time are recorded for subsequent correction and analysis. A data visualization tool (such as Matplotlib) is used to plot the attenuation diagram to visually display the attenuation characteristics of different angles, analyze the influence of different environmental factors (such as walls, obstacles, etc.) on signal attenuation, and organize and store the relevant data. Using the microphone position and sound source position data in the three - dimensional space model, the propagation path of the signal received by each microphone is calculated. According to the attenuation coefficient and the calculated distance of each microphone, spatial attenuation correction is performed. The corrected signal is reconstructed to generate a spatially attenuated compensated audio signal. Synthetic technology is used to mix the corrected signals in different directions to form a stereo or surround - sound effect. An audio processing software (such as MATLAB or Audacity) is used to implement signal reconstruction and output to ensure the maintenance of audio quality and the accuracy of the signal. The characteristics of the output signal are recorded, including the spectrum, dynamic range, signal - to - noise ratio, etc. of the signal, for subsequent analysis and optimization.

[0016] Step S3: Conduct an analysis of the audio rhythm fluctuations of the spatially attenuated compensated audio signal and perform in - depth audio semantic parsing to obtain the audio semantic features of the audio signal; In this embodiment, the spatially attenuated compensated audio signal is preprocessed, including noise removal and normalization, to improve the accuracy of subsequent analysis. A band-pass filter (such as a frequency range of 3 Hz to 20 kHz) is used to remove low-frequency noise and high-frequency interference. The short-time Fourier transform (STFT) is used to convert the signal into the time-frequency domain, with a window size of 1024 samples and an overlap rate of 50% to ensure good time and frequency resolution, so as to better capture the rhythm changes in the signal. A beat detection algorithm (such as dynamic time warping (DTW) or energy-based beat detection) is used to identify the beats in the audio signal. An energy threshold is set to ensure that only significant rhythm changes are captured. The rhythm characteristics of the audio signal are calculated, including beats per minute (BPM) and beat interval time. The librosa library in Python is used for rhythm analysis to obtain the time series of rhythm changes. The extracted rhythm characteristics are statistically analyzed to observe the stability and fluctuations of the rhythm, and the standard deviation and mean are calculated to evaluate the amplitude of rhythm fluctuations in the audio signal. A visualization tool (such as Matplotlib) is used to plot the rhythm fluctuation graph to visually display the trend and pattern of rhythm changes, helping to understand the structure of the audio signal. Spectral features (such as Mel-frequency cepstral coefficients (MFCC) and pitch features) are used to represent the audio features of the audio signal. Extraction parameters are set, such as the dimension of MFCC being 13, and the MFCC features of each frame are calculated and processed using a frame step (such as 512 samples). Through the extraction of features such as harmony, timbre, and dynamic range of the audio signal, a multi-dimensional feature vector of the audio is formed. A deep learning model (such as a convolutional neural network (CNN) or a long short-term memory network (LSTM)) is used to perform semantic parsing on the audio signal. The extracted features are input into the model for training, and the output of the model is the corresponding audio semantic label (such as emotion, style, etc.). A labeled audio dataset (such as the GTZAN dataset or Emo-DB) is prepared, and training parameters (such as a batch size of 64 and a learning rate of 0.001) are set for model training. After the model training is completed, a new spatially attenuated compensated audio signal is used for inference to generate the semantic features of the audio signal. These features include emotional states (such as happy, sad, angry, etc.) and music styles (such as classical, pop, rock, etc.). The generated audio semantic features are stored in structured data for subsequent analysis and application.

[0017] Step S4: Based on the audio semantic features of the audio signal, the user's audio emotion is speculated, and the evolution prediction of the audio demand is carried out, so as to generate the user's real-time audio demand prediction data; In this embodiment, a suitable emotion classification model is selected, such as a support vector machine (SVM), a random forest, or a deep learning model (such as a convolutional neural network (CNN) or a long short-term memory network (LSTM)). The input of the model is the extracted audio semantic features, and the output is the corresponding emotion label (such as happy, sad, angry, relaxed, etc.). Prepare a labeled audio dataset containing the emotion labels corresponding to each audio sample (such as the Emo-DB or RAVDESS dataset). Set the training parameters, such as the batch size (64), the learning rate (0.001), and the number of training epochs (100 epochs). Use the extracted audio semantic features to train the emotion model. Set the ratio of the training set to the test set to 80:20, and use the cross-validation method to evaluate the performance of the model. Evaluate the effect of the model through metrics such as the confusion matrix, accuracy, and F1 score to ensure that the model can accurately distinguish different emotion states. Perform hyperparameter tuning on the model to improve the accuracy and generalization ability of the model. Use the grid search or random search method to optimize the model parameters. After the model training is completed, use the semantic features of the new audio signal to infer the emotion. Input the features into the model to output the current emotion state of the user. Record the results of each inference, including the timestamp, the inferred emotion type, and its confidence level for subsequent analysis and optimization. Select a suitable time series prediction model, such as an autoregressive integrated moving average (ARIMA), a long short-term memory network (LSTM), or a deep learning model for time series. The input of the model is the user's emotion state and historical audio demand data, and the output is the prediction of future audio demands. Convert the user's emotion state into numerical features (such as emotion intensity scores), and combine them with the historical audio demand data to form a time series dataset. Train the demand evolution model, setting the training parameters, such as the learning rate (0.001), the batch size (64), and the number of training epochs (100 epochs). Use the historical audio demand data and emotion state data for training to ensure that the model can capture the relationship between emotion and demand. Use the cross-validation method to evaluate the performance of the model to ensure that the model has good prediction ability on unseen data. The evaluation metrics include the mean squared error (MSE), the root mean squared error (RMSE), and the mean absolute error (MAE). After the model training is completed, use the real-time user emotion state and historical demand data to predict the audio demand. Input these data into the model to generate the prediction results of future audio demands. Record the results of each prediction, including the prediction timestamp, the predicted audio content type (such as music type, volume demand, etc.), and its prediction confidence level.

[0018] Step S5: Perform all-round environmental noise perception on the spatially attenuated compensated audio signal, and then perform adaptive audio parameter regulation according to the user's real-time audio demand prediction data, so as to generate a dynamically regulated audio signal; In this embodiment, a multi-channel microphone array (e.g., 8 or more omnidirectional microphones) is configured at different angles to achieve 360-degree ambient sound collection, ensuring that the height and spacing of the microphone positions can cover the entire space and avoid blind spots. The sampling rate of each microphone is set to 48 kHz to improve the clarity and details of the collected audio. An audio recording software (such as Audacity or MATLAB) is used to simultaneously capture the ambient noise signals of each microphone, ensuring the synchronization of the signals of each microphone for subsequent processing. A sampling time window (such as 5 seconds) is set, and multiple samples are taken within this time to ensure sufficient noise data is obtained for analysis. The collected noise signals are analyzed in the time domain and frequency domain. The short-time Fourier transform (STFT) is used to extract the spectral characteristics of the noise, and the noise intensity and frequency distribution in each direction are calculated. Thresholds are set to identify and classify different types of ambient noise (such as traffic noise, human voices, natural sounds, etc.). The analysis results are visualized for further understanding and processing. An adaptive audio parameter regulation model is designed to respond to the user's needs and ambient sound in real time. The inputs of the model include the user's emotional state, audio demand prediction data, and current ambient noise characteristics. The audio parameters to be regulated are determined, such as volume, equalizer settings, sound effect modes, etc. Regulation strategies are set, for example, increasing the volume in a noisy environment or decreasing the volume in a quiet environment. The user's real-time audio demand prediction data is integrated with the omnidirectional ambient noise perception data through weighted average or other methods to determine the current optimal audio parameter settings. A machine learning model (such as random forest or support vector machine) is used to train the input data to predict the best audio regulation strategy. The integrated data is input into the audio regulation model to generate real-time audio parameter settings, and these parameters are applied in real time through an audio processing software (such as PureData or Max / MSP) to ensure the optimization of the user's audio experience. The parameter changes during the regulation process are recorded, such as the amplitude of volume increase during the ambient noise peak or the dynamic adjustment of equalizer settings, to ensure the accuracy and rapid response of the regulation. The regulated audio parameters are applied to the audio signal with spatial attenuation compensation for real-time reconstruction. An audio processing library (such as PyAudio or Librosa) is used for signal synthesis and output. According to the user's feedback and the changes in the regulation parameters, the output of the audio signal is adjusted in real time to ensure the optimization of the signal quality and user experience. A user feedback mechanism is established to collect the user's experience feedback in different environments to optimize the parameters and strategies of the regulation model. An online questionnaire or a real-time scoring system is adopted to ensure the timely response to the user's sound experience. The regulation model is updated and trained regularly to better adapt to the changes in the user's needs and different ambient noise characteristics, forming a closed-loop optimization system.

[0019] Step S6: Perform real-time playback processing based on the dynamically regulated audio signal, make an intelligent playback control decision, and construct an intelligent playback control model.

[0020] In this embodiment, the dynamically regulated audio signal is converted in format to ensure its compatibility with the audio playback device. Commonly used audio formats such as WAV or FLAC are selected to maintain the high fidelity of the audio. Audio processing software (such as Audacity or MATLAB) is used to preprocess the audio signal, including denoising, normalization, and volume equalization, to ensure the clarity and consistency of the signal. A digital audio workstation (DAW) or an audio player (such as VLC or a custom player) is used for real-time playback settings, ensuring that the player supports the dynamic regulation function and can respond in real time to changes in audio parameters. Playback parameters such as the sampling rate (e.g., 48 kHz) and buffer size are set to ensure low-latency audio playback. The buffer size is set to 256 samples to reduce playback latency without affecting the sound quality. During playback, the changes in the audio signal are monitored in real time, and the playback parameters are dynamically adjusted according to the ambient noise and user feedback. An audio processing library (such as PyAudio) is used to implement the real-time processing of the audio signal. The changes in various parameters during the real-time playback process are recorded, including volume, equalizer settings, and dynamic range, for subsequent analysis and optimization. User feedback data during the real-time playback process is collected, including the user's satisfaction with the audio signal, preferred types, and volume requirements, etc. Feedback information is collected through an online questionnaire or a real-time scoring system. The user feedback data is combined with the ambient noise information to form a multi-dimensional data set. The changing needs of users under different environmental conditions are analyzed, and key features such as the correlation between emotional states and audio preferences are extracted. A suitable machine learning model (such as a random forest, support vector machine, or deep learning model) is selected for intelligent decision-making. The input of the model is the real-time data features collected, and the output is the optimized audio playback strategy. A labeled data set is prepared to ensure that the model can learn the relationship between user preferences and audio parameters. Training parameters such as the learning rate (0.001), Batch size (64) and number of training epochs (100 epochs), use the collected multi-dimensional dataset to train the intelligent playback control model, adopt the K-fold cross-validation method to evaluate the model performance, ensure its generalization ability on unseen data, the evaluation metrics include accuracy, recall, and F1 score, etc., optimize the hyperparameters of the model, use grid search or random search methods to find the best parameter settings to improve the decision-making accuracy, integrate the trained intelligent playback control model into the audio playback system, ensure that the model can receive and process data in real time, use the API interface to connect the audio processing module and the intelligent control module for easy data transmission and processing, set the real-time feedback mechanism of the model to ensure that the model can adjust the playback strategy in a timely manner according to new user feedback, adopt a message queue or a stream processing framework (such as Kafka or Apache Flink) to implement real-time data processing, during each audio playback process, perform corresponding audio parameter adjustments according to the output of the intelligent playback control model, when the user's mood is "relaxed", automatically adjust the volume and equalizer settings to enhance comfort, record the results of each decision execution, including user feedback, playback effect, and parameter changes, to form a feedback loop for subsequent model optimization and adjustment.

[0021] In this embodiment, refer to Figure 2 , which is a schematic diagram of the detailed implementation steps of step S1. In this embodiment, the detailed implementation steps of step S1 include: Step S11: Perform real-time playback control based on the audio player, and collect real-time playback audio signals and spatially multi-angular reflected audio signals; Step S12: Calculate the signal reflection time for each angle of the spatially multi-angular reflected audio signals to obtain the reflection time for each spatial angle; Step S13: Estimate the distance of the spatial reflection object based on the reflection time for each spatial angle, thereby generating the distance of the reflection object for each angle in the space; Step S14: Analyze the spatial object distribution of the distance of the reflection object for each angle in the space to generate spatial object distribution data; Step S15: Locate the spatial position of each object in the spatial object distribution data to generate the position coordinates of each object in the space; Step S16: Perform three-dimensional audio spatial morphology modeling on the position coordinates of each object in the space to construct a three-dimensional audio spatial morphology model.

[0022] In this embodiment, a library that supports real-time audio playback (such as PyAudio or WebAudio API) is used to implement the playback of audio signals. The parameters of the audio player are configured, such as the sampling rate (e.g., 48 kHz), bit depth (e.g., 16 bits), and number of channels (e.g., stereo). According to the audio content to be played, a playlist is designed, and some specific test audio signals (such as white noise or pulse signals) are selected for playback to facilitate subsequent reflection analysis. A multi-channel audio interface (such as a USB audio interface) is used for signal acquisition to ensure the simultaneous acquisition of the real-time played audio signal and its reflection signal. The audio acquisition software is configured to obtain the reflection signals at different angles in real time, and the acquisition parameters are set, such as the sampling rate and buffer size (e.g., 1024 samples), to ensure that the signals can be recorded with high quality. During audio playback, multiple microphone arrays (such as circular or linear arrays) are used to capture the reflection signals at various angles in space. The placement position of each microphone should be carefully designed to cover all important directions in the target space. Real-time data stream processing is implemented, and the acquired audio signals are stored in a memory buffer for subsequent analysis. The multi-angle reflected audio signals acquired are preprocessed, including noise removal and normalization, to improve the accuracy of subsequent analysis. A signal processing library (such as SciPy) is used for filtering and signal enhancement. The short-time Fourier transform (STFT) is applied to the signals of each channel to extract frequency-domain features for better identification of the reflection signals. The window length is set to 1024 samples, and the overlap rate is 50%. By calculating the delay time between the direct audio signal and the reflected audio signal, the reflection time of each spatial angle is obtained. During the analysis process, the cross-correlation function is used to identify the time delay between the signals, R(τ)= , where \(x(t)\) is the direct signal, \(y(t)\) is the reflected signal, \(\tau\) is the delay time. Record the reflection time of each microphone to form an array of reflection times for subsequent use. Determine the speed of sound propagation in air, which is usually 343 meters per second at 20°C and needs to be fine-tuned according to different environmental conditions. According to the reflection time and the speed of sound propagation, use the following formula to calculate the distance \(d\) of the reflecting object at each spatial angle: \(d = v\cdot t / 2\), where \(d\) is the distance, \(v\) is the speed of sound propagation, and \(t\) is the reflection time. Note that since the sound wave needs to travel back and forth, it needs to be divided by 2. Based on the distance of the reflecting object, construct a three-dimensional spatial distribution model. Combine each distance value with the corresponding spatial angle to form a spatial coordinate point. The calculation method is: \((x,y,z)=(d\cdot\cos(\theta),d\cdot\sin(\theta),0)\), where \(\theta\) is the spatial angle and \(d\) is the object distance. Use a three-dimensional visualization tool (such as Matplotlib or Mayavi) to display the spatial object distribution map. Plot the distribution data of the reflecting objects in a three-dimensional coordinate system to more intuitively observe the spatial distribution characteristics of the objects. Conduct statistical analysis of the spatial object distribution. Record the number of objects and the distribution density in each direction, and generate a distribution histogram or heat map as needed to facilitate understanding of the object distribution in space. Design an object positioning algorithm. Calculate the spatial position of each reflecting object one by one through the reflection time and the known spatial angle. Use triangulation to infer the accurate position of the object through the reflected signals at different angles. For each angle, combine the distance of the reflecting object and the corresponding angle, and calculate the three-dimensional coordinates of the object according to the following formula: \((x,y,z)=(d\cdot\cos(\alpha)\cdot\cos(\beta),d\cdot\cos(\alpha)\cdot\sin(\beta),d\cdot\sin(\alpha))\), where \(\alpha\) is the elevation angle, \(\beta\) is the azimuth angle, and \(d\) is the distance of the reflecting object. Select a suitable three-dimensional modeling tool or library (such as Blender, Open3D or Unity) to model the three-dimensional audio space. According to the requirements, set the level of detail and visual effects of the model. According to the position coordinates of the objects, convert the position of each object in the three-dimensional space into a point in the model. Assign attributes (such as color, size, etc.) to each object to clearly identify it in the three-dimensional space. Use the API of the three-dimensional modeling tool to implement the creation and rendering of the model, ensuring that the position, shape and visual effects of each object in space can be correctly presented. Combine the audio signal information with the three-dimensional model to create an interactive experience of the audio space, and implement spatial audio rendering technologies such as HRTF (Head-Related Transfer Function) or Ambisonics to simulate a real audio experience in the three-dimensional model. Output the final three-dimensional model and conduct tests to ensure the compatibility and performance of the model on different devices.

[0023] In this embodiment, refer to Figure 3, which is a schematic diagram of the detailed implementation steps of step S2. In this embodiment, the detailed implementation steps of step S2 include: Step S21: Calculate the audio reflection amplitude of the spatially multi-angle reflected audio signal and extract the audio reflection amplitude value of each angle; Step S22: Conduct in-depth spectral mining on the spatially multi-angle reflected audio signal to generate reflected audio spectral features; Step S23: Perform step-by-step spatial angle reflection attenuation evolution based on the audio reflection amplitude value of each angle and the reflected audio spectral features to obtain the audio reflection attenuation feature of each angle; Step S24: Fit the audio reflection attenuation feature of each angle to the three-dimensional audio space form model for spatial audio attenuation to construct a three-dimensional audio space attenuation field; Step S25: Perform dynamic attenuation compensation calculation on the three-dimensional audio space attenuation field to generate the dynamic attenuation compensation parameter of each angle; Step S26: Perform spatial audio attenuation correction on the real-time played audio signal based on the dynamic attenuation compensation parameter of each angle to obtain a spatially attenuated compensated audio signal.

[0024] In this embodiment, the multi-angle reflected audio signal collected needs to be preprocessed, including denoising and normalization. A high-pass filter (such as a Butterworth filter) is used to remove low-frequency noise, and the cut-off frequency is set to 100 Hz to ensure the clarity of the signal. The audio signals of each angle are normalized to ensure that the amplitudes are within the same range. This is achieved by maximum value scaling of the signal. The short-time Fourier transform (STFT) is used to convert the time-domain signal into a frequency-domain signal for easy amplitude calculation. The window length is set to 1024 samples and the overlap rate is 50%. Calculate the amplitude spectrum of each frequency component. A(f)= , where R(f) is the frequency-domain signal and R∗(f) is its complex conjugate. Extract the amplitude value of the audio reflection signal of each angle to obtain the reflection amplitude value of each spatial angle. Record the amplitude value of each angle in a list for subsequent analysis. Extract the amplitude value of the audio reflection signal of each angle to obtain the reflection amplitude value of each spatial angle. Record the amplitude value of each angle in a list for subsequent analysis. Use STFT to perform spectral analysis on the audio reflection signal of each spatial angle. Select an appropriate window function (such as a Hamming window) to improve the spectral performance. Calculate the amplitude spectrum and phase spectrum of each frequency component for subsequent mining. Extract spectral features such as power spectral density (PSD), spectral centroid, spectral width, etc. Use functions in the librosa library for feature extraction and store the extracted features in a feature array to form a spectral feature set for each angle. Select a suitable attenuation model (such as an exponential attenuation model) to describe the attenuation characteristics of sound waves in space. A(d)= where A(d) is the attenuated amplitude, 0 is the initial amplitude, α is the attenuation coefficient, and d is the distance. Determine the grid structure of the 3D audio spatial form model and select a suitable 3D modeling tool (such as Blender or Open3D) for fitting. Determine the coordinate system and unit of the space to ensure the accuracy of the model. Apply the attenuation characteristics at each angle to the corresponding positions in the 3D space model. Smoothly map the attenuation characteristics into the space model through an interpolation method (such as cubic spline interpolation) to form a complete attenuation field. Calculate the corresponding attenuation value for each angle and use color or transparency to represent the attenuation intensity to form the visualization effect of the spatial attenuation field. Store the constructed 3D audio spatial attenuation field as a 3D model file (such as OBJ or STL) and perform visual display. Use a visualization tool to display the model effect and ensure its compatibility on different devices. Design a dynamic attenuation compensation model, considering the impact of time variation on attenuation. Use a dynamically adjusted attenuation model and set a time constant (such as 0.5 seconds) to calculate the compensation parameters in real time. For each spatial angle, calculate the compensation parameters according to the attenuation value of the real-time audio signal. C(d,t)=A(d,t)+ΔA(d,t), where C(d,t) is the compensated amplitude, ΔA(d,t) is the real-time compensation value, d is the distance, and t is the time. Apply the dynamic attenuation compensation parameters in the processing of the real-time played audio signal. Use an audio processing library (such as PyDub or Soundfile) to implement the dynamic adjustment of the signal. The processed audio signal should be played in real time to ensure that users can feel the improvement of audio quality. The pyaudio library can be used to implement real-time audio playback while monitoring the processing delay. Track the quality metrics of the audio signal (such as signal-to-noise ratio SNR, total harmonic distortion THD, etc.) through a real-time monitoring system (such as Prometheus) to ensure the effectiveness of the compensation process.

[0025] In this embodiment, refer to Figure 4 , which is a schematic diagram of the detailed implementation steps of step S3. In this embodiment, the detailed implementation steps of the said step S3 include: Step S31: Analyze the audio rhythm fluctuation of the spatially attenuated compensated audio signal to generate audio signal rhythm fluctuation characteristics; Step S32: Identify the audio tone change of the spatially attenuated compensated audio signal and extract the audio tone change characteristics; Step S33: Perceive the audio emotional color according to the audio tone change characteristics, thereby generating audio emotional color perception data; Step S34: Perform in-depth audio semantic parsing based on the audio emotional color perception data and the audio signal rhythm fluctuation characteristics, thereby obtaining the audio semantic characteristics of the audio signal.

[0026] In this embodiment, the spatially attenuated compensated audio signal is preprocessed, including noise removal and normalization. A band-pass filter (such as a Butterworth filter) is used to remove low-frequency and high-frequency noise to ensure the accuracy of subsequent analysis. The frequency range of the filter is set, for example, the low cut-off frequency is 20 Hz and the high cut-off frequency is 4000 Hz to retain the main components of the audio signal. An audio processing library (such as Librosa) is used to extract the rhythm features of the audio signal. The librosa.beat.beat_track() function is used to detect the beats of the audio signal, and parameters such as start_bpm are set to 60 to adapt to different types of audio signals. The timestamps of each beat are recorded, and the time intervals between adjacent beats are calculated to generate rhythm fluctuation features. These features include beat frequency, beat intensity, and beat change rate, etc. The extracted rhythm fluctuation features are stored in structured data (such as a Pandas DataFrame) for subsequent analysis and processing. Visualization tools such as Matplotlib are used to plot the rhythm fluctuation diagram to help understand the rhythm changes of the audio signal. Speech signal processing techniques are used to extract the acoustic features of the audio, such as pitch, volume, and prosody features. The librosa.pitches() in the librosa library is used to extract the pitch features, and the window length of each frame is set to 2048 samples with an overlap rate of 50%. At the same time, the change in volume is calculated by calculating the short-time energy to obtain the loudness change of the audio signal. Machine learning algorithms (such as support vector machines or random forests) are applied to train the extracted acoustic features to identify the tone changes in the audio. A suitable training dataset (such as EMO-DB or CREMA-D) is selected, and training parameters such as batch size and learning rate are set. After training, the model is used to infer the spatially attenuated compensated audio signal to identify the tone change features (such as happy, sad, or angry). The identified tone change features are recorded in structured data, including the change category and its corresponding timestamp, for subsequent analysis and application. An emotional color model is designed to map the tone changes to the emotional dimension. An emotion analysis model (such as Russell's emotion circle or Plutchik's emotion wheel) is used to define the emotional color. The relationship between different tone changes and emotional colors is set, for example, mapping "happy" to "positive" and "sad" to "negative". Based on the identified tone change features, the emotional color at each moment is calculated. Combining the tone change intensity with the emotion model, emotional color perception data is generated. For a happy tone change, its emotional intensity is marked as 0.8, while for sadness it is marked as 0.2. The calculated emotional color perception data is stored in structured data, recording the emotional state and intensity of each time period. Visualization tools such as Plotly or Seaborn are used to display the emotional color perception data.Generate a graph showing the change of emotional color over time to facilitate the analysis and understanding of the emotional characteristics of audio signals. Select a suitable deep learning model (such as Long Short-Term Memory Network LSTM) to perform deep semantic parsing on the audio signal. Set the model input as the emotional color perception data and the rhythm fluctuation characteristics, and the output as the semantic characteristics of the audio (such as theme, emotion, etc.). When training the model, use the labeled audio dataset and set the training parameters such as the learning rate (0.001), batch size (64), and number of training epochs (100 epochs). Apply the trained deep semantic parsing model to the audio signal with spatial attenuation compensation to extract the semantic characteristics of the audio signal. These characteristics include the emotional state, theme content, and related emotional color of the audio. Use the features output by the model for post-processing to ensure the accuracy of the extraction results.

[0027] In this embodiment, step S4 includes the following steps: Step S41: Based on the audio semantic characteristics of the audio signal, speculate on the user's audio emotion to generate the user's audio emotion characteristics; Step S42: Based on the user's audio emotion characteristics, analyze the current motion state to generate the user's current motion state characteristics; Step S43: Predict the evolution of the audio demand for the user's current motion state characteristics, thereby generating the user's real-time audio demand prediction data.

[0028] In this embodiment, a suitable emotion inference model is selected, such as a deep learning model based on a convolutional neural network (CNN) or a long short-term memory network (LSTM). The input of the model is the extracted audio semantic features, and the output is the probability distribution of different emotion states (such as happy, sad, angry, etc.). A well-labeled training dataset is prepared, such as Emo-DB or CREMA-D. Training parameters such as batch size (64) and learning rate (0.001) are set, and data preprocessing is performed to ensure the standardization of the input data. When training the model, the cross-entropy loss function is used to evaluate the gap between the model output and the true emotion labels. During the model training process, an early stopping strategy (EarlyStopping) is adopted to prevent overfitting, and the performance of the validation set is checked regularly. After training, the held-out test set is used to evaluate the model to ensure that the model has a high accuracy and recall rate. The classification effect of the model can be visualized through a confusion matrix. The trained model is applied to new audio signals to infer the audio emotion features of the user, and the output is the probability value of each emotion state. And the emotion with the highest probability value is selected as the current emotion state of the user. The inferred audio emotion features of the user are recorded in structured data for subsequent analysis. A motion state analysis model is designed, and machine learning algorithms (such as random forest or support vector machine) are used to infer the motion state of the user (such as stationary, walking, running, etc.). The input of the model is the audio emotion features of the user and is combined with other sensor data (such as accelerometer and gyroscope). A diverse training dataset is collected, including audio emotion features and sensor data in different motion states, ensuring the representativeness of the data. During the training process, the audio emotion features and sensor data are fused to form a comprehensive feature vector. Training parameters such as learning rate (0.01) And the batch size (32), use cross-validation techniques to evaluate the model performance, ensure that the model can accurately distinguish different motion states, use the ROC curve and AUC value to judge the classification ability of the model, apply the trained model to the real-time data stream, infer the user's current motion state, record the motion state characteristics of each time period, including the motion type and state changes, store the motion state characteristics in structured data for subsequent analysis and decision support, design a demand prediction model, use time series prediction methods (such as ARIMA or LSTM) to analyze the changes in the user's audio needs, the input of the model is the user's motion state characteristics and historical audio demand data, collect the audio demand data of the user in different motion states for training the model, ensure the diversity and continuity of the data, set an appropriate time window (such as 5 minutes) during the training process to capture the user's demand changes, use the mean squared error (MSE) as the loss function to evaluate the model performance, adjust the model parameters (such as the learning rate and the number of hidden layer units) to optimize the prediction effect, use 80% of the historical data for training and 20% for validation to ensure the generalization ability of the model on unseen data, apply the trained model to the real-time data stream, input the user's current motion state characteristics, predict the future audio needs, when the user is running, they need faster-paced music, while when stationary, they prefer light music, the generated audio demand prediction data should include the demand type and intensity within a certain period in the future and be stored in structured data for subsequent analysis and application.

[0029] In this embodiment, step S5 includes the following steps: Step S51: Perform an omni-directional environmental noise perception on the spatially attenuated compensated audio signal to generate an omni-directional environmental noise signal; Step S52: Perform an environmental noise frequency identification on the omni-directional environmental noise signal to obtain the environmental noise frequency characteristics; Step S53: Perform a precise analysis of the environmental scene based on the environmental noise frequency characteristics to generate the real-time audio environmental scene type; Step S54: Perform an adaptive audio parameter regulation based on the user's real-time audio demand prediction data and the real-time audio environmental scene type to generate a dynamically regulated audio signal.

[0030] In this embodiment, a multi-channel microphone array (such as a circular or stereo array) is used to collect ambient sound omnidirectionally, ensuring that the microphone arrangement can cover a 360-degree environment. High-quality microphones are selected to reduce the interference of background noise. The audio acquisition device is configured with a sampling rate of 48 kHz to ensure high-fidelity signals. The buffer size is set to 1024 samples for easy real-time processing and analysis. During the omnidirectional microphone acquisition process, the ambient noise signal is monitored and recorded in real time. Through audio processing software (such as Audacity or MATLAB), the audio signal is monitored in real time. Real-time signal processing techniques are used to extract the features of the acquired audio signal, generating an omnidirectional ambient noise signal. The short-time Fourier transform (STFT) is used to extract the spectral features of the audio signal for subsequent analysis. The generated omnidirectional ambient noise signal is stored in a database for subsequent processing. The noise signal features in each direction can be recorded using a structured data format (such as JSON or CSV). A visualization tool (such as Matplotlib) is used to plot the heat map of the ambient noise to visually display the noise intensity distribution in different directions. The fast Fourier transform (FFT) is used to perform spectral analysis on the omnidirectional ambient noise signal. Through the FFT, the time-domain signal is converted into a frequency-domain signal to identify the frequency components of the ambient noise. The window size of the FFT is set to 2048 samples, and the overlap rate is 50% to balance the frequency resolution and time resolution. The FFT processing is performed on the ambient noise signal in each direction to extract frequency features, including the main frequency, spectral energy distribution, and frequency bandwidth, etc. The amplitude and phase of each frequency component are calculated to identify the main frequency features of the noise, and the intensity of each frequency component is recorded. The extracted ambient noise frequency features are stored in structured data for subsequent analysis. The main frequency and its corresponding intensity value in each direction are recorded, and a visualization tool is used to plot the spectrogram to help analyze and understand the ambient noise features in different directions. An environmental scene analysis model is designed, and machine learning algorithms (such as decision trees or random forests) are used to classify different environmental scenes (such as urban, natural, indoor, etc.). The input of the model is the extracted ambient noise frequency features. A training data set of multiple environmental scenes is collected to ensure the diversity and representativeness of the data. The features of each scene are labeled. During the training process, training parameters are set, such as the learning rate (0.01) The number of trees (such as 100 trees), use the cross-validation method to evaluate the model performance, ensure that the model can accurately distinguish different environmental scenarios. After training, use the validation set to evaluate the accuracy and recall rate of the model, use the confusion matrix to visualize the classification effect, apply the trained model to the real-time collected environmental noise frequency characteristics, infer the current audio environmental scenario type, record the environmental scenario type and its confidence at each moment, design an audio parameter regulation model, determine the relationship between the key parameters of the audio signal (such as volume, equalizer settings, sound effect modes, etc.) and user needs and environmental scenarios, set the regulation strategy, for example, increase the volume in a noisy environment and decrease the volume in a quiet environment, collect historical data of user audio needs, analyze user preferences in different environmental scenarios, establish regulation rules, monitor the user's audio demand prediction data and environmental scenario type in real time, apply the regulation model to generate corresponding audio parameter settings, use a simple rule engine or a complex machine learning model to achieve this process. When the user is in an urban environment and needs high-energy music, increase the volume and low frequency, and at the same time decrease the volume and switch to natural sound effects in a natural environment. Generate a dynamically regulated audio signal according to the regulation result, use an audio processing library (such as PyAudio or Librosa) to output the regulated audio signal in real time to ensure that the user obtains the best audio experience, record the parameter changes during the regulation process, and store the regulation result in the database for subsequent analysis and optimization.

[0031] In this embodiment, step S6 includes the following steps: Step S61: Perform real-time playback processing based on the dynamically regulated audio signal and obtain the user's real-time playback feedback; Step S62: Perform deep reinforcement learning on the user's real-time playback feedback to generate user feedback reinforcement learning data; Step S63: Perform personalized user demand evolution based on the user feedback reinforcement learning data to obtain personalized user audio playback demand characteristics; Step S64: Make an intelligent playback control decision on the personalized user audio playback demand characteristics and construct an intelligent playback control model.

[0032] In this embodiment, an audio playback system that supports dynamic regulation (such as VLC or a custom audio player) is used to set the playback parameters of the audio signal, including volume, equalizer settings, and playback mode, ensuring that the player can respond in real time to changes in audio parameters. Set the audio content to be played and select representative tracks so that effective user feedback can be collected. Design a user feedback mechanism to collect feedback in multiple ways, such as through questionnaires, a real-time rating system (such as a 1 to 5-star rating), or emotion recognition technology (such as facial expression recognition), ensuring that the feedback process is simple and intuitive to increase user engagement. Set a feedback button on the player interface so that users can provide feedback on the audio content at any time during playback, record the timestamp of the feedback and the user's selection, and monitor the user feedback data in real time. Store the user feedback and the corresponding playback signal in a database, using a structured data format (such as JSON or CSV) to record the audio signal characteristics of each playback and the user's feedback information, ensuring the integrity and accuracy of the data. Regularly check the data quality to avoid data loss or errors. Design a deep reinforcement learning model. The model structure can adopt a deep Q-network (DQN) or a policy gradient method. The model input is the user's real-time feedback and audio playback characteristics, and the output is the action for each feedback (such as adjusting the volume, switching tracks, etc.). Set the state space, including the user's feedback information, the current audio playback characteristics, and the ambient noise situation. The action space includes audio parameter adjustments. Collect the user's feedback data and use this data to train the reinforcement learning model. Set the training parameters, such as the learning rate (0.001) and the discount factor (0.99), during the training process, the Experience Replay technique is adopted to store past experiences in a buffer and randomly sample to reduce the correlation of training and improve the generalization ability of the model. After the model training is completed, the model is used to infer new user feedback and generate reinforcement learning data for user feedback. This data includes the relationship between user feedback and the actions taken by the model to optimize future playback strategies. To build a demand evolution analysis model, clustering algorithms such as K-means or DBSCAN are used to analyze user feedback data to identify the audio features and demand changes preferred by users. Set clustering parameters such as the number of clusters and the distance metric method to ensure that the model can effectively distinguish the demand characteristics of different users. According to the clustering results, extract the personalized audio playback demand characteristics of each user. These characteristics include the music types, playback styles, and requested audio effects preferred by users. Use statistical methods such as mean and variance to analyze user feedback in different situations and evaluate the changing trends of user demands. Design an intelligent playback control decision model using deep learning algorithms such as LSTM or convolutional neural networks to predict users' audio playback demands. The input of the model is the personalized user audio playback demand characteristics, and the output is the recommended audio playback list or parameter settings. Collect users' historical playback data to ensure that the model can learn users' long-term preferences. During the training process, set training parameters such as the learning rate (0.001) and batch size (64), use the mean squared error (MSE) as the loss function to evaluate the model performance, and adopt the cross-validation method to evaluate the performance of the model on unseen data to ensure the generalization ability and accuracy of the model. Apply the trained intelligent playback control model to the real-time data stream, dynamically generate audio playback strategies according to users' personalized demand characteristics, consider users' environmental conditions and feedback, and intelligently adjust the playback content and parameters. Record the parameter changes during the intelligent playback control process and users' feedback to ensure that the model can optimize itself.

[0033] In this specification, an audio playback control device is provided for performing the audio playback control method described above, including: An audio space module for obtaining a real-time played audio signal and a spatially multi-angled reflected audio signal; calculating the signal reflection time for each angle of the spatially multi-angled reflected audio signal and performing three-dimensional audio space morphology modeling to construct a three-dimensional audio space morphology model; An audio attenuation correction module for performing the attenuation evolution of the spatially multi-angled reflected audio signal for each spatial angle and performing spatial audio attenuation correction based on the three-dimensional audio space morphology model to obtain a spatially attenuated compensated audio signal; An audio semantics module for performing audio rhythm fluctuation analysis on the spatially attenuated compensated audio signal and performing deep audio semantics parsing to obtain the audio semantics features of the audio signal; An audio demand prediction module, which is used to speculate on the user's audio emotion based on the audio semantic features of the audio signal, and conduct audio demand evolution prediction, so as to generate user real-time audio demand prediction data; An audio parameter regulation module, which is used to perform all-round environmental noise perception on the spatially attenuated compensated audio signal, and then perform adaptive audio parameter regulation according to the user real-time audio demand prediction data, so as to generate a dynamically regulated audio signal; An intelligent playback control module, which is used to perform real-time playback processing based on the dynamically regulated audio signal, and make an intelligent playback control decision to construct an intelligent playback control model.

[0034] The present invention also provides a computer device, including a memory and a processor. A computer program is stored in the memory, and when the processor executes the computer program, the steps of the audio playback control method described in any one of the above are implemented.

[0035] The present invention also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the audio playback control method described in any one of the above are implemented.

[0036] Those skilled in the art clearly understand that for the convenience and conciseness of description, the specific working processes of the above-described systems, systems, and units refer to the corresponding processes in the foregoing method embodiments, and will not be described herein again.

[0037] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it is stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present application essentially or the part that contributes to the prior art or all or part of the technical solution is embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (personal computer, server, or network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present application. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM), random access memories (RAM), magnetic disks, or optical discs that store program codes.

[0038] Therefore, from any point of view, the embodiments should be regarded as exemplary and non-limiting. The scope of the present invention is defined by the appended claims rather than the above description. Therefore, it is intended to cover all changes falling within the meaning and scope of the equivalent elements of the application documents within the present invention.

[0039] The above are only specific embodiments of the present invention, enabling those skilled in the art to understand or implement the present invention. Various modifications to these embodiments will be obvious to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to these embodiments shown herein, but rather will be accorded the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. An audio playback control method, characterized in that: The following steps are involved: Step S1: obtaining a real-time audio signal and a spatial multi-angle reflected audio signal; calculating the reflection time of each angle of the spatial multi-angle reflected audio signal and performing audio space three-dimensional morphology modeling on the audio signal to construct a three-dimensional audio space morphology model; Step S2: performing reflection attenuation evolution of the spatial multi-angle reflected audio signal one by one at spatial angles, and performing spatial audio attenuation correction based on a three-dimensional audio space morphology model, thereby obtaining a spatial attenuation compensated audio signal; Step S3: performing audio rhythm fluctuation analysis on the spatial attenuation compensated audio signal, and performing deep audio semantic analysis, so as to obtain audio semantic features of the audio signal; Step S4: inferring the user's audio emotion based on the audio semantic features of the audio signal, and predicting the evolution of the audio demand, thereby generating the user's real-time audio demand prediction data; Step S5: Perform omnidirectional environmental noise perception on the spatial attenuation compensated audio signal, and then perform adaptive audio parameter control according to the user's real-time audio demand prediction data, thereby generating a dynamically controlled audio signal; Step S6: Perform real-time playback processing based on the dynamically regulated audio signal, make intelligent playback control decisions, and build an intelligent playback control model.

2. The audio playback control method according to claim 1, characterized in that: The specific steps of step S1 are: Step S11: performing real-time playback control based on an audio player, and collecting real-time playback audio signals and spatial multi-angle reflected audio signals; Step S12: calculating the reflection time of each angle of the spatial multi-angle reflected audio signal, thereby obtaining the reflection time of each spatial angle; Step S13: Calculate the distance of the spatial reflection object based on the reflection time at each spatial angle, thereby generating the distance of the reflection object at each angle in the space; Step S14: performing spatial object distribution analysis on the distance of the reflection object at each angle in the space to generate spatial object distribution data; Step S15: positioning the spatial position of each object in the spatial object distribution data to generate the position coordinates of each object in the space; Step S16: Perform audio space three-dimensional morphology modeling on the position coordinates of each object in the space to construct a three-dimensional audio space morphology model.

3. The audio playback control method according to claim 1, characterized in that: The specific steps of step S2 are: Step S21: Calculate the audio reflection amplitude of the spatial multi-angle reflected audio signal, and extract the audio reflection amplitude of each angle; Step S22: performing deep spectrum mining on the spatial multi-angle reflected audio signals to generate reflected audio spectrum features; Step S23: performing reflection attenuation evolution at each spatial angle according to the audio reflection amplitude and reflected audio spectrum characteristics at each angle, and obtaining the audio reflection attenuation characteristics at each angle; Step S24: performing spatial audio attenuation fitting on the three-dimensional audio space morphology model for the audio reflection attenuation characteristics at each angle, and constructing a three-dimensional audio space attenuation field; Step S25: performing dynamic attenuation compensation calculation on the three-dimensional audio spatial attenuation field, thereby generating dynamic attenuation compensation parameters for each angle; Step S26: performing spatial audio attenuation correction on the real-time audio signal based on the dynamic attenuation compensation parameter of each angle, thereby obtaining a spatial attenuation compensated audio signal.

4. The audio playback control method according to claim 1, characterized in that: The specific steps of step S3 are: Step S31: performing audio rhythm fluctuation analysis on the spatial attenuation compensated audio signal to generate audio signal rhythm fluctuation features; Step S32: performing audio tone change recognition on the spatial attenuation compensated audio signal to extract audio tone change features; Step S33: performing audio emotion color perception according to the audio tone change characteristics, thereby generating audio emotion color perception data; Step S34: Perform deep audio semantic analysis based on the audio emotional color perception data and the rhythm fluctuation characteristics of the audio signal to obtain the audio semantic features of the audio signal.

5. The audio playback control method according to claim 1, characterized in that: The specific steps of step S4 are: Step S41: inferring the user audio emotion based on the audio semantic features of the audio signal to generate the user audio emotion features; Step S42: Analyze the current motion state based on the user's audio emotion characteristics to generate the user's current motion state characteristics; Step S43: Predict the evolution of audio demand based on the user's current motion state characteristics, thereby generating user real-time audio demand prediction data.

6. The audio playback control method according to claim 1, characterized in that: The specific steps of step S5 are: Step S51: performing omnidirectional environmental noise perception on the spatial attenuation compensated audio signal to generate an omnidirectional environmental noise signal; Step S52: performing environmental noise frequency recognition on the omnidirectional environmental noise signal to obtain environmental noise frequency characteristics; Step S53: accurately analyzing the environment scene based on the frequency characteristics of the environmental noise, thereby generating a real-time audio environment scene type; Step S54: Adaptively adjust audio parameters based on the user's real-time audio demand prediction data and the real-time audio environment scene type, thereby generating a dynamically adjusted audio signal.

7. The audio playback control method according to claim 1, characterized in that: The specific steps of step S6 are: Step S61: performing real-time playback processing based on the dynamically regulated audio signal, and obtaining real-time playback feedback from the user; Step S62: performing deep reinforcement learning on the user's real-time playback feedback to generate user feedback reinforcement learning data; Step S63: Perform personalized user demand evolution based on user feedback reinforcement learning data to obtain personalized user audio playback demand characteristics; Step S64: Make intelligent playback control decisions based on personalized user audio playback demand characteristics and build an intelligent playback control model.

8. An audio playback control device, characterized in that: Used to execute the audio playback control method as claimed in claim 1, comprising: The audio space module is used to obtain real-time audio signals and multi-angle reflected audio signals in space; calculate the reflection time of each angle of the multi-angle reflected audio signals in space and perform three-dimensional audio space morphology modeling to construct a three-dimensional audio space morphology model; An audio attenuation correction module is used to perform spatial angle reflection attenuation evolution on the spatial multi-angle reflected audio signal one by one, and perform spatial audio attenuation correction based on a three-dimensional audio space morphology model, thereby obtaining a spatial attenuation compensated audio signal; An audio semantic module is used to perform audio rhythm fluctuation analysis on the spatial attenuation compensated audio signal and perform deep audio semantic analysis to obtain audio semantic features of the audio signal; The audio demand prediction module is used to infer the user's audio emotions based on the audio semantic features of the audio signal and predict the evolution of audio demand, thereby generating real-time audio demand prediction data for the user; The audio parameter control module is used to perform all-round environmental noise perception on the spatial attenuation compensation audio signal, and then perform adaptive audio parameter control based on the user's real-time audio demand prediction data, thereby generating a dynamically controlled audio signal; The intelligent playback control module is used to perform real-time playback processing based on dynamically regulated audio signals, make intelligent playback control decisions, and build an intelligent playback control model.

9. A computer device comprising a memory and a processor, wherein a computer program is stored in the memory, wherein: When the processor executes the computer program, the steps of the audio playback control method according to any one of claims 1 to 7 are implemented.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the audio playback control method according to any one of claims 1 to 7 are implemented.

Citation Information

Cited By

  • Intelligent music playing method based on three-dimensional scene interaction

    CN120803391A

  • Intelligent music playing method based on three-dimensional scene interaction

    CN120803391B

  • Adaptive control method and device for audio equipment

    CN120972573A