AI-based scene recognition-based intelligent control method and system for digital audio sound field parameters

By combining AI scene recognition and particle swarm optimization, a digital audio sound field parameter intelligent control method has been developed, which solves the problem that existing technologies cannot adapt to complex environments and personalized listening needs, thereby improving sound quality and user experience.

CN122138094APending Publication Date: 2026-06-02SHENZHEN HANKE TECH CO LTD

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
SHENZHEN HANKE TECH CO LTD
Filing Date
2026-03-09
Publication Date
2026-06-02

AI Technical Summary

Technical Problem

Existing digital audio systems cannot dynamically adapt to complex environmental changes and personalized listening needs, resulting in poor sound quality and a poor user experience.

Method used

By acquiring environmental acoustic parameters and user-played content through AI scene recognition, and combining this with an auditory demand analyzer to generate optimized sound field control targets, the control parameters of the digital audio system are iteratively optimized using a particle swarm optimization algorithm to achieve intelligent control.

Benefits of technology

It achieves intelligent control throughout the entire process, from environmental perception to dynamic optimization and control, improving sound quality and user experience, and ensuring rapid convergence to the optimal solution in different scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122138094A_ABST
    Figure CN122138094A_ABST
Patent Text Reader

Abstract

This application discloses an AI-based method and system for intelligent control of digital speaker sound field parameters, relating to the field of digital professional audio equipment technology. The method includes: acquiring environmental acoustic parameters and user-played content, and identifying the playback scene to obtain the playback scene; performing auditory demand analysis to obtain an auditory demand solution; obtaining an initial sound field control target, and conducting an environmental impact assessment based on the environmental acoustic parameters to obtain an optimized sound field control target; iteratively optimizing the control parameters of the digital speaker in conjunction with the playback scene to obtain an optimized control scheme, and intelligently controlling the digital speaker. This solves the technical problem that existing digital speaker sound field parameter control methods cannot dynamically adapt to complex environmental changes and personalized auditory needs, resulting in poor sound quality and a poor user experience.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of digital professional audio equipment technology, specifically to a method and system for intelligent control of digital audio field parameters based on AI scene recognition. Background Technology

[0002] With the development of digital technology and the popularization of smart terminals, users' demands for sound quality and intelligent control of digital speakers are increasing. However, the sound field parameter control of existing digital speakers mostly relies on manual adjustment by the user, or can only be simply switched according to preset fixed scene modes, which has limitations.

[0003] On the one hand, ordinary users often lack professional acoustic knowledge and find it difficult to adjust various parameters according to the actual environment and playback content, resulting in poor sound quality.

[0004] On the other hand, preset modes cannot fully cover the complex and ever-changing actual environment and personalized auditory needs. They cannot dynamically adapt to changes in environmental acoustic parameters and users' differentiated auditory preferences for different playback content, thus making it difficult to achieve optimal control of the sound field and affecting the user's auditory experience. Summary of the Invention

[0005] This application provides an AI-based method and system for intelligent control of digital audio sound field parameters, which solves the technical problem that existing digital audio sound field parameter control cannot dynamically adapt to complex environmental changes and personalized listening needs, resulting in poor sound quality and user experience.

[0006] The technical solution to the above-mentioned technical problems in this application is as follows: In a first aspect, this application provides a method for intelligent control of digital audio field parameters based on AI scene recognition, the method comprising: Acquire environmental acoustic parameters and user-played content, and identify the playback scene based on the environmental acoustic parameters to obtain the playback scene; Based on the user's playback content and the playback scenario, auditory needs analysis is conducted to obtain auditory needs solutions; Based on the auditory demand scheme, the initial sound field control target is obtained, and based on the environmental acoustic parameters, an environmental impact assessment is conducted to obtain the optimized sound field control target. Based on the aforementioned sound field control objective and combined with the playback scenario, the control parameters of the digital audio system are iteratively optimized to obtain an optimized control scheme and intelligently control the digital audio system.

[0007] Secondly, this application provides an AI-based scene recognition-based intelligent control system for digital audio field parameters, including: The information acquisition module is used to acquire environmental acoustic parameters and user playback content, and to identify the playback scene based on the environmental acoustic parameters. The solution acquisition module is used to perform auditory demand analysis based on the user's playback content and the playback scenario, and to obtain auditory demand solutions. The environmental assessment module is used to obtain initial sound field control targets based on the auditory demand scheme, and to conduct environmental impact assessments based on environmental acoustic parameters to obtain optimized sound field control targets. The scheme optimization module is used to iteratively optimize the control parameters of the digital audio system based on the optimized sound field control target and the playback scenario, obtain an optimized control scheme, and intelligently control the digital audio system.

[0008] This application provides one or more technical solutions, which have at least the following technical effects or advantages: This application provides an AI-based method and system for intelligent control of digital speaker sound field parameters based on AI scene recognition. First, it acquires environmental acoustic parameters and the user's playback content, and identifies the current playback scene based on the environmental acoustic parameters. Second, using a trained auditory demand analyzer, it analyzes the user's auditory preferences and needs for the playback content in that scene, combined with the identified playback scene, thereby generating a targeted auditory demand solution. Then, based on the auditory demand solution, it determines an initial sound field control target, and then assesses the environmental impact of the initial target using environmental acoustic parameters, further optimizing the initial target to obtain an optimized sound field control target that better matches the actual environment. Finally, guided by the optimized sound field control target, and considering the playback scene, iteratively optimizes the control parameters of the digital speaker using a particle swarm optimization algorithm. Simultaneously, based on the typicality of the playback scene and the characteristics of the auditory demand solution, iterative optimization round thresholds and fitness thresholds are set to ensure rapid convergence to the optimal solution in different scenes, ultimately obtaining an optimized control scheme and intelligently controlling the digital speaker.

[0009] Through the above technical solution, this application realizes intelligent control of the entire process from environmental perception and demand analysis to dynamic optimization and control, effectively overcoming the limitations of traditional control methods and improving sound quality and user experience. Attached Figure Description

[0010] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0011] Figure 1This is a flowchart illustrating the intelligent control method for digital audio field parameters based on AI scene recognition provided in this application embodiment; Figure 2 This is a schematic diagram of the structure of the AI ​​scene recognition digital audio field parameter intelligent control system provided in the embodiments of this application.

[0012] The components represented by each number in the attached diagram are explained below: Information acquisition module 11, solution acquisition module 12, environmental assessment module 13, and solution optimization module 14. Detailed Implementation

[0013] This application provides an AI-based method and system for intelligent control of digital audio sound field parameters based on scene recognition. This system addresses the technical problem that existing digital audio sound field parameter control methods cannot dynamically adapt to complex environmental changes and personalized listening needs, resulting in poor sound quality and user experience.

[0014] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0015] In the description of this application, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Therefore, a feature defined as "first" or "second" may explicitly or implicitly include one or more features. In the description of this application, "multiple" means two or more, unless otherwise explicitly specified.

[0016] In the description of this application, the term "for example" is used to mean "used as an example, illustration, or description." Any embodiment described as "for example" in this application is not necessarily to be construed as being more preferred or advantageous than other embodiments. The following description is provided to enable any person skilled in the art to make and use this application. Details are set forth in the following description for purposes of explanation. It should be understood that those skilled in the art will recognize that this application can be made without using these specific details. In other instances, well-known structures and processes will not be described in detail to avoid unnecessarily obscuring the description of this application. Therefore, this application is not intended to be limited to the embodiments shown, but is consistent with the broadest scope of the principles and features disclosed in this application.

[0017] Example 1, as Figure 1As shown in the embodiments of this application, an intelligent control method for digital audio field parameters based on AI scene recognition is provided, including: S10: Obtain environmental acoustic parameters and user playback content, and identify the playback scene based on the environmental acoustic parameters to obtain the playback scene; In this embodiment, firstly, environmental acoustic parameters and user playback content are acquired. The environmental acoustic parameters can be collected in real time by acoustic sensors deployed on digital audio equipment, specifically including parameters such as environmental noise spectrum, room reverberation time, and sound field uniformity. The user playback content is acquired through an audio input interface or a streaming media playback software interface, including the audio file format, sampling rate, bit rate, content type, and corresponding audio characteristic parameters, such as spectral distribution, dynamic range, and rhythm characteristics.

[0018] Then, playback scene identification is performed based on the acquired environmental acoustic parameters. Specifically, this can be done by matching with a pre-set scene feature library, which contains environmental acoustic parameter feature templates corresponding to various typical playback scenes, such as family living rooms, bedrooms, conference rooms, outdoor squares, and small cinemas. The real-time collected environmental acoustic parameters are compared with the templates in the scene feature library to calculate similarity, and the scene corresponding to the template with the highest similarity is selected as the identified playback scene.

[0019] Specifically, step S10 in the method includes: Acquire environmental acoustic parameters and user-played content; Input the ambient acoustic parameters into the playback scene recognizer to obtain the playback scene.

[0020] In this embodiment of the application, firstly, environmental acoustic parameters are collected by an acoustic sensor array, which includes an omnidirectional microphone and a directional microphone. The omnidirectional microphone is used to collect overall acoustic characteristics such as the ambient noise spectrum and sound field uniformity, while the directional microphone is used to assist in locating the direction of the sound source and judging the influence of spatial structure on sound wave reflection.

[0021] At the same time, the system reads metadata information of the user's playback content, such as audio encoding format, sampling frequency, and number of channels, through the playback control module of the digital audio system or the interface of an external storage device, and extracts the temporal and frequency domain features of the playback content through audio parsing algorithms.

[0022] Secondly, the collected environmental acoustic parameters are standardized and preprocessed to eliminate dimensional differences before being input into the trained playback scene recognizer. The playback scene recognizer is a classification model built on a deep neural network. Its input layer receives the preprocessed environmental acoustic parameter vector, the hidden layer extracts high-dimensional features through convolution and pooling operations, and the output layer uses the softmax activation function to output the probability distribution of each playback scene category. The category with the highest probability value is selected as the final recognition result, i.e., the playback scene is obtained.

[0023] The training process of the playback scene recognizer uses a labeled environmental acoustic parameter sample set for supervised learning. The labels are the corresponding actual playback scene categories. The network parameters are optimized through the backpropagation algorithm until the recognition accuracy of the model on the validation set reaches a preset threshold, such as 95%, to ensure the accuracy and robustness of playback scene recognition.

[0024] The acquisition of the playback scene recognizer includes: Obtain the sample environment acoustic parameter set; The acoustic parameter set of the sample environment is labeled to obtain playback scene labels, which are then integrated into a playback scene label set, which includes various environmental labels. Construct a playback scene recognizer, using sample environmental acoustic parameters as input and playback scene labels as supervision, and train the playback scene recognizer until convergence.

[0025] In this embodiment of the application, firstly, environmental acoustic parameter samples under different typical playback scenarios are collected to construct a sample environmental acoustic parameter set.

[0026] Specifically, data was collected using acoustic sensor arrays of the same model as those actually deployed in various pre-set typical scenarios, such as family living rooms, bedrooms, conference rooms, outdoor plazas, and small cinemas. Environmental acoustic parameters were collected for different time periods and under different environmental interference conditions in each scenario to ensure the diversity and representativeness of the samples. The number of samples was determined based on the complexity of the scenario; for example, no fewer than 500 sets of sample data were collected for each typical scenario.

[0027] Then, each set of sample data in the sample environmental acoustic parameter set is manually labeled and assigned a corresponding playback scene label, forming a playback scene label set. The playback scene labels specifically include various environmental labels such as family living room, bedroom, study, conference room, lecture hall, outdoor square, park, small cinema, KTV room, and car interior, covering common digital audio usage scenarios.

[0028] Next, a playback scene recognizer based on a deep neural network was constructed. The network structure adopted an improved CNN-LSTM hybrid model, namely a convolutional neural network-long short-term memory network model. The CNN part includes convolutional layers, batch normalization layers, and pooling layers, used to extract local spatial features from the spectrogram or temporal waveform of environmental acoustic parameters. The LSTM part includes a forget gate, input gate, and output gate, used to capture the dynamic features of environmental acoustic parameters changing over time, suitable for processing parameters with temporal characteristics such as reverberation time and noise variations. The input layer of the network was set to match the dimension of the preprocessed environmental acoustic parameter vector, and the output layer used the softmax function, with the output dimension corresponding to the number of categories of the playback scene labels.

[0029] Finally, the constructed playback scene recognizer is trained using the sample environment acoustic parameters as input and the corresponding playback scene labels as supervision signals. During training, the difference between the predicted and true labels is calculated using the cross-entropy loss function, and the parameters are updated using the Adam optimizer. The learning rate is initially set to 0.001, and a learning rate decay strategy is adopted, reducing the learning rate to 0.9 times its original value every 10 training epochs.

[0030] Simultaneously, the sample environmental acoustic parameter set is divided into a training set, a validation set, and a test set in a 7:2:1 ratio. The training set is used for model parameter learning, the validation set is used to evaluate model performance and determine early stopping during training (training stops when the validation set accuracy no longer improves after 5 consecutive epochs), and the test set is used to finally evaluate the model's generalization ability. Training continues until the model's accuracy on the validation set reaches a preset threshold (e.g., 95%) and its accuracy on the test set is no less than 92%. At this point, the model converges, and the trained playback scene recognizer is obtained.

[0031] S20: Based on the user's playback content and playback scenario, conduct auditory demand analysis and obtain auditory demand solutions; In this embodiment, after acquiring the user's playback content and the identified playback scenario, auditory demand analysis is performed to generate an auditory demand solution. The auditory demand solution is a comprehensive description of the user's ideal auditory experience for a specific playback content in a specific scenario, including specific demand indicators such as sound quality preference, volume dynamic range, channel balance, low-frequency enhancement, and vocal clarity.

[0032] Specifically, step S20 in the method includes: Obtain an auditory demand analyzer, which is trained based on a set of sample user playback content, a set of sample playback scenarios, and a set of sample auditory demand solutions. Input the user's playback content and playback scenario into the auditory demand analyzer to obtain auditory demand solutions.

[0033] In this embodiment, firstly, an auditory demand analyzer is constructed. This analyzer is a multimodal deep learning model that integrates content features and scene features. Its training data includes a set of sample user playback content, a set of sample playback scenes, and a corresponding set of sample auditory demand schemes.

[0034] Furthermore, the sample user playback content set includes different types of audio data, such as classical music, pop songs, movie soundtracks, news broadcasts, and audiobooks. Each type contains multiple representative audio samples, and their audio features, such as Mel spectrograms, MFCC coefficients, rhythm intensity, and pitch distribution, are extracted. The sample playback scene set refers to the various typical scenes involved in the aforementioned playback scene identification. The sample auditory demand solution set is obtained through questionnaires, user interviews, and annotation by professional acoustic engineers. For each group of sample user playback content and sample playback scene, corresponding auditory demand indicators are labeled.

[0035] Secondly, the audio features of the content played by the sample users and the category features of the sample playback scenes are fused. The category of the sample playback scene is converted into a vector form through one-hot encoding, and concatenated with the audio feature vector of the content played by the sample users as the input of the auditory demand analyzer. The network structure of the auditory demand analyzer adopts an architecture combining a Transformer encoder and a fully connected layer. The self-attention mechanism of the Transformer encoder can capture the long-term dependencies within the audio features, as well as the interaction between audio features and scene features. The output fused feature vector is mapped to the various index dimensions of the auditory demand scheme through the fully connected layer. The model parameters are optimized by comparing it with the sample auditory demand scheme through the mean squared error loss function and using gradient descent.

[0036] During model training, the sample dataset is also divided into training set, validation set and test set. The model's prediction error for auditory demand indicators is monitored through the validation set. When the error is lower than the preset threshold, such as the average absolute error of each indicator being less than 0.5dB or 5%, training is stopped, and a trained auditory demand analyzer is obtained.

[0037] S30: Based on the auditory demand solution, obtain the initial sound field control target, and based on the environmental acoustic parameters, conduct an environmental impact assessment to obtain the optimized sound field control target; In this embodiment of the application, based on the auditory demand scheme, it is first transformed into a specific initial sound field control target. This target covers various core parameters of digital audio, including frequency equalization curve (EQ), volume reference value, channel separation, reverberation depth, bass enhancement coefficient, and vocal clarity compensation value.

[0038] For example, if the auditory requirements scheme requires "clear and prominent vocals with full low frequencies", the initial sound field control target can be set as follows: increase the gain of the 300-3000Hz frequency band by 2-3dB, increase the gain of the 80-200Hz low frequency band by 1-2dB, and adjust the reverberation depth to 0.8-1.2s to avoid blurring of vocals.

[0039] Subsequently, an environmental impact assessment is conducted based on the environmental acoustic parameters obtained in step S10 to revise the initial sound field control target and obtain an optimized sound field control target. The core of the environmental impact assessment lies in analyzing the interference and attenuation effects of the current environmental acoustic parameters on the initial sound field control target. For example, an excessively long room reverberation time may cause high-frequency signals to attenuate too quickly, and specific frequency peaks in the environmental noise spectrum may mask the details of the played content.

[0040] Finally, through multi-factor weighted comprehensive correction, an optimized sound field control target adapted to the current environment is generated.

[0041] Specifically, step S30 in the method includes: Based on the auditory demand scheme, the initial sound field control target is obtained, which includes the sound pressure target and the reverberation time target. An environmental impact assessment is conducted based on environmental acoustic parameters to obtain the degree of environmental impact. The environmental impact assessment is based on the comprehensive impact of environmental parameters on reverberation time targets, sound pressure level control targets, and frequency band gain control targets. Based on the degree of environmental impact, the initial sound field control target is optimized to obtain the optimized sound field control target.

[0042] In this embodiment, firstly, based on the auditory requirements scheme, it is broken down into quantifiable initial sound field control targets, specifically including sound pressure level (SPL) targets and reverberation time targets. The SPL target refers to the desired SPL level across various frequency bands of the playback content. For example, for playback content primarily featuring human voices, the SPL target for the 1kHz-4kHz band is set to 75±2dB to ensure the audibility of human voices; for classical music, the dynamic range target for the SPL across the entire 20Hz-20kHz band might be set to above 80dB to preserve rich musical details. The reverberation time target is determined based on the playback scenario and content type. For example, when playing light music in a family bedroom, the reverberation time target can be set to 0.3-0.5s to create a clear and compact listening experience; while when playing movie soundtracks in a small cinema setting, the reverberation time target can be appropriately extended to 0.8-1.0s to enhance spatial immersion.

[0043] Secondly, an environmental impact assessment is conducted based on environmental acoustic parameters to obtain the degree of environmental impact. The environmental impact assessment analyzes the comprehensive impact of environmental parameters on reverberation time targets, sound pressure level control targets, and frequency band gain control targets.

[0044] In the specific assessment process, an influence model between environmental parameters and various control targets is established, and each influencing factor is assigned a weight. For example, the weight of reverberation time difference is 0.3, the weight of the difference between the peak value of the noise spectrum and the target sound pressure level is 0.4, the weight of sound field uniformity deviation is 0.2, and the weight of the influence of temperature and humidity on sound speed is 0.1. The comprehensive environmental impact value is obtained by weighted summation. The value ranges from 0 to 1. The larger the environmental impact value, the more significant the influence of the environment on the initial control target.

[0045] Finally, based on the calculated environmental impact, the initial sound field control target is optimized to obtain the optimized sound field control target. The optimization process employs a dynamic adjustment algorithm, which compensates for or corrects parameters such as sound pressure level, reverberation time, and frequency band gain in the initial target according to the direction and degree of influence of different environmental parameters.

[0046] For example, when the ambient noise is high in a certain frequency band, the sound pressure level target for that frequency band is increased based on the environmental impact calculation results. The increase amount is ΔL=k(N-L0), where k is the impact coefficient, N is the noise sound pressure level in that frequency band, and L0 is the initial sound pressure level target. If the ambient reverberation time is longer than the target reverberation time, the reverberation suppression parameters are adjusted through the adaptive filter of the digital audio system, and the optimized reverberation time target is corrected to T1=Tt+α(Te-Tt), where α is the correction coefficient determined based on the environmental impact, Te is the measured reverberation time in the environment, and Tt is the initial reverberation time target. The value of α makes the optimized reverberation time target closer to the auditory needs in the actual environment.

[0047] Furthermore, regarding the frequency band gain control target, if the environmental acoustic parameters indicate that the room has severe sound absorption in the high-frequency band, such as above 8kHz, resulting in excessively rapid sound pressure level decay, then the gain target in the high-frequency band will be adjusted upwards based on the degree of environmental impact, for example, from the initial +1dB to +3dB, to compensate for the loss caused by environmental absorption, ensure that the sound pressure level in each frequency band reaches the optimized target value, and ultimately achieve the control of the sound field parameters.

[0048] S40: Based on the goal of optimizing sound field control and combined with the playback scenario, the control parameters of digital audio are iteratively optimized to obtain an optimized control scheme and intelligently control the digital audio.

[0049] In this embodiment, based on the optimized sound field control target and the identified playback scenario, the control parameters of the digital audio system are dynamically adjusted through a multi-parameter collaborative iterative optimization algorithm to generate the final optimized control scheme.

[0050] Among them, the multi-parameter collaborative iterative optimization algorithm is an adaptive optimization method that comprehensively considers the sound field control target, hardware device performance constraints and user subjective listening feedback. Its core lies in constructing a mapping relationship model between control parameters and the optimized sound field control target, and making the actual output sound field as close as possible to the optimized target through iterative optimization.

[0051] Specifically, step S40 in the method includes: Based on the degree of environmental impact, obtain the initial optimization step size; Based on the particle swarm optimization algorithm, with the optimization of sound field control as the optimization objective, the control parameters of digital audio are iteratively optimized to obtain an optimized control scheme and intelligently control the digital audio.

[0052] In this embodiment, the initial optimization step size is first obtained based on the degree of environmental impact. The initial optimization step size needs to be dynamically adjusted according to the significance of the environment's influence on the sound field control target, in order to balance optimization efficiency and control accuracy.

[0053] Specifically, when the environmental impact value is high, such as greater than 0.6, it indicates that the current environment significantly interferes with the initial control target, and the parameter correction requirement is significant. In this case, the initial optimization step size is set to a larger value, such as 0.15-0.2, so that the control parameters can quickly approach the optimization target and shorten the iteration convergence time. When the environmental impact value is low, such as less than 0.3, the environmental interference is small, and the initial control target is already highly feasible. In this case, the initial optimization step size is set to a smaller value, such as 0.05-0.1, to avoid excessive parameter adjustment leading to overshoot or oscillation, and to ensure the stability of the control process. For moderate impact cases with an environmental impact value between 0.3 and 0.6, a moderate initial optimization step size, such as 0.1-0.15, is used, and can be fine-tuned according to the fluctuation of specific environmental parameters. For example, if the environmental noise spectrum fluctuates greatly, the step size can be appropriately increased by 0.02-0.03 to enhance the algorithm's adaptability to dynamic environments.

[0054] Then, based on the particle swarm optimization algorithm, with the optimization of sound field control as the optimization objective, the control parameters of the digital audio system are iteratively optimized to obtain an optimized control scheme and intelligently control the digital audio system.

[0055] Among them, the particle swarm optimization algorithm is a global optimization algorithm based on swarm intelligence. It finds the optimal solution by simulating the information sharing and cooperative behavior of birds in the process of foraging.

[0056] In this application, various control parameters of the digital audio system, such as the gain of each EQ band, volume reference value, reverberation depth parameter, channel balance coefficient, and bass enhancement coefficient, are regarded as particles in a particle swarm. The position vector of each particle corresponds to a specific combination of control parameters. During the initialization of the particle swarm, the initial position of the particles is randomly generated within the range of control parameter values ​​allowed by the digital audio hardware, and the initial velocity is set according to the dynamic range of the parameters. For example, for the EQ band gain range of -12dB to +12dB, the initial velocity can be set to ±2dB / iteration.

[0057] Furthermore, the algorithm's fitness function is designed to measure the degree of closeness between the actual output sound field and the optimized sound field control target under the current combination of control parameters. Specifically, it includes: calculating the root mean square error (RMSE) between the sound pressure level curve of the actual sound field and the optimized sound pressure target curve, the absolute error between the actual reverberation time and the optimized reverberation time target, and the weighted sum of the deviations between the actual gain values ​​and the optimized gain targets for each frequency band.

[0058] The smaller the fitness function value, the better the current combination of control parameters. In each iteration, for each combination of control parameters represented by a particle, the actual output sound field parameters are obtained through the simulation model of the digital audio system or the real-time feedback module, substituted into the fitness function to calculate its fitness value, and the individual extreme value and the global extreme value are updated accordingly.

[0059] The iteration terminates when the preset maximum number of iterations is reached, such as 50, or when the change in the global optimal fitness value is less than a preset threshold (e.g., 0.5%) over five consecutive iterations. When the termination condition is met, the combination of control parameters corresponding to the global optimal position is the optimized control scheme.

[0060] Finally, the optimized control scheme is converted into control commands that can be executed by the digital audio system. These commands are then sent to the DSP processing unit of the digital audio system through the control interface. The operating parameters of the internal equalizer, reverb, volume controller, and other modules are adjusted in real time to achieve intelligent control of the sound field of the digital audio system. This ensures that the actual playback effect can accurately match the user's auditory needs for specific playback content in the current scenario.

[0061] Among them, based on the particle swarm optimization algorithm, with the optimization of sound field control as the objective, the control parameters of the digital audio system are iteratively optimized to obtain an optimized control scheme, and the digital audio system is intelligently controlled, including: Obtain the initial control scheme, which includes volume intensity, filter parameters, and frequency band gain; Based on the initial optimization step size, the initial control scheme is iteratively optimized using the particle swarm optimization algorithm until the iteration optimization round threshold is reached or an optimized control scheme with a fitness greater than the fitness threshold is obtained. The optimized control scheme is then used to intelligently control the digital audio system.

[0062] In this embodiment, firstly, an initial control scheme is obtained based on the optimized sound field control target and a set of basic control parameters preset within the hardware parameter range of the digital audio device.

[0063] The initial volume intensity is set to 60%-70% of the loudness corresponding to the optimized sound pressure target to avoid the initial volume being too high or too low. The filter parameters include the cutoff frequency of the low-pass filter, the cutoff frequency of the high-pass filter, and the center frequency and bandwidth of the band-pass filter. The initial values ​​are adapted according to the type of content being played. For example, when playing classical music, the initial cutoff frequency of the low-pass filter is set to 18kHz to preserve high-frequency details, while when playing rock music, it can be appropriately reduced to 16kHz to reduce high-frequency noise. The initial value of the frequency band gain is directly mapped to the frequency band gain target in the optimized sound field control target. For example, if the 300-3000Hz frequency band needs to be boosted by 2dB in the optimization target, the EQ gain parameter of that frequency band in the initial control scheme is set to +2dB.

[0064] Secondly, based on the initial optimization step size obtained above, the initial control scheme is iteratively optimized using the particle swarm optimization algorithm. During the iteration process, each particle represents a set of control schemes including volume intensity, filter parameters, and frequency band gain. The particle's position update is guided by individual extrema and global extrema, while the velocity update is dynamically adjusted in combination with the initial optimization step size and parameters such as inertia weight, cognitive coefficient, and social coefficient.

[0065] For example, if a particle's current fitness value is higher than its historical best value, its position update will be more inclined to move towards the individual's extreme value, and the movement step size is determined by the initial optimization step size and the cognitive coefficient. At the same time, the position information of the global extreme value affects the movement direction of the entire particle swarm through the social coefficient, prompting the group to search for a better region.

[0066] Furthermore, the iterative optimization process continues until a preset threshold for the number of iterations is reached, such as 30 iterations, or until an optimized control scheme with a fitness value greater than the preset fitness threshold is obtained through multiple consecutive iterations. The fitness threshold is set according to the accuracy requirements of the actual application scenario; for example, a fitness threshold of 0.9 is used, where a smaller fitness function value is better. Here, 0.9 is a normalized relative threshold, indicating that the matching degree between the actual sound field and the target sound field reaches over 90%. When any of the above conditions are met, the iteration stops. At this point, the combination of control parameters corresponding to the globally optimal particle is the final optimized control scheme, which is then converted into control commands for the digital speaker, enabling intelligent control of the digital speaker's sound field.

[0067] Furthermore, the iterative optimization of the initial control scheme based on the particle swarm optimization algorithm also includes: During the iterative optimization process, the fitness similarity of multiple solutions is calculated. The fitness is obtained based on the reciprocal of the deviation between the control result and the control target. When the deviation between the control result and the control target is 0, the fitness is 1. Two solutions with a fitness similarity greater than a similarity threshold are obtained. The better solution with the higher fitness is retained, and the optimization adjustment weight is obtained based on the fitness of the inferior solution with the lower fitness. The optimization adjustment weight is proportional to the fitness of the inferior solution. The optimal solution is assigned optimization adjustment weights and iteratively optimized. The similarity threshold is obtained based on the typicality of the playback scenario and auditory requirements.

[0068] In this embodiment, firstly, during the iterative optimization process, for the control schemes corresponding to all particles in the particle swarm, the fitness similarity between any two solutions is calculated. The fitness similarity is obtained by calculating the cosine similarity of the fitness values ​​of the two solutions or the normalized value of the Euclidean distance. For example, when using cosine similarity, the fitness vectors of the two solutions are calculated by vector cosine value, which includes equal components such as the root mean square error of sound pressure level, the absolute error of reverberation time, and the weighted sum of frequency band gain deviation, resulting in a similarity value ranging from 0 to 1. The closer the value is to 1, the more similar the fitness characteristics of the two solutions are.

[0069] Specifically, the fitness is calculated as before, based on the reciprocal of the deviation between the control result and the control target. When the deviation is 0, that is, the actual sound field completely matches the target sound field, the fitness takes the maximum value of 1, which represents the optimal solution.

[0070] Secondly, a similarity threshold is set, which is dynamically adjusted based on the typicality of the playback scenario and the auditory requirements. For example, in typical scenarios like home theaters where sound field accuracy is extremely high, the auditory requirements are clear and standardized, and the similarity threshold is set to 0.85 to rigorously filter similar solutions and avoid redundant searches. However, in atypical scenarios such as outdoor gatherings with complex environmental interference, the ambiguity of the auditory requirements is higher, and the similarity threshold can be reduced to 0.7, allowing for a certain degree of solution diversity to cope with environmental uncertainties. When the fitness similarity of two solutions is detected to be greater than the set similarity threshold, they are determined to be similar solutions.

[0071] At this point, the better solution with the higher fitness (i.e., the smaller fitness function value) is retained, while the optimization adjustment weight is obtained based on the fitness value of the inferior solution with the smaller fitness. The formula for calculating the optimization adjustment weight is ω=(f worst / f best )×β, where f worst f is the fitness value of the inferior solution. bestLet ω be the fitness value of the optimal solution, and β be the scenario correction coefficient. Specifically, β = 0.3 for typical scenarios and β = 0.5 for atypical scenarios. ω ranges from 0 to 1 and is directly proportional to the fitness value of the inferior solution; that is, the worse the fitness of the inferior solution, the lower the fitness value. worst The larger the value of ω, the greater the need for adjustment to the optimal solution to avoid getting trapped in local optima.

[0072] Finally, the retained optimal solutions are assigned an optimization adjustment weight ω for a new round of iterative optimization. Specifically, the weight ω is introduced into the position update formula of the optimal solution to adjust its velocity update strategy: v new =ω×v old +c1×r1(p best -x)+c2×r2(g best -x), where v old The old speed is represented by c1 and c2, which are the cognitive and social coefficients, respectively. r1 and r2 are random numbers between 0 and 1. p best For the individual extreme value location, g best The location represents the global extremum. By introducing ω, the optimal solution, while retaining its advantages, dynamically adjusts its exploration capability based on the "quality" of the inferior solution. When the fitness of the inferior solution is poor, i.e., ω is large, the speed update of the optimal solution increases, enhancing the exploration of new regions; when the fitness of the inferior solution is close to that of the optimal solution, i.e., ω is small, the speed update of the optimal solution tends to be conservative, focusing on local fine-grained search.

[0073] Specifically, iterative optimization of the initial control scheme based on the particle swarm optimization algorithm also includes: Based on the playback scenario, evaluate the scene iteration effect and obtain the threshold for the iteration optimization round; Based on the auditory needs scheme, an auditory fitness assessment is conducted to obtain the fitness threshold; Calculate the frequency of the playback scene in the playback scene tag, and calculate the frequency of the auditory demand solution in the sample auditory demand solution set. Perform a weighted sum to obtain the typicality of the playback scene and the auditory demand solution.

[0074] In this embodiment, firstly, based on the playback scenario, the effect of scene iteration is evaluated to obtain the threshold for the iteration optimization round. Specifically, for different playback scenarios, a mapping relationship model between scene characteristics and the iteration optimization round threshold is pre-constructed using experimental data.

[0075] The scene characteristics include environmental complexity, sound field stability, and user sensitivity to control response speed. For example, real-time live streaming requires rapid convergence, while background playback allows for a slightly longer optimization time. After obtaining the label of the current playback scene through the scene recognition module, the historical optimization data corresponding to that scene is retrieved to analyze the optimization effect improvement curves under different iteration rounds. For example, for a stable environment with minimal interference in a home bedroom, experimental data shows that after 30 iterations, the improvement in fitness value is less than 5% of the initial improvement. Therefore, the iteration optimization round threshold for this scene is set to 30. For a dynamically changing in-vehicle scene, to ensure rapid response to changes in road noise during vehicle operation, experiments show that 20 iterations are sufficient to achieve a usable sound field control effect. Therefore, the iteration optimization round threshold is set to 20.

[0076] Secondly, based on the auditory demand scheme, an auditory adaptation assessment is conducted to obtain the adaptation threshold. The auditory demand scheme includes a specific description of the user's sound quality preferences, such as "clear vocals," "powerful bass," and "surround immersion," with different sound field parameter weights corresponding to these preferences. The auditory adaptation assessment breaks down the user's auditory demand scheme into quantifiable sound field target parameter indicators and, in conjunction with a psychoacoustic model, determines the allowable deviation range for different indicators, thereby setting the adaptation threshold.

[0077] Finally, the frequency of the playback scene in the playback scene tag is calculated, and the frequency of the auditory demand solution in the sample auditory demand solution set is calculated. These are then weighted and summed to obtain the typicality of the playback scene and auditory demand solution. The frequency of the playback scene tag is determined by counting the number of times each scene tag is activated within a preset time period. For example, if the "Home Theater" scene appeared 15 times and the "Outdoor Picnic" scene appeared 5 times in the past 30 days, then the frequency of "Home Theater" is 15 / (15+5) = 0.75.

[0078] Furthermore, the frequency of occurrence of auditory demand solutions is determined by matching the current auditory demand solution with each solution in the sample auditory demand solution set, using a text similarity algorithm based on semantic analysis or feature vector cosine similarity, and counting the proportion of sample solutions with a similarity higher than 80% to the total number of samples.

[0079] For example, if the sample set contains 100 auditory demand solutions for "heavy metal music preference", and the current user's demand solution has a similarity of 85%, then its frequency of occurrence is 100 / total number of samples. The weighted summation formula for typicality is: Typicality = α × Scene Occurrence Frequency + (1-α) × Auditory Demand Solution Occurrence Frequency, where α is the scene weight coefficient, with a value ranging from 0.3 to 0.7, dynamically adjusted according to the application scenario type.

[0080] Furthermore, for scenarios with fixed hardware devices, such as smart speakers, the scenario has a greater impact on sound field control, so α is taken as 0.7; for portable devices, such as headphones, user auditory needs play a dominant role, so α is taken as 0.3.

[0081] For example, the frequency of the "home theater" scenario is 0.6, the frequency of the "high-definition surround sound" auditory requirement is 0.7, and α is set to 0.5. Then, the typicality = 0.5 × 0.6 + 0.5 × 0.7 = 0.65. The higher the typicality value, the more common the current scenario and requirement combination is, and the richer the optimization experience. In this case, the similarity threshold can be appropriately increased to reduce redundant calculations. Conversely, a low typicality value, such as less than 0.3, indicates a rare scenario or personalized requirement. The similarity threshold should be lowered to retain more diverse solutions to explore potential optimization directions.

[0082] In summary, compared to existing technologies, this application deeply integrates AI scene recognition with particle swarm optimization algorithms, using the deviation between sound field parameters and user auditory needs as the core optimization objective, and constructs a dynamically adaptive iterative optimization mechanism. By pre-setting an initial control scheme and dynamically adjusting the optimization step size, similarity threshold, and iteration termination conditions in combination with scene characteristics, precise optimization of volume intensity, filter parameters, and frequency band gain is achieved.

[0083] In summary, the embodiments of this application have at least the following technical effects: This application provides an AI-based method for intelligently controlling the sound field parameters of digital speakers through scene recognition. First, it acquires environmental acoustic parameters and the content being played by the user, and then identifies the current playback scene based on the environmental acoustic parameters. Second, using a trained auditory demand analyzer, it analyzes the user's auditory preferences and needs for the content in that scene, combining the user's playback content with the identified playback scene, thereby generating a targeted auditory demand solution. Subsequently, it determines an initial sound field control target based on the auditory demand solution, and then assesses the environmental impact of the initial target using the environmental acoustic parameters, further optimizing the initial target to obtain an optimized sound field control target that better matches the actual environment. Finally, guided by the optimized sound field control target and considering the playback scene, iteratively optimizes the control parameters of the digital speaker using a particle swarm optimization algorithm. Simultaneously, based on the typicality of the playback scene and the characteristics of the auditory demand solution, iterative optimization round thresholds and fitness thresholds are set to ensure rapid convergence to the optimal solution in different scenes, ultimately obtaining an optimized control scheme and intelligently controlling the digital speaker.

[0084] Through the above technical solution, this application realizes intelligent control of the entire process from environmental perception and demand analysis to dynamic optimization and control, effectively overcoming the limitations of traditional control methods and improving sound quality and user experience.

[0085] Example 2, as Figure 2 As shown, based on the same inventive concept as the AI ​​scene recognition digital audio field parameter intelligent control method provided in Embodiment 1, this application also provides an AI scene recognition digital audio field parameter intelligent control system, including: The information acquisition module 11 is used to acquire environmental acoustic parameters and user playback content, and to identify the playback scene based on the environmental acoustic parameters to obtain the playback scene. The solution acquisition module 12 is used to perform auditory demand analysis based on the user's playback content and playback scenario, and to obtain auditory demand solutions. The environmental assessment module 13 is used to obtain the initial sound field control target based on the auditory demand scheme, and to conduct an environmental impact assessment based on environmental acoustic parameters to obtain the optimized sound field control target. The scheme optimization module 14 is used to iteratively optimize the control parameters of the digital audio system based on the sound field control target and the playback scenario, obtain the optimized control scheme, and intelligently control the digital audio system.

[0086] In one embodiment, the information acquisition module 11 is specifically used for: Acquire environmental acoustic parameters and user-played content; Input the ambient acoustic parameters into the playback scene recognizer to obtain the playback scene.

[0087] Furthermore, in one embodiment of the application, the acquisition of the playback scene recognizer includes: Obtain the sample environment acoustic parameter set; The acoustic parameter set of the sample environment is labeled to obtain playback scene labels, which are then integrated into a playback scene label set, which includes various environmental labels. Construct a playback scene recognizer, using sample environmental acoustic parameters as input and playback scene labels as supervision, and train the playback scene recognizer until convergence.

[0088] In one embodiment, the solution acquisition module 12 is specifically used for: Obtain an auditory demand analyzer, which is trained based on a set of sample user playback content, a set of sample playback scenarios, and a set of sample auditory demand solutions. Input the user's playback content and playback scenario into the auditory demand analyzer to obtain auditory demand solutions.

[0089] In one embodiment, the environmental assessment module 13 is specifically used for: Based on the auditory demand scheme, the initial sound field control target is obtained, which includes the sound pressure target and the reverberation time target. An environmental impact assessment is conducted based on environmental acoustic parameters to obtain the degree of environmental impact. The environmental impact assessment is based on the comprehensive impact of environmental parameters on reverberation time targets, sound pressure level control targets, and frequency band gain control targets. Based on the degree of environmental impact, the initial sound field control target is optimized to obtain the optimized sound field control target.

[0090] In one embodiment, the scheme optimization module 14 is specifically used for: Based on the degree of environmental impact, obtain the initial optimization step size; Based on the particle swarm optimization algorithm, with the optimization of sound field control as the optimization objective, the control parameters of digital audio are iteratively optimized to obtain an optimized control scheme and intelligently control the digital audio.

[0091] Furthermore, in one embodiment, based on the particle swarm optimization algorithm, with the optimization objective of sound field control as the optimization goal, the control parameters of the digital audio system are iteratively optimized to obtain an optimized control scheme, and the digital audio system is intelligently controlled, including: Obtain the initial control scheme, which includes volume intensity, filter parameters, and frequency band gain; Based on the initial optimization step size, the initial control scheme is iteratively optimized using the particle swarm optimization algorithm until the iteration optimization round threshold is reached or an optimized control scheme with a fitness greater than the fitness threshold is obtained. The optimized control scheme is then used to intelligently control the digital audio system.

[0092] Furthermore, the iterative optimization of the initial control scheme based on the particle swarm optimization algorithm also includes: During the iterative optimization process, the fitness similarity of multiple solutions is calculated. The fitness is obtained based on the reciprocal of the deviation between the control result and the control target. When the deviation between the control result and the control target is 0, the fitness is 1. Two solutions with a fitness similarity greater than a similarity threshold are obtained. The better solution with the higher fitness is retained, and the optimization adjustment weight is obtained based on the fitness of the inferior solution with the lower fitness. The optimization adjustment weight is proportional to the fitness of the inferior solution. The optimal solution is assigned optimization adjustment weights and iteratively optimized. The similarity threshold is obtained based on the typicality of the playback scenario and auditory requirements.

[0093] Specifically, iterative optimization of the initial control scheme based on the particle swarm optimization algorithm also includes: Based on the playback scenario, evaluate the scene iteration effect and obtain the threshold for the iteration optimization round; Based on the auditory needs scheme, an auditory fitness assessment is conducted to obtain the fitness threshold; Calculate the frequency of the playback scene in the playback scene tag, and calculate the frequency of the auditory demand solution in the sample auditory demand solution set. Perform a weighted sum to obtain the typicality of the playback scene and the auditory demand solution.

[0094] It should be noted that the order of the embodiments described above is merely for descriptive purposes and does not represent the superiority or inferiority of the embodiments. Furthermore, the above description focuses on specific embodiments of this specification. Additionally, the processes depicted in the accompanying drawings do not necessarily require a specific or sequential order to achieve the desired results. In some implementations, multitasking and parallel processing are possible or may be advantageous.

[0095] The above are merely preferred embodiments of this application and are not intended to limit this application. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the protection scope of this application.

[0096] This specification and accompanying drawings are merely illustrative examples of this application and are intended to cover any and all modifications, variations, combinations, or equivalents within the scope of this application. Clearly, those skilled in the art can make various alterations and modifications to this application without departing from its scope. Therefore, if such modifications and modifications fall within the scope of this application and its equivalents, this application intends to include such modifications and modifications.

Claims

1. An AI-based method for intelligent control of digital audio field parameters based on scene recognition, characterized in that: include: Acquire environmental acoustic parameters and user-played content, and identify the playback scene based on the environmental acoustic parameters to obtain the playback scene; Based on the user's playback content and the playback scenario, auditory needs analysis is conducted to obtain auditory needs solutions; Based on the auditory demand scheme, the initial sound field control target is obtained, and based on the environmental acoustic parameters, an environmental impact assessment is conducted to obtain the optimized sound field control target. Based on the aforementioned sound field control objective and combined with the playback scenario, the control parameters of the digital audio system are iteratively optimized to obtain an optimized control scheme and intelligently control the digital audio system.

2. The intelligent control method for digital audio field parameters based on AI scene recognition according to claim 1, characterized in that, Acquire ambient acoustic parameters and user-played content, and identify the playback scene based on the ambient acoustic parameters to obtain the playback scene, including: Acquire environmental acoustic parameters and user-played content; The environmental acoustic parameters are input into the playback scene recognizer to obtain the playback scene.

3. The intelligent control method for digital audio field parameters based on AI scene recognition according to claim 2, characterized in that, Acquiring the playback scene recognizer includes: Obtain the sample environment acoustic parameter set; The sample environmental acoustic parameter set is labeled to obtain playback scene tags, which are then integrated into a playback scene tag set, wherein the playback scene tags include multiple environmental tags; A playback scene recognizer is constructed by using the acoustic parameters of the sample environment as input and the playback scene labels as supervision, and training the playback scene recognizer until convergence.

4. The intelligent control method for digital audio field parameters based on AI scene recognition according to claim 1, characterized in that, Based on the user's playback content and the playback scenario, auditory needs analysis is performed to obtain auditory needs solutions, including: An auditory demand analyzer is obtained, wherein the auditory demand analyzer is trained and obtained based on a set of sample user playback content, a set of sample playback scenarios, and a set of sample auditory demand solutions. The user's playback content and the playback scenario are input into the auditory demand analyzer to obtain an auditory demand solution.

5. The intelligent control method for digital audio field parameters based on AI scene recognition according to claim 1, characterized in that, Based on the auditory demand scheme, initial sound field control targets are obtained, and based on environmental acoustic parameters, an environmental impact assessment is conducted to obtain optimized sound field control targets, including: Based on the auditory demand scheme, an initial sound field control target is obtained, wherein the initial sound field control target includes a sound pressure target and a reverberation time target; An environmental impact assessment is conducted based on environmental acoustic parameters to obtain the degree of environmental impact. The environmental impact assessment is based on the comprehensive impact of environmental parameters on reverberation time targets, sound pressure level control targets, and frequency band gain control targets. Based on the environmental impact level, the initial sound field control target is optimized to obtain the optimized sound field control target.

6. The intelligent control method for digital audio field parameters based on AI scene recognition according to claim 1, characterized in that, Based on the aforementioned sound field control objective and in conjunction with the playback scenario, the control parameters of the digital audio system are iteratively optimized to obtain an optimized control scheme. This allows for intelligent control of the digital audio system, including: Based on the environmental impact level, obtain the initial optimization step size; Based on the particle swarm optimization algorithm, with the optimized sound field control target as the optimization objective, the control parameters of the digital audio system are iteratively optimized to obtain an optimized control scheme and intelligently control the digital audio system.

7. The intelligent control method for digital audio field parameters based on AI scene recognition according to claim 6, characterized in that, Based on the particle swarm optimization algorithm, with the optimization of sound field control as the objective, the control parameters of the digital audio system are iteratively optimized to obtain an optimized control scheme, enabling intelligent control of the digital audio system, including: Obtain an initial control scheme, wherein the initial control scheme includes volume intensity, filter parameters, and frequency band gain; Based on the initial optimization step size, the initial control scheme is iteratively optimized using the particle swarm optimization algorithm until the iteration optimization round threshold is reached or an optimized control scheme with a fitness greater than the fitness threshold is obtained. The optimized control scheme is then used to intelligently control the digital audio system.

8. The intelligent control method for digital audio field parameters based on AI scene recognition according to claim 7, characterized in that, The initial control scheme is iteratively optimized based on the particle swarm optimization algorithm, and the optimization further includes: During the iterative optimization process, the fitness similarity of multiple solutions is calculated. The fitness is obtained based on the reciprocal of the deviation between the control result and the control target. When the deviation between the control result and the control target is 0, the fitness is 1. Two solutions with a fitness similarity greater than a similarity threshold are obtained. The better solution with the higher fitness is retained, and the optimization adjustment weight is obtained based on the fitness of the inferior solution with the lower fitness. The optimization adjustment weight is proportional to the fitness of the inferior solution. The optimal solution is assigned the optimization adjustment weight and iteratively optimized, wherein the similarity threshold is obtained based on the typicality of the playback scenario and the auditory demand scheme.

9. The intelligent control method for digital audio field parameters based on AI scene recognition according to claim 8, characterized in that, The initial control scheme is iteratively optimized based on the particle swarm optimization algorithm, and the optimization further includes: Based on the playback scenario, the scene iteration effect is evaluated, and the threshold of the iteration optimization round is obtained; Based on the aforementioned auditory requirement scheme, an auditory fitness assessment is performed to obtain the fitness threshold. Calculate the frequency of occurrence of the playback scene in the playback scene tag, and calculate the frequency of occurrence of the auditory demand scheme in the sample auditory demand scheme set. Perform a weighted summation to obtain the typicality of the playback scene and the auditory demand scheme.

10. An AI-based scene recognition-based intelligent control system for digital audio field parameters, characterized in that: The method for intelligent adjustment of digital audio field parameters for performing AI scene recognition as described in any one of claims 1-9 includes: The information acquisition module is used to acquire environmental acoustic parameters and user playback content, and to identify the playback scene based on the environmental acoustic parameters. The solution acquisition module is used to perform auditory demand analysis based on the user's playback content and the playback scenario, and to obtain auditory demand solutions. The environmental assessment module is used to obtain initial sound field control targets based on the auditory demand scheme, and to conduct environmental impact assessments based on environmental acoustic parameters to obtain optimized sound field control targets. The scheme optimization module is used to iteratively optimize the control parameters of the digital audio system based on the optimized sound field control target and the playback scenario, obtain an optimized control scheme, and intelligently control the digital audio system.