Autonomous audio control system based on artificial intelligence

By adopting artificial intelligence technology in the audio control system, using policy networks and value function networks for audio adjustment, and integrating user feedback optimization control strategies, the problem that traditional audio control systems cannot dynamically adapt to environmental changes and poor adaptability is solved, and personalized and accurate audio experience and system adaptability and robustness are achieved.

CN120161776AInactive Publication Date: 2025-06-17WENZHOU POLYTECHNIC

Patent Information

Application Number
CN202510641271.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-19
Publication Date
2025-06-17
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Traditional audio control systems cannot dynamically adapt to environmental changes, the control process is rigid, and the integration of user feedback is insufficient, so they cannot automatically learn and optimize control strategies, and have poor adaptability.

Method used

The audio autonomous control system based on artificial intelligence is adopted to realize the comprehensive perception of the user's current environment and state through the policy network and the value function network, integrate the user feedback signal into the loss function of the model, and optimize the control strategy.

Benefits of technology

It realizes personalization and accuracy of audio adjustment, improves audio processing capabilities, provides users with real-time, efficient and comfortable personalized audio experience, and enhances the adaptability and robustness of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120161776A_ABST
    Figure CN120161776A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of audio management, and particularly discloses an artificial intelligence-based audio autonomous control system, which comprises an audio monitoring unit, an environment acquisition unit, a processing unit and an output unit. According to the scheme, the comprehensive perception of the current environment and state of the user is realized, and the control process is refined through the strategy network and the value function network, so that the audio adjustment is more personalized and accurate, the audio processing capability is effectively improved, and real-time, efficient and comfortable personalized audio experience is provided for the user; user feedback signals are integrated into loss functions of a model, a system is further optimized through strategy adjustment, and strategy optimization and user feedback are better balanced by defining various loss functions and reasonably adjusting weights of the loss functions, so that an optimal control strategy is found, effectiveness of the control strategy is ensured, and adaptability and robustness of the system are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of audio management, and specifically refers to an audio autonomous control system based on artificial intelligence. Background Art

[0002] Audio applications involve various scenarios such as smart homes, in-vehicle audio systems, and personal audio devices. The importance that users attach to the audio experience makes it necessary for audio control systems to possess intelligent and adaptive control capabilities. Traditional audio control systems are usually based on simple rule settings and predefined scenarios, with a rigid control process that cannot dynamically adapt to environmental changes; general audio control strategies have insufficient integration of user feedback, cannot automatically learn and optimize control strategies, and have poor adaptability. Summary of the Invention

[0003] In view of the above situation, to overcome the defects of the prior art, the present invention provides an audio autonomous control system based on artificial intelligence. Aiming at the problem that traditional audio control systems are usually based on simple rule settings and predefined scenarios, with a rigid control process that cannot dynamically adapt to environmental changes, this solution realizes the comprehensive perception of the user's current environment and state, and refines the control process through a policy network and a value function network, making audio adjustment more personalized and precise, effectively improving the audio processing ability, and providing users with a real-time, efficient, and comfortable personalized audio experience; aiming at the problem that general audio control strategies have insufficient integration of user feedback, cannot automatically learn and optimize control strategies, and have poor adaptability, this solution integrates the user feedback signal into the loss function of the model, further optimizes the system through policy adjustment, defines multiple loss functions and reasonably adjusts their weights to better balance policy optimization and user feedback, thereby finding the optimal control strategy, ensuring the effectiveness of the control strategy, and improving the self-adaptability and robustness of the system.

[0004] An audio autonomous control system based on artificial intelligence provided by the present invention includes an audio monitoring unit, an environment acquisition unit, a processing unit, and an output unit;

[0005] The audio monitoring unit monitors the background service of the audio device, collects audio parameters from the background service, and sends them to the processing unit;

[0006] The environment acquisition unit uses sensors to capture environmental data, collects interactive data manually input by the user, and sends the environmental parameters and interactive data to the processing unit;

[0007] The processing unit is provided with an MCU, which has a built-in memory, a microprocessor, and peripheral interfaces. Audio parameters, environmental parameters, and interaction data together serve as input signals. The peripheral interfaces receive the input signals and store them in the memory. Control strategies are deployed in the memory. The microprocessor reads the input signals for preprocessing, executes the control strategies to generate control instructions, saves the control instructions in the memory, and then transmits them to the output unit through the peripheral interfaces;

[0008] The output unit adjusts the audio parameters according to the control instructions, outputs the audio parameters to the audio device, and displays them on the user interaction interface;

[0009] Control strategies are deployed in the memory. The design of the control strategies includes the following steps:

[0010] Step S1: Environmental feature extraction, converting environmental data into a vector form to represent environmental features;

[0011] Step S2: Audio feature extraction, converting audio parameters into a vector form to represent audio features;

[0012] Step S3: Autonomous control, constructing a feature mapping model to represent the relationship between environmental features and audio features, and formulating control strategies;

[0013] Step S4: Strategy adjustment, using the interaction data as a feedback signal to adjust the control strategies and output control instructions.

[0014] Further, in step S1, the environmental features include noise features, motion states, and geographical locations. The specific steps are as follows:

[0015] Step S11: Noise feature extraction, using the fast Fourier transform to convert noise data from the time domain to the frequency domain, analyzing the main frequency components of the noise, extracting the corresponding frequency intensities and performing weighted summation to define the noise feature value. The formula used is as follows: ;

[0016] In the formula, represents the noise feature value, represents the number of main frequency components, represents different noise intensities, represents the weights corresponding to different noise intensities;

[0017] Step S12: Motion state encoding, encoding the motion state of the user, defining the stationary state encoding as 0, the walking state encoding as 1, and the running state encoding as 2;

[0018] Step S13: Geographical location extraction, dividing the geographical location into longitude values and latitude values;

[0019] Step S14: Environmental feature combination. The noise feature value, motion state encoding, and geographical location are linearly combined to obtain the environmental feature. The formula used is as follows: ;

[0020] In the formula, represents the environmental feature, represents the noise feature value, represents the motion state encoding, represents the longitude value, represents the latitude value.

[0021] Furthermore, in step S2, the audio parameters include sound intensity, sampling rate, music rhythm, and equalization settings, which are combined into a vector form to represent the audio feature. The formula used is as follows: ;

[0022] In the formula, represents the audio feature, represents the sound intensity, represents the sampling rate, represents the music rhythm, represents the equalization setting.

[0023] Furthermore, in step S3, the autonomous control includes the following steps:

[0024] Step S31: Establish an MDP model. Use the Markov decision process to model the control strategy. The Markov decision process includes a state space, an action space, a state transition probability, a reward function, and a discount factor. Use the environmental feature and the audio feature as the state space and the action space of the MDP model respectively. A group of vectors in the environmental feature is used as a state in the state space, and a group of vectors in the audio feature is used as an action in the action space. The formula used is as follows: ;

[0025] In the formula, represents the Markov decision process, represents the state space, represents the action space, represents the state transition probability, represents the reward function, represents the discount factor;

[0026] Step S32: Define the policy network. Input a state in the state space and output the probability of executing the corresponding action under the given state;

[0027] Step S33: Define the value function network. Input a state in the state space and output the estimated value of the current state;

[0028] Step S34: Define the loss function, and calculate the policy loss, value function loss, and policy change loss as the overall loss function. The formulas used are as follows: ;

[0029] In the formula, represents the parameters of the policy network and the value function network, represents the overall loss function, represents the policy loss, represents the value function loss, represents the policy change loss, is a hyperparameter for weighing the loss, is the entropy of the policy;

[0030] Step S35: Define the optimizer and use the Adam optimizer to optimize the overall loss function;

[0031] Step S36: Model training. Iteratively train the feature mapping model, and use the optimizer to update the parameters of the policy network and the value function network until convergence. The formulas used are as follows: ;

[0032] In the formula, represents the time step, represents the learning rate, represents the deviation of the first moment estimate, represents the deviation of the second moment estimate, is a constant for stabilizing the numerical value.

[0033] Furthermore, Step S4, policy adjustment, includes the following steps:

[0034] Step S41: Obtain user feedback and use the interaction data as the feedback signal. The formula used is as follows: ;

[0035] In the formula, represents the feedback signal, represents the weighted summation operation on the interaction data, represents the number of times the user performs the interaction operation, represents the index of the interaction operation, represents the weight of the th interaction operation data, represents the adjustment value of the user for the

[0036] th interaction operation; ;

[0037] In the formula, represents the modified overall loss function, represents the weight of the feedback signal;

[0038] Step S43: Policy optimization. When the modified overall loss function reaches the optimal value, the action output by the policy network is used as the control instruction for output.

[0039] The beneficial effects achieved by the present invention using the above solution are as follows:

[0040] (1) Aiming at the problem that traditional audio control systems are usually based on simple rule settings and predefined scenarios, the control process is rigid and cannot dynamically adapt to environmental changes. This solution realizes the comprehensive perception of the user's current environment and state, and refines the control process through the policy network and the value function network, making the audio adjustment more personalized and accurate, effectively improving the audio processing ability, and providing users with a real-time, efficient, and comfortable personalized audio experience.

[0041] (2) Aiming at the problem that general audio control strategies lack integration of user feedback, cannot automatically learn and optimize control strategies, and have poor adaptability. This solution integrates the user feedback signal into the loss function of the model, further optimizes the system through policy adjustment, and defines multiple loss functions and reasonably adjusts their weights to better balance policy optimization and user feedback, thereby finding the optimal control strategy, ensuring the effectiveness of the control strategy, and improving the self-adaptability and robustness of the system. Description of the Drawings

[0042] Figure 1 is a schematic diagram of an audio autonomous control system based on artificial intelligence proposed by the present invention;

[0043] Figure 2 is a schematic diagram of the control strategy.

[0044] The drawings are used to provide a further understanding of the present invention, and constitute a part of the specification. Together with the embodiments of the present invention, they are used to explain the present invention, and do not constitute a limitation to the present invention. Detailed Embodiments

[0045] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments; based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0046] Example 1. Refer to Figure 1, an audio autonomous control system based on artificial intelligence provided by the present invention includes an audio monitoring unit, an environment acquisition unit, a processing unit, and an output unit;

[0047] The audio monitoring unit monitors the background service of the audio device, collects audio parameters from the background service, and sends them to the processing unit;

[0048] The environment acquisition unit uses sensors to capture environmental data, collects interactive data manually input by the user, and sends the environmental data and interactive data to the processing unit;

[0049] The processing unit is provided with an MCU. The MCU is built-in with a memory, a microprocessor, and a peripheral interface. The audio parameters, environmental data, and interactive data are used as input signals together. The peripheral interface receives the input signals and stores the input signals in the memory. A control strategy is deployed in the memory. The microprocessor reads the input signals for preprocessing, executes the control strategy to generate control instructions, saves the control instructions in the memory, and then transmits the control instructions to the output unit through the peripheral interface;

[0050] The output unit adjusts the audio parameters according to the control instructions, outputs the audio parameters to the audio device, and displays them on the user interaction interface.

[0051] Embodiment 2, refer to Figure 1 and 2 , based on the above embodiment, a control strategy is deployed in the memory. The design of the control strategy includes the following steps:

[0052] Step S1: Environment feature extraction, converting environmental data into a vector form to represent environmental features;

[0053] Step S2: Audio feature extraction, converting audio parameters into a vector form to represent audio features;

[0054] Step S3: Autonomous control, constructing a feature mapping model to represent the relationship between environmental features and audio features, and formulating a control strategy;

[0055] Step S4: Strategy adjustment, using the interactive data as a feedback signal to adjust the control strategy and output control instructions.

[0056] Embodiment 3, refer to Figure 1 and Figure 2 , based on the above embodiment, in step S1, the environmental features include noise features, motion states, and geographical locations. The specific steps are as follows:

[0057] Step S11: Noise feature extraction, including the following steps:

[0058] Step S111: Time-frequency conversion. Use the fast Fourier transform to convert the noise data from the time domain to the frequency domain to obtain the frequency-domain signal of the noise. The formula used is as follows: ;

[0059] In the formula, represents the frequency of the noise, represents the frequency-domain signal of the noise, represents time, represents the time-domain signal of the noise, represents the natural constant, is the imaginary unit;

[0060] Step S112: Frequency extraction. Analyze the main frequency components of the noise from the frequency-domain signal of the noise and extract the corresponding frequency intensities;

[0061] Step S113: Eigenvalue definition. Perform weighted summation on the frequency intensities and define it as the noise eigenvalue. The formula used is as follows: ;

[0062] In the formula, represents the noise eigenvalue, represents the number of main frequency components, represents different noise intensities, represents the weights corresponding to different noise intensities;

[0063] Step S12: Motion state encoding. Encode the motion state of the user. Define the stationary state encoding as 0, the walking state encoding as 1, and the running state encoding as 2;

[0064] Step S13: Geographic location extraction. Divide the geographic location into longitude value and latitude value;

[0065] Step S14: Environmental feature combination. Linearly combine the noise eigenvalue, motion state encoding, and geographic location to obtain the environmental feature. The formula used is as follows: ;

[0066] In the formula, represents the environmental feature, represents the noise eigenvalue, represents the motion state encoding, represents the longitude value, represents the latitude value.

[0067] Example 4. Refer to Figure 1 and Figure 2, this embodiment is based on the above embodiment. In step S2, the audio features include sound intensity, sampling rate, music rhythm, and equalization settings. The specific steps are as follows:

[0068] Step S21: Equalization setting. Set 3 frequency points and corresponding gain attenuation values, which respectively represent high-frequency, mid-frequency, and low-frequency. Convert them into vector form, and the formula used is as follows: ;

[0069] In the formula, represents the equalization setting, respectively represent high-frequency, mid-frequency, and low-frequency, respectively represent the gain attenuation values corresponding to high-frequency, mid-frequency, and low-frequency;

[0070] Step S22: Audio feature combination. Combine the sound intensity, sampling rate, music rhythm, and equalization setting of the audio parameters into vector form to represent the audio features. The formula used is as follows: ;

[0071] In the formula, represents the audio features, represents the sound intensity, represents the sampling rate, represents the music rhythm, represents the equalization setting.

[0072] Embodiment Five. Refer to Figure 1 and Figure 2 , this embodiment is based on the above embodiment. In step S3, the autonomous control includes the following steps:

[0073] Step S31: Establish an MDP model. Use the Markov decision process to model the control strategy. The Markov decision process includes a state space, an action space, a state transition probability, a reward function, and a discount factor. Use the environmental features and audio features as the state space and action space of the MDP model respectively. A group of vectors in the environmental features is used as a state in the state space, and a group of vectors in the audio features is used as an action in the action space. The formula used is as follows: ;

[0074] In the formula, represents the Markov decision process, represents the state space, represents the action space, represents the state transition probability, represents the reward function, represents the discount factor;

[0075] Step S32: Define a policy network, which consists of an input layer, a hidden layer, and an output layer. The input layer inputs a state in the state space. Four input nodes are designed to represent four variables of environmental features. The hidden layer is connected using the tanh activation function, and sixty-four hidden nodes are designed to calculate the mapping relationship. The output layer designs four output nodes to represent four parameters of audio features, and uses the softmax activation function to output the probability of performing the corresponding action in a given state;

[0076] Step S33: Define a value function network, which consists of an input layer, a hidden layer, and an output layer. The input layer inputs a state in the state space. Four input nodes are designed to represent four variables of environmental features. The hidden layer is connected using the tanh activation function, and sixty-four hidden nodes are designed to calculate the mapping relationship. The output layer designs one output node to output the estimated value of the current state;

[0077] Step S34: Define a loss function, and calculate the policy loss, value function loss, and policy change loss as the overall loss function;

[0078] Step S35: Define an optimizer, and use the Adam optimizer to optimize the overall loss function;

[0079] Step S36: Model training. Iteratively train the feature mapping model, and use the optimizer to update the parameters of the policy network and the value function network until convergence. The formula used is as follows: ;

[0080] In the formula, represents the parameters of the policy network and the value function network, represents the time step, represents the learning rate, represents the deviation of the first-order moment estimate, represents the deviation of the second-order moment estimate, is a constant for stable values.

[0081] By performing the above operations, for the problem that traditional audio control systems are usually based on simple rule settings and predefined scenarios, and the control process is rigid and unable to dynamically adapt to environmental changes, this solution realizes the comprehensive perception of the user's current environment and state, and refines the control process through the policy network and the value function network, making the audio adjustment more personalized and accurate, effectively improving the audio processing ability, and providing users with a real-time, efficient, and comfortable personalized audio experience.

[0082] Example 6. Refer to Figure 1 and Figure 2 , this example is based on the above example. In step S34, define a loss function, including the following steps:

[0083] Step S341: Policy loss. Use the CLIP clipping term as the policy loss, and the formula is as follows: ; ;

[0084] In the formula, represents the policy loss, represents the time step at the expectation of, represents the probability ratio of the new and old policies, represents the estimation of the advantage function, is the hyperparameter for clipping, is the current policy, is the old policy, represents the action, represents the state;

[0085] Step S342: Value function loss. Use the mean squared error to represent the value function loss, and the formula is as follows: ;

[0086] In the formula, represents the value function loss, is the estimation of the current value function, is the target value calculated from the cumulative reward;

[0087] Step S343: Policy change loss. Use the KL divergence penalty term to control the relative change between the new and old policies, and the formula is as follows: ;

[0088] In the formula, represents the policy change loss, is the hyperparameter for adjusting the penalty weight, represents the KL divergence, represents the probability distribution of the action under the state;

[0089] Step S344: Overall loss. The overall loss function is the sum of the policy loss, value function loss, and policy change loss, and the formula is as follows: ;

[0090] In the formula, represents the overall loss function, is the hyperparameter for weighing the loss, is the entropy of the policy.

[0091] Example Seven. Refer to Figure 1 and Figure 2, based on the above embodiment, in step S35, the definition of the optimizer includes the following steps:

[0092] Step S351: Calculate the first moment estimate of the gradient, and the formula used is as follows: ;

[0093] In the formula, is the decay coefficient, is the first moment estimate of the gradient at the current time step, is the first moment estimate of the gradient at the previous time step, is the gradient at the current time step;

[0094] Step S352: Correct the bias of the first moment estimate, and the formula used is as follows: ;

[0095] In the formula, represents the bias of the first moment estimate;

[0096] Step S353: Second moment estimate, calculate the second moment estimate of the gradient, and the formula used is as follows: ;

[0097] In the formula, is the decay coefficient, is the second moment estimate of the gradient at the current time step, is the second moment estimate of the gradient at the previous time step;

[0098] Step S354: Correct the bias of the second moment estimate, and the formula used is as follows: ;

[0099] In the formula, represents the bias of the second moment estimate.

[0100] Embodiment Eight, refer to Figure 1 and Figure 2 , based on the above embodiment, in step S4, the strategy adjustment includes the following steps:

[0101] Step S41: Obtain user feedback, and use the interaction data as the feedback signal. The formula used is as follows: ;

[0102] In the formula, represents the feedback signal, represents the weighted summation operation on the interaction data, represents the number of times the user performs interaction operations, represents the index of the interaction operation. represents the weight of the th interaction operation data, represents the adjustment value of the user for the th interaction operation;

[0103] Step S42: Modify the loss function, and integrate the feedback signal into the overall loss function of the feature mapping model. The formula used is as follows: ;

[0104] In the formula, represents the modified overall loss function, represents the weight of the feedback signal;

[0105] Step S43: Policy optimization. When the modified overall loss function reaches the optimal value, the action output by the policy network is used as the control instruction for output.

[0106] By performing the above operations, aiming at the problem that the general audio control policy has insufficient integration of user feedback, cannot automatically learn and optimize the control policy, and has poor adaptability, this solution integrates the user feedback signal into the loss function of the model, further optimizes the system through policy adjustment, and better balances policy optimization and user feedback by defining multiple loss functions and reasonably adjusting their weights, so as to find the optimal control policy, ensure the effectiveness of the control policy, and improve the self - adaptability and robustness of the system.

[0107] It should be noted that in this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non - exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article or device.

[0108] Although the embodiments of the present invention have been shown and described, for those of ordinary skill in the art, it can be understood that various changes, modifications, substitutions and variations can be made to these embodiments without departing from the principles and spirit of the present invention. The scope of the present invention is defined by the appended claims and their equivalents.

[0109] The above describes the present invention and its embodiments. Such description is not restrictive. What is shown in the drawings is only one of the embodiments of the present invention, and the actual structure is not limited thereto. Generally speaking, if those of ordinary skill in the art are inspired by it and, without departing from the gist of the present invention, design similar structural modes and embodiments to this technical solution without creative efforts, they shall fall within the protection scope of the present invention.

Claims

1. An audio autonomous control system based on artificial intelligence, characterized by: It includes an audio monitoring unit, an environment collection unit, a processing unit, and an output unit; The audio monitoring unit monitors the background service of the audio device, collects audio parameters from the background service and sends them to the processing unit; The environment acquisition unit uses sensors to capture environment data, collects interaction data manually input by users, and sends the environment parameters and interaction data to the processing unit; The processing unit is provided with an MCU, the MCU has a built-in memory, a microprocessor and a peripheral interface, the audio parameters, the environmental parameters and the interaction data are collectively used as input signals, the peripheral interface receives the input signals and stores the input signals in the memory, the control strategy is deployed in the memory, the microprocessor reads the input signals for preprocessing, executes the control strategy to generate control instructions, stores the control instructions in the memory, and then transmits them to the output unit through the peripheral interface; The output unit adjusts the audio parameters according to the control instruction, outputs the audio parameters to the audio device, and displays them on the user interaction interface; The control strategy is deployed in the memory, and the design of the control strategy includes the following steps: Step S1: extracting environmental features, converting environmental data into vector form to represent environmental features; Step S2: audio feature extraction, converting audio parameters into vector form to represent audio features; Step S3: autonomous control, constructing a feature mapping model to represent the relationship between environmental features and audio features, and formulating a control strategy; Step S4: Strategy adjustment, using the interaction data as a feedback signal to adjust the control strategy and output control instructions.

2. The audio autonomous control system based on artificial intelligence according to claim 1, characterized in that: In step S1, the environmental characteristics include noise characteristics, motion status, and geographic location. The specific steps are as follows: Step S11: Noise feature extraction, using fast Fourier transform to convert noise data from time domain to frequency domain, analyzing the main frequency components of the noise, extracting the corresponding frequency intensity and performing weighted summation to define the noise feature value; Step S12: motion state coding, coding the motion state of the user, defining the static state as 0, the walking state as 1, and the running state as 2; Step S13: extracting the geographical location, dividing the geographical location into longitude and latitude values; Step S14: Environmental feature combination, linearly combining the noise feature value, motion state code, and geographic location to obtain environmental features.

3. The audio autonomous control system based on artificial intelligence according to claim 2, characterized in that: In step S2, the audio parameters include sound intensity, sampling rate, music rhythm and equalization setting, which are combined into a vector form to represent the audio features.

4. The audio autonomous control system based on artificial intelligence according to claim 3 is characterized in that: In step S3, the autonomous control includes the following steps: Step S31: Establish an MDP model, and use a Markov decision process to model the control strategy. The Markov decision process includes a state space, an action space, a state transition probability, a reward function, and a discount factor. Environmental features and audio features are used as the state space and action space of the MDP model respectively. A set of vectors in the environmental features is used as a state in the state space, and a set of vectors in the audio features is used as an action in the action space. Step S32: define a policy network, input a state in the state space, and output the probability of executing a corresponding action in a given state; Step S33: define a value function network, input a state in the state space, and output an estimated value of the current state; Step S34: define a loss function, and calculate the strategy loss, value function loss, and strategy change loss as the overall loss function; Step S35: define an optimizer and use the Adam optimizer to optimize the overall loss function; Step S36: Model training, iteratively training the feature mapping model, using the optimizer to update the parameters of the policy network and the value function network until convergence.

5. The audio autonomous control system based on artificial intelligence according to claim 4 is characterized in that: Step S4, strategy adjustment, includes the following steps: Step S41: Obtain user feedback and use the interaction data as a feedback signal. The formula used is as follows: ; In the formula, Represents the feedback signal, represents the weighted sum operation of the interaction data, Indicates the number of user interaction operations. Indicates the index of the interaction operation, Indicates The weight of the interactive operation data, Indicates the user's The adjustment value of the interaction operation; Step S42: modify the loss function and integrate the feedback signal into the overall loss function of the feature mapping model. The formula used is as follows: ; In the formula, represents the modified overall loss function, represents the weight of the feedback signal, represents the strategy loss, represents the value function loss, represents the strategy change loss, is a hyperparameter that weighs the loss, is the entropy of the strategy; Step S43: Strategy optimization. When the modified overall loss function reaches the optimal value, the action output by the strategy network is output as a control instruction.

Citation Information

Patent Citations

  • Music necklace apparatus for automatically adjusting rhythm during running

    CN105595550A

  • Headphone control method, headphone and computer readable storage medium

    CN109765784A

  • Volume adjustment method and device for audio device and storage medium

    CN111930336A

  • Volume regulation and control method and device based on massage chair and mobile terminal and storage medium

    CN118519603A

  • Federal reinforcement learning system, method and equipment for multi-agent trusted interactive decision control

    CN118982061A

Cited By

  • Control method of intelligent sound system

    CN121262501A