A broadcast volume self-adaptive adjustment method, device, equipment, medium and product
By preprocessing and extracting features from the zoned audio signals, and combining this with reinforcement learning algorithms to adaptively adjust broadcast parameters, the problem of inaccurate separation of noise and broadcast signals was solved, achieving optimal broadcast propagation effects and listener experience in complex acoustic environments.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-10-18
- Publication Date
- 2026-04-07
AI Technical Summary
In existing technologies, noise and broadcast signals are not accurately separated, and the adaptive adjustment of broadcast parameters cannot meet user needs. In particular, in complex acoustic environments, broadcast sound is easily covered by noise, and listeners cannot accurately receive broadcast information.
By acquiring and preprocessing the ambient noise and broadcast signal, extracting comprehensive features, measuring the average sound pressure level, and using reinforcement learning algorithms to adaptively adjust broadcast parameters, including volume, gain ratio, and frequency response, to achieve the optimal parameter strategy.
Effectively separating noise and broadcast signals in complex acoustic environments ensures optimal transmission of broadcast signals, satisfies the listener's auditory experience, and guarantees the acquisition of critical information.
Smart Images

Figure CN119420442B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of broadcasting equipment technology, and in particular to a method, apparatus, device, medium and product for adaptive adjustment of broadcast volume. Background Technology
[0002] Today, broadcasting systems are widely used in various fields. In practical applications, whether in open spaces such as outdoor plazas, stations, and parks, or public areas such as schools, airports, and subways, broadcasting systems face complex acoustic environments. These areas are often accompanied by huge crowds, and the resulting noise cannot be ignored. When the noise is too loud, the broadcast sound is often easily masked, making it difficult for listeners to accurately receive and understand the broadcast information.
[0003] Currently, when environmental noise is too loud, adjusting the broadcast volume to cover the noise is often done manually or automatically by existing technologies. Manual adjustment requires a lot of manpower, and existing automatic adjustment often cannot separate noise from the broadcast. The broadcast parameters involved in the adjustment process are few, and it cannot adapt to changes in noise level in real time. Moreover, the adjustment effect is difficult to meet the listening experience of the audience. Summary of the Invention
[0004] This invention provides a method, apparatus, device, medium, and product for adaptive adjustment of broadcast volume, in order to solve the problems of inaccurate noise and broadcast separation technology in the prior art, and the inability of the adaptive adjustment effect of broadcast parameters to meet user needs.
[0005] In a first aspect, embodiments of the present invention provide a method for adaptive adjustment of broadcast volume, comprising:
[0006] Acquire the partition sound signal and preprocess the acquired partition sound signal to obtain partition sound data;
[0007] Feature extraction is performed on the acquired sound data to obtain comprehensive features. Based on the acquired comprehensive features, environmental noise and broadcast signals are separated. The average sound pressure level of the separated environmental noise and broadcast signals is measured respectively.
[0008] Based on the acquired environmental noise and the average sound pressure level of the broadcast signal, the type of broadcast content, the collected environmental data, and the reinforcement learning algorithm, the optimal parameter strategy corresponding to the broadcast is adaptively adjusted.
[0009] Based on the optimal parameter strategy, the current parameters of the broadcast signal are adjusted to the parameters corresponding to the optimal parameter strategy.
[0010] As a preferred embodiment of the first aspect, the integrated features include spectral features, temporal features, and sound signal energy features.
[0011] As a preferred embodiment of the first aspect, the separation of environmental noise and broadcast signals based on the acquired comprehensive features specifically includes the following steps:
[0012] When the broadcast system is not working, the partition sound data is the first ambient noise of the corresponding partition;
[0013] When the broadcasting system is working, it uses real-time acquired zone sound data to input the extracted comprehensive features into a preset sound classification model to obtain classification information for the zone sound data.
[0014] For the parts classified as noise, residual broadcast signal components are removed through filtering.
[0015] For the portion classified as broadcast signal, residual noise components are removed through filtering.
[0016] The partitioned sound data and the corresponding classification label results are input into the separation algorithm, which outputs the separated second environmental noise and the first broadcast signal.
[0017] As a preferred embodiment of the first aspect, the measurement of the average sound pressure level specifically includes: dividing the second ambient noise and the first broadcast signal into time windows, and calculating the average sound pressure level of all time serial ports.
[0018] As a preferred embodiment of the first aspect, the adaptive adjustment strategy for the optimal parameters corresponding to the broadcast based on the acquired ambient noise and the average sound pressure level of the broadcast signal, the type of broadcast content, the collected environmental data, and the reinforcement learning algorithm includes:
[0019] The environmental state of the reinforcement learning algorithm is defined based on the current ambient noise and the average sound pressure level of the broadcast signal, the type of broadcast content, and the collected environmental data; wherein, the collected environmental data includes the flow of people and time within the partition;
[0020] The action space of the reinforcement learning algorithm is defined according to the broadcast parameters; wherein, the action space includes: adjusting the volume, gain ratio and corresponding frequency response of the broadcast signal;
[0021] Based on the environmental state, the action space, and the reward function in the reinforcement learning algorithm, determine the reward value corresponding to the specific action currently being performed on the broadcast signal parameters;
[0022] Based on the reward value, the current parameter strategy of the broadcast signal is optimized to obtain the optimal parameter strategy.
[0023] As an optional embodiment of the first aspect, adjusting the current parameters of the broadcast signal to the parameters corresponding to the optimal parameter strategy based on the optimal parameter strategy includes:
[0024] The optimal parameter strategy is sent to the broadcast to be adjusted, and the volume, gain ratio and frequency response corresponding to the broadcast to be adjusted in the optimal parameter strategy are determined.
[0025] Based on the dynamic adjustment algorithm, the current parameters of the broadcast to be adjusted are adjusted to the parameters corresponding to the optimal parameter strategy.
[0026] Secondly, embodiments of the present invention provide a broadcast volume adaptive adjustment device, the broadcast volume adaptive adjustment device including a real-time acquisition module, a separation measurement module, a strategy determination module and a broadcast adjustment module;
[0027] The real-time acquisition module is used to acquire partitioned sound signals and preprocess the acquired partitioned sound signals to obtain partitioned sound data.
[0028] The separation measurement module extracts comprehensive features from the acquired sound data, separates environmental noise and broadcast signals based on the acquired comprehensive features, and measures the corresponding average sound pressure level of the separated environmental noise and broadcast signals respectively.
[0029] The strategy determination module is used to adaptively adjust the optimal parameter strategy corresponding to the broadcast based on the acquired ambient noise and the average sound pressure level of the broadcast signal, the type of broadcast content, the collected environmental data and the reinforcement learning algorithm.
[0030] The broadcast adjustment module is used to adjust the current parameters of the broadcast signal to the parameters corresponding to the optimal parameter strategy according to the optimal parameter strategy.
[0031] Thirdly, embodiments of the present invention also provide an electronic device, the electronic device comprising:
[0032] One or more processors;
[0033] Storage medium used to store one or more programs;
[0034] When the one or more programs are executed by the one or more processors, the one or more processors implement the broadcast volume adaptive adjustment method according to any embodiment of the present invention.
[0035] Fourthly, embodiments of the present invention also provide a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the broadcast volume adaptive adjustment method described in any embodiment of the present invention.
[0036] Fifthly, embodiments of the present invention also provide a computer program product, including a computer program that, when executed by a processor, implements the broadcast volume adaptive adjustment method as described in any embodiment of the present invention.
[0037] This invention provides a method, apparatus, electronic device, storage medium, and product for adaptive adjustment of broadcast volume. It acquires sound signals from different zones, preprocesses the acquired zone sound signals to obtain zone sound data, and extracts features from the acquired sound data to obtain comprehensive features. Based on these comprehensive features, noise and broadcast can be separated more effectively. The average sound pressure level (SPL) of the separated environmental noise and broadcast signal is measured separately, and the SPL data is used as the basis for adjusting broadcast parameters and environmental control to ensure the quality of sound propagation and help evaluate the propagation effect of the broadcast signal. Combining the acquired data on the average SPL of the environmental noise within the zone and the broadcast signal, as well as the type of broadcast content, a reinforcement learning algorithm is used to determine the corresponding optimal parameter strategy. Based on dynamic adjustment technology, the broadcast parameters are adaptively and dynamically adjusted to ensure that the auditory experience is satisfied in complex acoustic environments, to ensure the best propagation effect of the broadcast signal under different conditions, and to guarantee that the audience can obtain key information. Attached Figure Description
[0038] To more clearly illustrate the technical solution of this application, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.
[0039] Figure 1 A flowchart of a method for adjusting broadcast volume provided in this application;
[0040] Figure 2 This is a schematic diagram of a broadcast volume adjustment method provided by the present invention;
[0041] Figure 3 This is a schematic diagram of the structure of an electronic device provided by the present invention. Detailed Implementation
[0042] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0043] Example 1
[0044] Please refer to Figure 1 This is a flowchart illustrating an embodiment of a broadcast volume adaptive adjustment method provided by the present invention, including steps S01-S04, each step of which is as follows:
[0045] Step S01: Acquire the partition sound signal and preprocess the acquired partition sound signal to obtain partition sound data;
[0046] In this embodiment, high-quality audio sensors, including but not limited to probes and microphone arrays, are placed at appropriate locations within each zone to acquire sound signals within each zone. An accurate representation of the sound signal is provided by setting an appropriate sampling frequency; in a preferred embodiment, a sampling frequency of 44.1 kHz or 48 kHz is selected, and the sampling bit depth is set to 16 bits or 24 bits. The sound signals acquired in real time by the audio sensors are preprocessed to provide a good foundation for subsequent separation of environmental noise and broadcast signals.
[0047] The preprocessing includes DC component removal, filtering, and normalization. DC component removal is achieved by subtracting the average signal value from the sampled values, bringing the average value close to zero and preventing interference from DC components in subsequent frequency analysis and feature extraction. For example, in a station environment, there may be DC offsets caused by equipment or environmental factors; removing DC components improves the accuracy of subsequent processing. Filters remove noise outside a specific frequency range, such as electromagnetic interference from equipment in a station environment, highlighting the frequency range of broadcast signals and crowd noise.
[0048] Step S02: Extract features from the acquired sound data to obtain comprehensive features, separate environmental noise and broadcast signals based on the acquired comprehensive features, and measure the corresponding average sound pressure level of the separated environmental noise and broadcast signals respectively.
[0049] In a preferred embodiment, by extracting comprehensive features, detailed features of the sound signal in terms of frequency, time domain, and energy are obtained. These features provide detailed information for subsequent separation and classification, enabling the classifier to accurately determine whether the sound signal belongs to a broadcast signal or noise. The extraction of comprehensive features can make subsequent classification and separation better adapt to different environments and achieve more effective separation.
[0050] First, a Fast Fourier Transform (FFT) is performed on the preprocessed audio data. The FFT yields the corresponding frequency domain coefficients representing the frequency and phase of different frequency components. The FFT analysis then yields the audio data's spectrum, i.e., the energy distribution of different frequency components. The formula for calculating the FFT is: Where X[k] is the complex value at the corresponding frequency point in the frequency domain after the Fast Fourier Transform, which contains the amplitude and phase information of that frequency component; x[n] represents the amplitude value of the original sound data at discrete time point n, k represents the frequency index in the frequency domain, and j is the imaginary unit. It is a complex exponential function, which plays a key role in the Fast Fourier Transform (FFT) in converting time-domain signals to the frequency domain. It determines the phase rotation angle of different frequency components.
[0051] The spectrum of the sound signal is obtained by fast Fourier transform, which is the energy distribution of different frequency components. The dominant frequency is determined by finding the frequency component with the highest energy in the spectrum. In a preferred embodiment, the spectrum obtained by fast Fourier transform is |X[k]|, where |X[k]| represents the amplitude of the frequency domain coefficients. By traversing all frequency points, the frequency value corresponding to the frequency point with the largest amplitude is found.
[0052] The formula for calculating the clock speed is as follows: Where f0 represents the dominant frequency, i.e., the frequency component with the highest energy in the sound signal, and f represents the index value of the frequency point with the largest amplitude in the spectrum after k0 undergoes a Fast Fourier Transform. s The sampling frequency represents the number of times the sound signal is sampled per unit time. In a specific embodiment, the sampling frequency is set to 44100Hz, which means that the sound signal is sampled 44100 times per second. N represents the length of the sound signal.
[0053] Determine the frequency range where the broadcast signal may appear, such as [f1, f2]. Calculate the total energy within this frequency range using the following formula: Among them, E b This represents the total energy of the broadcast signal within a specific frequency range [f1, f2], where k represents the frequency index. Here, k takes values within the range [f1, f2], meaning only frequency points within the specific range are considered. The total energy of the entire spectrum is calculated to determine the proportion of the broadcast signal's energy in that frequency range. The formula for calculating the total spectrum energy is: Here, the value of k ranges from 0 to N-1, and the energy proportion of the broadcast signal is...
[0054] The standard deviation of energy is used to measure energy fluctuations. The formula for calculating the standard deviation is as follows: A smaller energy standard deviation indicates smaller energy fluctuations, which may be due to broadcast signals. Conversely, a larger energy standard deviation indicates larger energy fluctuations, which may be due to noise.
[0055] For the acquired signal x[n], the zero-crossing rate of the sound signal is calculated. The rhythm of the broadcast may cause the zero-crossing rate to change regularly within a certain time interval, while the zero-crossing rate of environmental noise is usually more random. The formula for calculating the zero-crossing rate is as follows: Where sgn(x) is the sign function, sgn(x) = 1 when x ≥ 0, and sgn(x) = -1 when x < 0; where ZCR represents the zero-crossing rate of the sound signal, N represents the length of the sound signal, and x[n] represents the amplitude value of the original sound signal at discrete time point n.
[0056] Specifically, the separation of environmental noise and broadcast signals based on the acquired comprehensive features includes the following steps:
[0057] When the broadcast system is not working, the partition sound data is the first ambient noise of the corresponding partition;
[0058] When the broadcasting system is working, it uses real-time acquired zone sound data to input the extracted comprehensive features into a preset sound classification model to obtain classification information for the zone sound data.
[0059] For the parts classified as noise, residual broadcast signal components are removed through filtering.
[0060] For the portion classified as broadcast signal, residual noise components are removed through filtering.
[0061] The partitioned sound data and the corresponding classification label results are input into the separation algorithm, which outputs the separated second environmental noise and the first broadcast signal.
[0062] Optionally, a random forest algorithm is used to construct a sound classification model. The sound data with extracted comprehensive features in real time is organized into a format processed by the random forest algorithm. In a specific embodiment, there is a set of sound segments, each represented by a feature vector. These vectors include features such as dominant frequency, frequency range energy percentage, and energy standard deviation. The feature vectors are combined to form a feature matrix X. Each sound segment also has a corresponding label indicating whether it belongs to environmental noise or a broadcast signal. These labels form a target vector y. The random forest is an ensemble learning algorithm composed of multiple decision trees. First, the number of decision trees in the random forest is determined. For training each decision tree, a subset of samples with replacement is randomly selected from the feature matrix X and the target vector y to construct a subset. Simultaneously, when splitting at each node, a subset of features is randomly selected for consideration instead of using all features. This increases the diversity of the decision trees and reduces the risk of overfitting. Each decision tree starts from the root node and grows continuously based on the selected features and splitting conditions until certain stopping conditions are met, such as reaching the maximum depth or the number of samples in a node being less than a certain threshold.
[0063] For a newly received sound segment, its feature vector is input into the trained random forest. Each decision tree classifies the segment and provides a corresponding prediction result. The final classification result is determined by voting from all decision trees. If most decision trees believe that the segment is a broadcast signal, the segment is classified as a broadcast signal. Conversely, if most decision trees believe that the segment is ambient noise, it is classified as ambient noise.
[0064] Based on the classification results, the sound segments are labeled and stored in their respective datasets. The mixed sound data is then processed using the independent component analysis algorithm combined with the classification results of random forest as prior knowledge to improve the accuracy of separation.
[0065] The mixed audio data is separated using filters. For the broadcast signal, a bandpass filter is used to filter the dataset labeled "broadcast," retaining only components within the broadcast signal's frequency range and removing mixed noise components to obtain the first broadcast signal. For the dataset labeled "noise," a bandstop filter is used to block the broadcast signal components, retaining only the noise portion to obtain the second ambient noise. The signal-to-noise ratio (SNR) of the broadcast signal dataset and the separated noise components is compared; a higher SNR indicates better separation.
[0066] More specifically, the measurement of the average sound pressure level includes: dividing the second ambient noise and the first broadcast signal into time windows, and calculating the average sound pressure level of all time serial ports.
[0067] In a preferred embodiment, an initial time window size is first determined. This size can be selected based on a preliminary analysis of the sound signal. For example, a relatively small time window can be set to more sensitively capture rapid changes in the signal at the beginning. Real-time analysis is performed on the separated second ambient noise and the first broadcast signal. For the second ambient noise, the total energy of the noise signal within each time window is calculated. If the energy changes significantly within several consecutive windows, it indicates that the noise has changed significantly during this period, and the window size can be reduced for more refined analysis. If the energy changes more than a preset change threshold within several adjacent windows, the window size is halved. The amplitude changes of the noise signal are monitored. When the amplitude changes drastically, it indicates that the noise situation may have changed significantly, and the window size can also be reduced for better capture. For the first broadcast signal, the frequency stability of the broadcast signal is analyzed. When the frequency changes significantly in a short period of time, the window can be reduced to track frequency dynamics. For example, when the main frequency or main frequency range shifts significantly within several consecutive windows, the window size can be adjusted. Considering the rhythm and periodicity of the broadcast signal, if the rhythm suddenly speeds up or slows down, the window size is adaptively adjusted to adapt to signal changes.
[0068] The root mean square (RMS) value is calculated for the dynamically adjusted noise signal time window. The formula for calculating the RMS value is as follows: Where, n i Let k be the sampled value of the noise signal within the corresponding window, and k be the number of sampling points within that window. Convert the corresponding root mean square (RMS) value to sound pressure level using the following formula: The sound pressure level (SPL) calculation results of the noise signal are continuously updated as the window is dynamically adjusted. A weighted average of the SPL for all dynamically adjusted noise signal time windows is then calculated. The weights can be determined based on the window size or the stability of the noise signal; for example, smaller windows are given higher weights because they better reflect rapid noise changes. The formula for calculating the average SPL is: Among them, L pni w represents the noise sound pressure level within multiple time windows. ni The corresponding weights.
[0069] For the dynamically adjusted broadcast signal time window, the root mean square (RMS) value is also calculated. The formula for calculating the root mean square (RMS) value is as follows: Among them, b i The sampled value of the noise signal within the corresponding window is given by , where l is the number of sampling points within that window. The corresponding root mean square (RMS) value is converted to sound pressure level using the following formula: The sound pressure level (SPL) calculation results of the broadcast signal are continuously updated as the window dynamically adjusts. A weighted average of the SPL across all windows of the broadcast signal is calculated. The formula for calculating the average SPL of the broadcast signal is: Among them, L pbi w represents the sound pressure level of the broadcast signal within multiple time windows. bi The corresponding weights are used. By combining dynamic window segmentation and signal separation, the average sound pressure level of environmental noise such as background noise and broadcast signals can be accurately measured. It should be noted that when the broadcast system is not working, the average sound pressure level of the first environmental noise can be directly obtained through the average sound pressure level formula.
[0070] Step S03: Based on the acquired ambient noise and the average sound pressure level of the broadcast signal, the type of broadcast content, the collected environmental data, and the reinforcement learning algorithm, adaptively adjust the optimal parameter strategy corresponding to the broadcast, including:
[0071] The environmental state of the reinforcement learning algorithm is defined based on the current ambient noise and the average sound pressure level of the broadcast signal, the type of broadcast content, and the collected environmental data; wherein, the collected environmental data includes the flow of people and time within the partition;
[0072] The action space of the reinforcement learning algorithm is defined according to the broadcast parameters; wherein, the action space includes: adjusting the volume, gain ratio and corresponding frequency response of the broadcast signal;
[0073] Based on the environmental state, the action space, and the reward function in the reinforcement learning algorithm, determine the reward value corresponding to the specific action currently being performed on the broadcast signal parameters;
[0074] Based on the reward value, the current parameter strategy of the broadcast signal is optimized to obtain the optimal parameter strategy.
[0075] In one specific embodiment, a Q-value table is initialized for all possible combinations of environmental states and actions. The environmental states include multiple dimensions such as average sound pressure level, broadcast content type, pedestrian traffic level, and time. The action space includes volume level, gain ratio, and frequency response type. Initially, all values in the Q-value table are set to a small random value. The environmental states, factors in the action space, and actions are quantified and encoded for processing in the reinforcement learning algorithm. The metrics used to determine the influence of the reward function may include broadcast intelligibility, signal-to-noise ratio, and adaptability to pedestrian traffic and time.
[0076] More specifically, if the broadcast content is audio, speech recognition software can be used to assess the proportion of words or sentences that listeners can correctly recognize under different broadcast parameters. Higher speech recognition accuracy under specific actions results in higher rewards. When the broadcast content contains important information, the broadcast signal must stand out against ambient noise to ensure listeners can hear it. This is reflected by calculating the signal-to-noise ratio (SNR) between the broadcast signal and ambient noise; a higher SNR indicates a more prominent broadcast signal relative to ambient noise. Different SNR thresholds can be set, with corresponding rewards based on the actual SNR. For example, a higher SNR than a preset first threshold results in a higher reward; a lower SNR than a preset second threshold results in a lower reward or penalty. Furthermore, corresponding SNR thresholds can be set based on the type of broadcast content; for example, for music, the corresponding thresholds are the third and fourth thresholds. In areas with high foot traffic, ambient noise usually increases, requiring increased broadcast volume or adjustments to other parameters to ensure audibility. Different rewards can be given based on the level of foot traffic. For example, as pedestrian traffic increases from low to high, the reward value is gradually increased to encourage the algorithm to increase the broadcast volume or gain ratio. Environmental noise levels and people's sensitivity to broadcasts may differ at different times of day. For instance, during the day, environmental noise may be higher, requiring a higher volume; while at night, people may prefer a lower volume to avoid disturbing others. Rewards can be given based on time categories. For example, at a specific time, such as at night, lowering the volume might earn a reward, while increasing the volume during the day might earn a reward. The reward function R can be expressed by the following formula:
[0077] R=w1×C+w2×S+w3×P+w4×T;
[0078] Where R is the total reward value, C is the reward related to signal-to-noise ratio, S is the reward related to broadcast intelligibility, P is the reward related to pedestrian traffic, and T is the reward related to time; w1, w2, w3, and w4 are the corresponding weighting coefficients used to adjust the importance of each factor in the reward function. These weighting coefficients can be dynamically adjusted according to actual conditions. For example, if broadcast intelligibility becomes more important during a certain time period, the weight of the relevant reward can be increased. The reward function parameters are adaptively adjusted and updated to adapt to changes in the environment and needs. When determining the reward function, the balance between multiple objectives should be considered. Multi-objective optimization techniques, such as the Pareto principle, can be used to find the optimal trade-off between different objectives.
[0079] Based on this, a specific action is randomly selected from the action space as the starting action, and a greedy algorithm is used for balance exploration. The selected starting action is executed, and the new environment state generated after the starting action is observed in real time. The reward value after each specific action in the action space is calculated according to the set reward function. Considering the impact of each specific action in the action space on the environment state, risk penalty or reward can be added when updating the reward value Q. The Q value is updated using the following formula:
[0080] Q(s,a)=Q(s,a)+α*(r+γ*max(Q(s',a'))-Q(s,a));
[0081] Where Q(s,a) represents the reward value of taking action a in state s, and is an estimate of the long-term expected cumulative reward of the specific action taken in this environmental state; α represents the learning rate; r represents the immediate reward obtained after taking action a; γ represents the discount factor, used to measure the importance of future rewards; max(Q(s',a')) represents the Q value of the action with the largest Q value among all possible actions a' in the new state s', representing the expected reward estimate of the best future action; based on this, the reward value Q is continuously updated until the reward value Q stabilizes, at which point the optimal parameter strategy for the broadcast signal is output. When the environmental state changes, the volume, gain ratio, and frequency response of the broadcast signal are adjusted according to the optimal strategy determined by the Q value table to achieve adaptive broadcast parameter adjustment.
[0082] S04: Adjusting the current parameters of the broadcast signal to the parameters corresponding to the optimal parameter strategy based on the optimal parameter strategy, including:
[0083] The optimal parameter strategy is sent to the broadcast to be adjusted, and the volume, gain ratio and frequency response corresponding to the broadcast to be adjusted in the optimal parameter strategy are determined.
[0084] Based on the dynamic adjustment algorithm, the current parameters of the broadcast to be adjusted are adjusted to the parameters corresponding to the optimal parameter strategy.
[0085] It should be noted that after receiving the optimal parameter strategy, the broadcast to be adjusted needs to parse the strategy information to determine the volume, gain ratio, and frequency response to be used. For example, the strategy may instruct the broadcast in a specific area to set the volume to a medium level (e.g., assuming the volume range is 0-10, and the medium level is 5), the gain ratio to 1.5, and to enhance the frequency response in the mid-to-low frequency band to improve speech clarity.
[0086] The dynamic adjustment algorithm aims to smoothly adjust the current parameters of the broadcast to the parameters specified by the optimal parameter strategy, avoiding sudden changes that may cause discomfort to the audience. In this embodiment, a linear interpolation-based method is used, with each adjustment increasing or decreasing by an increment ΔP, which is gradually increased over a certain period of time. The parameters of the broadcast signal are effectively adjusted according to the optimal parameter strategy to adapt to different environments and broadcast needs, thereby improving the broadcast effect and the audience experience.
[0087] This invention provides a method, apparatus, device, medium, and product for adaptive adjustment of broadcast volume. It acquires sound signals from different zones, preprocesses the acquired zone sound signals to obtain zone sound data, extracts features from the acquired sound data to obtain comprehensive features, and uses these comprehensive features to more effectively separate noise and broadcast signals. It measures the average sound pressure level (SPL) of the separated environmental noise and broadcast signal, using the SPL data as a basis for adjusting broadcast parameters and environmental control, ensuring the quality of sound propagation and helping to evaluate the propagation effect of the broadcast signal. Combining the acquired data on the average SPL of the environmental noise within the zone and the broadcast signal, as well as the type of broadcast content, it determines the corresponding optimal parameter strategy using a reinforcement learning algorithm. Based on dynamic adjustment technology, it adaptively and dynamically adjusts the broadcast parameters to ensure a satisfactory auditory experience for listeners in complex acoustic environments, ensures the best propagation effect of the broadcast signal under different conditions, and guarantees listeners' access to key information.
[0088] Example 2
[0089] Please refer to Figure 2 This is a schematic diagram of a broadcast volume adaptive adjustment device provided in an embodiment of this application. The device can be implemented by software and / or hardware and is generally integrated into any electronic device with network communication capabilities, including but not limited to: servers, computers, personal digital assistants, etc. Figure 2 The present application provides a broadcast volume adaptive adjustment device, which includes: a real-time acquisition module 201, a separation measurement module 202, a strategy determination module 203, and a broadcast adjustment module 204.
[0090] The real-time acquisition module 201 is used to acquire partition sound signals and preprocess the acquired partition sound signals to obtain partition sound data.
[0091] The separation measurement module 202 extracts comprehensive features from the acquired sound data, separates the ambient noise and broadcast signal based on the acquired comprehensive features, and measures the corresponding average sound pressure level of the separated ambient noise and broadcast signal respectively.
[0092] The strategy determination module 203 is used to adaptively adjust the optimal parameter strategy corresponding to the broadcast based on the acquired environmental noise and the average sound pressure level of the broadcast signal, the type of broadcast content, the collected environmental data and the reinforcement learning algorithm.
[0093] The broadcast adjustment module 204 is used to adjust the current parameters of the broadcast signal to the parameters corresponding to the optimal parameter strategy according to the optimal parameter strategy.
[0094] Example 3
[0095] Figure 3 This is a schematic diagram of an electronic device provided in an embodiment of the present invention. The electronic device 10 is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices (e.g., helmets, glasses, watches, etc.), and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the invention described and / or claimed herein.
[0096] like Figure 3 As shown, the electronic device 10 includes at least one processor 11 and a memory, such as a read-only memory (ROM) 12 or a random access memory (RAM) 13, communicatively connected to the at least one processor 11. The memory stores computer programs executable by the at least one processor. The processor 11 can perform various appropriate actions and processes based on the computer program stored in the ROM 12 or loaded from storage unit 18 into the RAM 13. The RAM 13 may also store various programs and data required for the operation of the electronic device 10. The processor 11, ROM 12, and RAM 13 are interconnected via a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.
[0097] Multiple components in electronic device 10 are connected to I / O interface 15, including: input unit 16, such as keyboard, mouse, etc.; output unit 17, such as various types of displays, speakers, etc.; storage unit 18, such as disk, optical disk, etc.; and communication unit 19, such as network card, modem, wireless transceiver, etc. Communication unit 19 allows electronic device 10 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.
[0098] Processor 11 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. Processor 11 performs the various methods and processes described above, such as the automatic frequency modulation method for a wireless microphone.
[0099] In some embodiments, the automatic broadcast volume adjustment method may be implemented as a computer program tangibly contained in a computer-readable storage medium, such as storage unit 18. In some embodiments, part or all of the computer program may be loaded and / or installed on electronic device 10 via ROM 12 and / or communication unit 19. When the computer program is loaded into RAM 13 and executed by processor 11, one or more steps of the automatic broadcast volume adjustment method described above may be performed. Alternatively, in other embodiments, processor 11 may be configured by any other suitable means (e.g., by means of firmware) to perform an automatic frequency modulation method for a wireless microphone.
[0100] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), payload-programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.
[0101] Computer programs used to implement the methods of the present invention may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when executed by the processor, the computer programs cause the functions / operations specified in the flowcharts and / or block diagrams to be performed. The computer programs may be executed entirely on a machine, partially on a machine, or as a standalone software package, partially on a machine and partially on a remote machine, or entirely on a remote machine or server.
[0102] In the context of this invention, a computer-readable storage medium can be a tangible medium that may contain or store a computer program for use by or in conjunction with an instruction execution system, apparatus, or device. A computer-readable storage medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination thereof. Alternatively, a computer-readable storage medium may be a machine-readable signal medium. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.
[0103] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the electronic device. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).
[0104] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as data servers), or computing systems that include middleware components (e.g., application servers), or computing systems that include frontend components (e.g., user computers with graphical user interfaces or web browsers through which users can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., communication networks). Examples of communication networks include local area networks (LANs), wide area networks (WANs), blockchain networks, and the Internet.
[0105] A computing system can include clients and servers. Clients and servers are generally located far apart and typically interact through communication networks. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or cloud host, which is a hosting product within the cloud computing service system to address the shortcomings of traditional physical hosts and VPS services, such as high management difficulty and weak business scalability.
[0106] Example 4
[0107] This application also provides a computer program product, including a computer program that, when executed by a processor, implements the broadcast volume adaptive adjustment method as provided in any embodiment of this application. In implementing the computer program product, computer program code for performing the operations of this invention can be written in one or more programming languages or a combination thereof. Programming languages include object-oriented programming languages such as Java, Smalltalk, and C++, as well as conventional procedural programming languages such as "C" or similar languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0108] It should be understood that the various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps described in this invention can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution of this invention can be achieved, and this is not limited herein.
[0109] The specific embodiments described above further illustrate the purpose, technical solution, and beneficial effects of the present invention. It should be understood that the above descriptions are merely specific embodiments of the present invention and are not intended to limit the scope of protection of the present invention. In particular, it should be noted that any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention for those skilled in the art.
Claims
1. A method for adaptive adjustment of broadcast volume, characterized in that, include: Acquire the partition sound signal, and preprocess the acquired partition sound signal to obtain partition sound data; Feature extraction is performed on the acquired partitioned sound data to obtain comprehensive features. Based on the acquired comprehensive features, environmental noise and broadcast signals are separated, and the corresponding average sound pressure levels of the separated environmental noise and broadcast signals are measured respectively. The comprehensive features include spectral features, time-domain features, and sound signal energy features. Based on the acquired ambient noise and the average sound pressure level of the broadcast signal, the type of broadcast content, the collected environmental data, and the reinforcement learning algorithm, the optimal parameter strategy corresponding to the broadcast is adaptively adjusted. This adaptive adjustment of the optimal parameter strategy includes: defining the environmental state of the reinforcement learning algorithm based on the current ambient noise and the average sound pressure level of the broadcast signal, the type of broadcast content, and the collected environmental data; wherein the collected environmental data includes the pedestrian flow and time within the area; and defining the action space of the reinforcement learning algorithm according to the broadcast signal parameters; wherein the action space includes: adjusting the volume, gain, and corresponding frequency response of the broadcast signal. Based on the environmental state, the action space, and the reward function in the reinforcement learning algorithm, determine the reward value corresponding to the current specific action of the broadcast signal parameters; optimize the current parameter policy of the broadcast signal based on the reward value to obtain the optimal parameter policy; adjust the current parameters of the broadcast signal to the parameters corresponding to the optimal parameter policy based on the optimal parameter policy; adjusting the current parameters of the broadcast signal to the parameters corresponding to the optimal parameter policy based on the optimal parameter policy includes: sending the optimal parameter policy to the broadcast to be adjusted, and determining the volume, gain, and frequency response corresponding to the broadcast to be adjusted in the optimal parameter policy; adjusting the current parameters of the broadcast to be adjusted to the parameters corresponding to the optimal parameter policy based on a dynamic adjustment algorithm.
2. The method for adaptive adjustment of broadcast volume according to claim 1, characterized in that, The separation of environmental noise and broadcast signals based on the acquired comprehensive features specifically includes the following steps: When the broadcast system is not working, the partition sound data is the first ambient noise of the corresponding partition; When the broadcast system is working, the extracted comprehensive features are input into a preset sound classification model to obtain classification information for the regional sound data; For the parts classified as noise, residual broadcast signal components are removed through filtering. For the portion classified as broadcast signal, residual noise components are removed through filtering. The partitioned sound data and the corresponding classification label results are input into the separation algorithm, which outputs the separated second environmental noise and the first broadcast signal.
3. The method for adaptive adjustment of broadcast volume according to claim 2, characterized in that, The measurement of the average sound pressure level specifically includes: dividing the second ambient noise and the first broadcast signal into time windows, and calculating the average sound pressure level corresponding to all time windows.
4. A broadcast volume adaptive adjustment device, characterized in that, It includes a real-time acquisition module, a separation measurement module, a strategy determination module, and a broadcast adjustment module; The real-time acquisition module is used to acquire partitioned sound signals and preprocess the acquired partitioned sound signals to obtain partitioned sound data. The separation measurement module is used to extract features from the acquired partitioned sound data to obtain comprehensive features, separate environmental noise and broadcast signals based on the acquired comprehensive features, and measure the corresponding average sound pressure level of the separated environmental noise and broadcast signals respectively; the comprehensive features include spectral features, time-domain features, and sound signal energy features; The strategy determination module is used to adaptively adjust the optimal parameter strategy corresponding to the broadcast based on the acquired ambient noise and the average sound pressure level of the broadcast signal, the type of broadcast content, the collected environmental data, and the reinforcement learning algorithm. The adaptive adjustment of the optimal parameter strategy based on the acquired ambient noise and the average sound pressure level of the broadcast signal, the type of broadcast content, the collected environmental data, and the reinforcement learning algorithm includes: defining the environmental state of the reinforcement learning algorithm based on the current ambient noise and the average sound pressure level of the broadcast signal, the type of broadcast content, and the collected environmental data; wherein the collected environmental data includes the pedestrian flow and time within the partition; defining the action space of the reinforcement learning algorithm according to the broadcast signal parameters; wherein the action space includes: adjusting the volume, gain, and corresponding frequency response of the broadcast signal; determining the reward value corresponding to the current specific action of the broadcast signal parameters based on the environmental state, the action space, and the reward function in the reinforcement learning algorithm; and optimizing the current parameter strategy of the broadcast signal based on the reward value to obtain the optimal parameter strategy. The broadcast adjustment module is used to adjust the current parameters of the broadcast signal to the parameters corresponding to the optimal parameter strategy based on the optimal parameter strategy. The adjustment of the current parameters of the broadcast signal to the parameters corresponding to the optimal parameter strategy based on the optimal parameter strategy includes: sending the optimal parameter strategy to the broadcast to be adjusted, and determining the volume, gain, and frequency response corresponding to the broadcast to be adjusted in the optimal parameter strategy; and adjusting the current parameters of the broadcast to be adjusted to the parameters corresponding to the optimal parameter strategy based on a dynamic adjustment algorithm.
5. An electronic device, characterized in that, include: One or more processors; Storage device for storing one or more programs; When the one or more programs are executed by the one or more processors, the one or more processors implement a broadcast volume adaptive adjustment method as described in any one of claims 1-3.
6. A storage medium containing computer-executable instructions, characterized in that, The computer-executable instructions, when executed by a computer processor, are used to perform a broadcast volume adaptive adjustment method as described in any one of claims 1-3.
7. A computer program product comprising a computer program that, when executed by a processor, implements a broadcast volume adaptive adjustment method according to any one of claims 1-3.
Citation Information
Patent Citations
Method and device for intelligent perception on television
CN103686009A
Speech separation method and device, mobile terminal and computer readable storage medium
CN110808061A
Volume adjustment method and device, electronic equipment and storage medium
CN113126952A