A sound field control method, system, device and computer readable storage medium
Patent Information
- Application Number
- CN202610789137.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-02
- Publication Date
- 2026-08-18
AI Technical Summary
现有技术仍多依赖预设的静态滤波参数,无法针对发声主体的生理特性及声学表现进行动态补偿,导致音频还原度不佳,难以满足用户对高保真、个性化音质的期待
Smart Images

Figure CN122598679A_ABST
Abstract
Description
Technical Field
[0001] This specification relates to the field of audio system control, and in particular to a sound field control method, system, device, and computer-readable storage medium. Background Technology
[0002] With the popularization of voice interaction technology, microphone acquisition systems have been widely used in smart homes, conference systems, professional recording and other fields. However, traditional microphone sound field design often adopts fixed or single frequency response curve settings (i.e., fixed gain mode), which makes it difficult to adapt to the individual differences of different voice-producing subjects.
[0003] For example, there are significant differences in vocal tract characteristics among individuals, not only in fundamental frequency fluctuations caused by gender, but also widely existing across different age groups, pronunciation habits, and even artistic voices. In single-person and multi-person application scenarios, due to the lack of in-depth analysis of real-time timbre characteristics, fixed gain mode is prone to information redundancy or perceptual loss in certain frequency bands of the audio signal (such as harsh high frequencies or thin low frequencies). In addition, in dynamic environments where multiple people take turns speaking, fixed gain mode cannot dynamically increase the gain according to the timbre characteristics of the current speaker, making it difficult to maintain the naturalness and consistency of speech.
[0004] Furthermore, with the deep integration of voice interaction and remote collaboration technologies, microphone acquisition systems are placing higher demands on the fine-grained processing of sound quality. Existing technologies still largely rely on preset static filtering parameters, failing to dynamically compensate for the physiological characteristics and acoustic performance of the speaker, resulting in poor audio fidelity and failing to meet users' expectations for high-fidelity, personalized sound quality.
[0005] Therefore, there is an urgent need for a sound field control method, system, device, and computer-readable storage medium that can adapt to the individual differences of different sound-producing subjects and achieve dynamic optimization of the sound field. Summary of the Invention
[0006] This specification provides one or more embodiments of a sound field control method, comprising: acquiring an audio signal of a sound-producing subject; generating audio features characterizing the sound-producing attributes of the sound-producing subject based on the audio signal; generating sound field control parameters based on the audio features; and controlling an audio device to output audio based on the sound field control parameters.
[0007] This specification provides one or more embodiments of a sound field control system, comprising: an acquisition module configured to acquire an audio signal of a sound-producing subject; a first generation module configured to generate audio features characterizing the sound-producing attributes of the sound-producing subject based on the audio signal; a second generation module configured to generate sound field control parameters based on the audio features; and a control module configured to control an audio device to output audio based on the sound field control parameters.
[0008] This specification provides one or more embodiments of a sound field control device, including: at least one processor and at least one memory; the at least one memory is used to store computer instructions; and the at least one processor is used to execute at least a portion of the computer instructions to implement the sound field control method.
[0009] This specification provides one or more embodiments of a computer-readable storage medium that stores computer instructions. When a computer reads the computer instructions from the storage medium, the computer executes the sound field control method. Attached Figure Description
[0010] This specification will be further described by way of exemplary embodiments, which will be described in detail with reference to the accompanying drawings. These embodiments are not limiting; in these embodiments, the same reference numerals denote the same structures, wherein:
[0011] Figure 1 These are schematic diagrams illustrating application scenarios of the sound field control system according to some embodiments of this specification; Figure 2 This is an exemplary block diagram of a sound field control system according to some embodiments of this specification; Figure 3 This is one of the exemplary flowcharts of a sound field control method according to some embodiments of this specification; Figure 4 This is a second exemplary flowchart of a sound field control method according to some embodiments of this specification; Figure 5 This is the third exemplary flowchart of a sound field control method according to some embodiments of this specification; Figure 6 This is an exemplary schematic diagram illustrating the generation of emotional fluctuation information according to some embodiments of this specification. Detailed Implementation
[0012] To more clearly illustrate the technical solutions of the embodiments in this specification, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are merely some examples or embodiments of this specification. For those skilled in the art, these drawings can be applied to other similar scenarios without creative effort. Unless obvious from the context or otherwise specified, the same reference numerals in the drawings represent the same structures or operations.
[0013] It should be understood that the terms “system,” “device,” “unit,” and / or “module” used herein are one way to distinguish different components, elements, parts, sections, or assemblies at different levels. However, if other terms can achieve the same purpose, they may be replaced by other expressions.
[0014] As indicated in this specification and claims, unless the context clearly indicates otherwise, the words "a," "an," "an," and / or "the" do not specifically refer to the singular and may also include the plural. Generally speaking, the terms "comprising" and "including" only indicate the inclusion of expressly identified steps and elements, which do not constitute an exclusive list, and the method or apparatus may also include other steps or elements.
[0015] Flowcharts are used in this specification to illustrate the operations performed by the system according to embodiments of this specification. It should be understood that the preceding or following operations are not necessarily performed in exact order. Instead, the steps can be processed in reverse order or simultaneously. Furthermore, other operations can be added to these processes, or one or more steps can be removed from them.
[0016] Figure 1 This is a schematic diagram illustrating an application scenario of a sound field control system according to some embodiments of this specification. In some embodiments, such as Figure 1 As shown, the application scenario 100 of the sound field control system (hereinafter referred to as application scenario 100) may include an audio acquisition device 110, a processor 120, a network 130, and an audio playback device 140.
[0017] In some embodiments, application scenario 100 may include various voice acquisition scenarios that require dynamic adaptation to speaker characteristics, such as voice interaction between a single speaker or multiple speakers. Examples include conference systems, cinemas, concerts, broadcasting systems, and lecterns.
[0018] The audio acquisition device 110 is used to acquire audio signals and convert them into digital or electrical signals that can be processed by a processor. For example, the audio acquisition device may include a microphone, a microphone array, etc. In some embodiments, the audio acquisition device 110 can be flexibly arranged according to the application scenario 100. For example, the audio acquisition device 110 can be arranged on a conference room table, in a vehicle ceiling, around a stage, or inside a user's smart device (such as a mobile phone, tablet, smart speaker, VR headset, etc.).
[0019] The processor 120 can communicate with the audio acquisition device 110 and the audio playback device 140 to process data and / or information acquired from the audio acquisition device 110. In some embodiments, the processor 120 can be a local or remote component relative to the sound field control system. In some embodiments, the processor 120 can be a single server or a group of servers. In some embodiments, the processor 120 can include one or more processors (e.g., a single-chip processor or a multi-chip processor). In some embodiments, the processor 120 (or all or part of its functions) can be part of the audio playback device 140. In some embodiments, the processor can process the audio signals acquired by the audio acquisition device 110, generate sound field control parameters, and send the generated control parameters to the audio playback device 140.
[0020] Network 130 can connect audio acquisition device 110, processor 120, audio playback device 140, and / or connect to external resources. Network 130 enables communication between audio acquisition device 110, processor 120, audio playback device 140, and / or external resources, facilitating the exchange of data and / or information. In some embodiments, network 130 can be any one or more of wired or wireless networks. For example, network 130 may include cable networks, fiber optic networks, telecommunications networks, the Internet, local area networks (LANs), wide area networks (WANs), wireless local area networks (WLANs), metropolitan area networks (MANs), public switched telephone networks (PSTNs), Bluetooth networks, ZigBee networks, near field communication (NFC), device internal buses, device internal wiring, cable connections, etc., or any combination thereof.
[0021] Audio playback device 140 is used to play audio. In some embodiments, audio playback device 140 may be a speaker system (such as a speaker cabinet) including filters (such as digital equalization filters), wearable audio devices (such as headphones, smart helmets, etc.), etc. In some embodiments, audio playback device 140 may receive sound field control parameters generated by a processor and convert audio signals into sound output according to the sound field control parameters, thereby realizing control over sound field spatial distribution, frequency response, dynamic range, etc.
[0022] In some embodiments, the audio acquisition device 110 can acquire the audio signal of the sound-producing subject and transmit the audio signal to the processor 120 through the network 130; the processor 120 can extract the audio features in the audio signal, thereby generating sound field control parameters, and send the sound field control parameters to the audio playback device 140 through the network 130; the audio playback device 140 plays the audio signal based on the sound field control parameters to achieve dynamic compensation of the physiological characteristics and acoustic performance of the sound-producing subject.
[0023] For further explanation of the aforementioned audio signals, sound-producing entities, audio characteristics, and sound field control parameters, please refer to [link to relevant documentation]. Figure 3 And its related descriptions.
[0024] It should be noted that the above description of application scenario 100 is for convenience only and should not be construed as limiting this specification to the scope of the illustrated embodiments. It is understood that those skilled in the art, after understanding the principles of this system, may arbitrarily combine the various modules or construct subsystems connected to other modules without departing from these principles. In some embodiments, the modules may share a single storage module, or each module may have its own separate storage module. Such modifications are all within the scope of this specification.
[0025] This specification provides a sound field control system through several embodiments. The sound field control system can be widely applied to various speech acquisition scenarios requiring dynamic adaptation to speaker characteristics, covering fields such as intelligent conferencing and remote communication, smart speakers and voice interaction, hearing aids and hearing assistance, broadcasting and professional recording, and in-vehicle voice systems. By recognizing the speaker's timbre characteristics in real time and automatically adjusting the microphone's frequency response curve, the sound field control system can maintain clear speech and a natural listening experience even in complex dynamic environments where different genders and voices alternate, significantly improving user experience and system adaptability.
[0026] Figure 2 This is an exemplary block diagram of a sound field control system according to some embodiments of this specification.
[0027] In some embodiments, such as Figure 2 As shown, the sound field control system 200 includes an acquisition module 211, a first generation module 212, a second generation module 213, and a control module 214.
[0028] In some embodiments, the acquisition module 211 is configured to acquire the audio signal of the sound-producing body.
[0029] In some embodiments, the acquisition module 211 can acquire the audio signal of the sound-producing subject collected by the audio acquisition device (such as the audio acquisition device 110) in real time, and convert the audio signal into a digital audio stream and transmit it to the first generation module 212.
[0030] In some embodiments, the first generation module 212 is configured to generate audio features that characterize the vocal attributes of the vocal subject based on the audio signal.
[0031] In some embodiments, the first generation module 212 may include at least a spectrum analysis unit (not shown in the figure) and a timbre recognition unit (not shown in the figure), wherein a neural network model is configured on the timbre recognition unit.
[0032] The neural network model can be composed of a gender extraction model and a feature extraction model. For more information on gender extraction and feature extraction models, please refer to [link to relevant documentation]. Figure 3 The description in step 320.
[0033] In some embodiments, the spectrum analysis unit can perform a Fast Fourier Transform (FFT) on the audio signal in real time to obtain the audio signal's spectrum, extract the frequency distribution within the spectrum, and thus determine data such as high-frequency bands and low-frequency bands. For further explanation of high-frequency bands and low-frequency bands, see [link to relevant documentation]. Figure 3 The relevant description in step 330.
[0034] In some embodiments, the timbre recognition unit may use a preset algorithm to analyze the spectrum of the audio signal in real time, determine the acoustic features of the audio signal, input the acoustic features into the neural network model, and then determine the audio features that characterize the vocal attributes of the vocal subject.
[0035] The preset algorithm may include the Mel-Frequency Cepstral Coefficients (MFCC) algorithm, and the time-frequency characteristics may include MFCC.
[0036] In some embodiments, the second generation module 213 is configured to generate sound field control parameters based on audio features.
[0037] In some embodiments, the second generation module 213 can calculate the gain or attenuation of each frequency band of each audio signal based on audio characteristics, and then generate sound field control parameters.
[0038] In some embodiments, the control module 214 is configured to control the audio device to output audio based on sound field control parameters.
[0039] In some embodiments, the control module 214 can dynamically configure the parameters of the digital equalizer filter in the audio device, thereby adjusting each frequency band for audio output.
[0040] For further explanation of the above content, please refer to Figure 3 And its related descriptions.
[0041] In some embodiments, such as Figure 2 As shown, the sound field control system 200 also includes a marker generation module 221, an energy determination module 222, a first adjustment module 223, and a transition processing module 224.
[0042] In some embodiments, the tag generation module 221 is configured to tag the sound-emitting subject and generate a sound-emitting subject tag.
[0043] In some embodiments, the energy determination module 222 is configured to determine the energy percentage of the sound-producing entity in the current audio signal.
[0044] In some embodiments, the first adjustment module 223 is configured to adjust the sound field control parameters corresponding to different audio frequency bands based on the energy ratio.
[0045] In some embodiments, the transition processing module 224 is configured to generate transition parameters within a preset time window when a switch of the sound-generating subject mark is detected, and to perform smooth transition processing on the sound field control parameters of the corresponding audio frequency band based on the transition parameters.
[0046] For more information on this section, please see [link / reference]. Figure 4 And its related descriptions.
[0047] In some embodiments, such as Figure 2 As shown, the sound field control system 200 also includes a division module 231, a verification module 232, an emotion determination module 233, and a second adjustment module 234.
[0048] In some embodiments, the segmentation module 231 is configured to segment audio features into identity audio features and emotional audio features.
[0049] In some embodiments, the verification module 232 is configured to identify the speaker's identifier based on the speaker's identity audio features and perform consistency verification of the speaker.
[0050] In some embodiments, the emotion determination module 233 is configured to determine the emotional fluctuation information of the speaker based on the emotional audio features of the speaker that have passed the consistency check.
[0051] In some embodiments, the emotion determination module 233 is further configured to: extract duration features, energy envelope features, and fundamental frequency perturbation values of the vocal subject in the audio signal; determine a first fluctuation amplitude of the duration features, a second fluctuation amplitude of the energy envelope features, and a third fluctuation amplitude of the fundamental frequency perturbation values; and generate emotion fluctuation information based on the first fluctuation amplitude, the second fluctuation amplitude, and the third fluctuation amplitude. For more information on this section, please refer to [link / reference]. Figure 6 And its related descriptions.
[0052] In some embodiments, the second adjustment module 234 is configured to adjust the sound field control parameters of different audio frequency bands based on emotional fluctuation information.
[0053] For more information on this section, please see [link / reference]. Figure 5 And its related descriptions.
[0054] In some embodiments, part or all of the sound field control system 200 may be integrated into the processor 120.
[0055] It should be noted that the above description of the sound field control system and its modules is for convenience only and should not be construed as limiting this specification to the scope of the illustrated embodiments. It is understood that those skilled in the art, after understanding the principles of the system, may arbitrarily combine the various modules or construct subsystems connected to other modules without departing from these principles. In some embodiments, Figure 2 The acquisition module 211, first generation module 212, second generation module 213, control module 214, tag generation module 221, energy determination module 222, first adjustment module 223, transition processing module 224, division module 231, verification module 232, emotion determination module 233, and second adjustment module 234 disclosed herein can be different modules within a single system, or a single module can implement the functions of two or more of the aforementioned modules. For example, the modules can share a single storage module, or each module can have its own separate storage module. Such variations are all within the scope of protection of this specification.
[0056] This specification provides an embodiment of a sound field control method.
[0057] Figure 3 This is one of the exemplary flowcharts of a sound field control method according to some embodiments of this specification. Figure 3 As shown, process 300 includes the following steps. In some embodiments, process 300 may be executed by a generation control system (such as sound field control system 200) or a processor (such as processor 120). The following description uses a processor to execute the sound field control method.
[0058] Step 310: Obtain the audio signal of the sound-producing subject.
[0059] The source of sound refers to the object that emits the sound. For example, sources of sound include people, cars, pets, and electrical appliances.
[0060] An audio signal refers to a digital or electrical signal containing the sound of at least one source. In some embodiments, the processor can acquire an audio signal containing the sound of at least one source in real time using an audio acquisition device (such as audio acquisition device 110). For further details regarding audio acquisition device 110, see [link to documentation]. Figure 1 The description in the text.
[0061] The processor can also obtain the audio signal of the sound-producing body through various means, such as reading data stored in the storage device.
[0062] Step 320: Based on the audio signal, generate audio features that characterize the vocal attributes of the vocal subject.
[0063] Vocal attributes are used to characterize the physiological traits and vocal state of the vocal subject. For example, vocal attributes may include the gender of the vocal subject, timbre characteristics (such as timbre, pitch, loudness (i.e., volume)), etc.
[0064] Audio features refer to physical quantities that can characterize the physiological characteristics and vocal state of a vocal subject. For example, audio features may include the gender of the vocal subject and acoustic features that characterize the timbre of the vocal subject.
[0065] In some embodiments, acoustic features may include the fundamental frequency distribution of the audio signal (such as a sequence of fundamental frequency distribution over time, the mean fundamental frequency, fundamental frequency variation, etc.), sound pressure level, frequency distribution, energy envelope, MFCC, etc. The fundamental frequency characterizes the pitch of the sound-producing element. The sound pressure level characterizes the loudness of the sound-producing element. The MFCC reflects the spectral envelope and formant structure, and can characterize the voice signature (or timbre) of the sound-producing element.
[0066] In some embodiments, the processor can extract acoustic features that characterize the timbre of the sound-producing subject from the audio signal in a variety of ways.
[0067] For example, a processor can analyze the amplitude of an audio signal to obtain its energy envelope; convert the audio signal from a time-domain signal to a frequency-domain signal using FFT to obtain its spectrum, thereby determining the frequency distribution and extracting the sound pressure level; extract the fundamental frequency distribution from the audio signal using the autocorrelation function; analyze the spectrum of the audio signal using the MFCC algorithm to extract the MFCC; and construct acoustic features to characterize the timbre of the sound-producing subject using spectral distribution, energy envelope, sound pressure level, fundamental frequency distribution, and MFCC.
[0068] The amplitude value refers to the digital sample value after the audio signal is converted into a digital signal. It can reflect the magnitude of the instantaneous voltage (or sound pressure) of the audio signal at a certain moment. The preset time period is determined based on actual needs.
[0069] For example, the processor can also input audio signals into the feature extraction model to obtain the acoustic features output by the feature extraction model.
[0070] In some embodiments, the feature extraction model can be a machine learning model. For example, one or more of the following: a deep neural network (DNN) or recurrent neural network (RNN) model, or other custom models.
[0071] In some embodiments, the feature extraction model can be obtained from multiple first training samples with first training labels. For example, the first training samples are input into an initial feature extraction model; a loss function is constructed using the output of the initial feature extraction model and the first training labels; the parameters of the initial feature extraction model are iteratively updated using the loss function until a preset condition is met, at which point the iteration stops, resulting in a trained feature extraction model. The preset condition includes one of the following: the iteration reaches the maximum number of iterations or the loss function converges.
[0072] The first training sample may include sample audio signals, which can be obtained based on historical data. The first training label is the acoustic feature of the sample corresponding to the first training sample, which can be obtained through manual annotation.
[0073] In some embodiments, the processor can determine the gender of the speaker in a variety of ways based on acoustic characteristics.
[0074] For example, when the average fundamental frequency is higher than the first fundamental frequency threshold and the second formant (F2) frequency is higher than the first resonant frequency threshold (i.e., the timbre of the vocalist is too high), the processor determines that the vocalist is female. When the average fundamental frequency is lower than the second fundamental frequency threshold and the F2 frequency is lower than the second resonant frequency threshold (i.e., the timbre of the vocalist is too low), the processor determines that the vocalist is male.
[0075] Among them, there are a first fundamental frequency threshold, a second fundamental frequency threshold, a first resonant frequency threshold, and a second resonant frequency threshold, wherein the first fundamental frequency threshold is greater than the second fundamental frequency threshold, and the first resonant frequency threshold is greater than the second resonant frequency threshold.
[0076] For example, a processor can input MFCC into a gender determination model, and the gender determination model can output the gender of the speaker.
[0077] In some embodiments, the gender extraction model can be a machine learning model. For example, one or more of the following: a deep neural network (DNN) or recurrent neural network (RNN) model, or other custom models.
[0078] In some embodiments, the gender extraction model can be obtained from multiple second training samples with second training labels. The training process of the gender extraction model is similar to that of the feature extraction model.
[0079] The second training sample may include MFCC extracted from the sample audio signal, which can be obtained based on historical data. The second training label is the actual gender of the speaker in the sample audio signal corresponding to the second training sample, which can be obtained through manual annotation.
[0080] In some embodiments, the processor may determine audio features based on the acoustic characteristics and gender of the speaker.
[0081] Step 330: Generate sound field control parameters based on audio features.
[0082] Sound field control parameters refer to the parameters used to control the sound output of audio devices. For example, sound field control parameters include the gain (such as enhancement or suppression) and gain amplitude of energy in different frequency bands when the audio device outputs sound.
[0083] As an example only, when the sound field control parameter for frequency band A is -3dB, it means that the energy of frequency band A is suppressed by 3dB. When the sound field control parameter for frequency band B is +5dB, it means that the energy of frequency band B is enhanced by 5dB. When the sound field control parameter for frequency band C is 0dB, it means that the energy of frequency band C is neither suppressed nor enhanced.
[0084] Among these, an audio device can be a device that outputs sound. For example, an audio device can be... Figure 1 The audio playback device 140.
[0085] In some embodiments, the processor can generate sound field control parameters in a variety of ways based on audio features.
[0086] For example, the processor can construct a frequency response vector based on the fundamental frequency, spectral envelope, formant structure, sound pressure level, frequency distribution, etc. in the audio features; and search in the preset parameter library to use the reference sound field control parameters that meet the first preset requirements as the sound field control parameters corresponding to the audio feature.
[0087] In some embodiments, the audio signal acquired by the processor may include audio signals from one or more sound-producing entities. When only one sound-producing entity is included, the processor may directly use the sound field control parameters determined based on the audio characteristics of that sound-producing entity as the sound field control parameters for controlling the output of the audio device.
[0088] When multiple sound-producing entities are involved, the processor can use the sound field control parameters corresponding to the audio characteristics of the sound-producing entity with the highest energy proportion in the current audio signal, or whose energy proportion exceeds a preset threshold for a duration not less than a preset duration, as the sound field control parameters for controlling the output of the audio device. For an explanation of energy proportion, see [link to relevant documentation]. Figure 4 The corresponding content.
[0089] It should be noted that the following Figure 3 The same principle applies to other methods of determining sound field control parameters.
[0090] In some embodiments, the preset parameter library includes reference frequency response vectors and reference sound field control parameters corresponding to multiple different reference audio features. The reference frequency response vectors can be constructed based on the audio features of historical audio signals, and the specific construction method is the same as that of the frequency response vectors.
[0091] In some embodiments, the preset parameter library can be constructed based on historical or experimental data. For example, for a frequency response vector (i.e., a reference frequency response vector) corresponding to a certain audio feature in historical data, the processor can use the sound field control parameter most frequently used by users as the reference sound field control parameter corresponding to that reference frequency response vector, thereby constructing the preset parameter library. As another example, the processor can construct different reference frequency response vectors based on different audio features, and randomly generate multiple different sound field control parameters for each reference frequency response vector to control the audio device's audio output. Multiple users can then rate the parameters, and the sound field control parameter with the highest average rating can be used as the reference sound field control parameter corresponding to that reference frequency response vector.
[0092] In some embodiments, the first preset requirement is set based on actual needs. For example, the first preset requirement may include minimizing the vector distance between the frequency response vector and the reference frequency response vector, or ensuring that the vector distance is less than a distance threshold. The distance threshold is set based on experience.
[0093] For example, in response to the detection that the timbre of the main sound source is too high, the processor can automatically determine the suppression amplitude of each frequency band of the audio signal. In response to the detection that the timbre of the main sound source is too low, the processor can automatically determine the enhancement amplitude of each frequency band of the audio signal; based on the suppression amplitude and enhancement amplitude, the sound field control parameters of the audio signal are adjusted. The suppression or enhancement amplitude of each frequency band of the audio signal corresponding to different timbres can be preset based on historical experience.
[0094] For example, in response to the detection that the speaker's timbre is high and the speaker is female, the processor identifies the frequency bands in the audio signal with frequencies higher than a first frequency threshold as high-frequency bands and determines the suppression amplitude of the high-frequency bands to reduce the sharpness of the sound. In response to the detection that the speaker's timbre is low and the speaker is male, the processor can identify the frequency bands in the audio signal with frequencies lower than a second frequency threshold as low-frequency bands and determine the enhancement amplitude of the low-frequency bands to enhance the fullness of the sound.
[0095] The first frequency threshold and the second frequency threshold are set based on experience, with the first frequency threshold being greater than the second frequency threshold.
[0096] For example, in response to a female voice, the processor can automatically determine the suppression level for high-frequency bands. In response to a male voice, the processor can automatically determine the enhancement level for low-frequency bands. This setting makes the audio signal more natural and comfortable for both genders.
[0097] In some embodiments, the processor can adjust the sound field control parameters corresponding to each frequency band based on the enhancement or suppression amplitude of each frequency band. The suppression and enhancement amplitudes are determined empirically, and the suppression or enhancement amplitudes corresponding to different frequency bands can be the same or different.
[0098] Step 340: Based on the sound field control parameters, control the audio device to output audio.
[0099] In some embodiments, the processor can load the generated sound field control parameters into the register of the filter of the audio device. The filter enhances or suppresses each frequency band of the audio signal and reconstructs the spectrum distribution in real time. The digital-to-analog converter of the audio device converts the processed audio signal into an analog signal and drives the sound-producing unit (such as a loudspeaker) of the audio device to achieve physical sound production.
[0100] In some embodiments of this specification, by acquiring the audio characteristics of the sound-producing subject and dynamically generating sound field control parameters, it is possible to quantify and correct the directional enhancement and suppression compensation of individual frequency response deviations, effectively solving the problems of timbre distortion, sharpness or muddiness of audio devices with single frequency response curve settings, ensuring that the audio output maintains high recognizability and listening comfort in complex environments, and providing users with an intelligent high-fidelity sound quality experience.
[0101] It is worth noting that in complex sound fields where multiple speakers take turns speaking, the voices of multiple speakers overlap. If the system cannot identify the current speaker and uses a fixed, one-size-fits-all sound field parameter, it will lead to chaotic sound quality and the speaker's voice being drowned out. At the same time, when the speaker switches, the sound field control parameters are prone to abrupt changes, resulting in a noticeable auditory jump and seriously affecting the continuity of the listening experience.
[0102] Therefore, it is necessary to determine the contribution of each speaker in the current speaking session in order to dynamically adjust the sound field control parameters. This ensures that the main speaker receives priority optimization in multi-person mixed scenarios, while also considering the listening experience of other participants. In addition, when a speaker switch is detected, the sound field control parameters need to be smoothly transitioned to avoid abrupt noise and ensure that the clarity and consistency of sound quality are maintained in complex interactive environments.
[0103] Figure 4 This is a second exemplary flowchart of a sound field control method according to some embodiments of this specification. Figure 4 As shown, process 400 includes the following steps. In some embodiments, process 400 may be executed by a generation control system (such as sound field control system 200) or a processor (such as processor 120). The following description uses a processor to execute the sound field control method.
[0104] Step 410: Mark the sound-producing entity to generate a sound-producing entity marker.
[0105] A voice entity identifier refers to a unique digital identity assigned to different voice entities.
[0106] In some embodiments, the processor can utilize the Independent Component Analysis (ICA) algorithm to deconstruct the audio signal containing multiple sound sources into several independent individual audio tracks. Simultaneously, the processor can also utilize the Time Difference of Arrival (TDOA) principle to determine the spatial location of each sound source, and use this spatial location as the verification basis for the ICA algorithm (i.e., aligning the individual audio tracks with their spatial locations) to improve separation accuracy. Subsequently, a unique sound source marker is assigned to each spatially aligned individual audio track. Here, an individual audio track can be understood as the audio signal corresponding to a specific sound source.
[0107] In some embodiments, the processor can load the speaker markers of different speakers and the corresponding audio features of individual audio tracks into a historical database. The method of determining sub-audio features based on individual audio tracks is similar to the method of determining audio features based on audio signals; see [link to documentation] for details. Figure 3 The relevant description in step 320.
[0108] In some embodiments, the processor can analyze the audio features corresponding to each individual audio track in real time. By matching a historical database, when an audio feature reappears, it is associated with an existing speaker tag to achieve identity locking for non-continuous speech. When an audio feature does not match the historical database, a speaker tag is assigned to the individual audio track corresponding to that audio feature, and the audio feature of the individual audio track and its corresponding speaker tag are loaded into the historical database.
[0109] Step 420: Determine the energy percentage of the sound-producing element in the current audio signal.
[0110] Energy ratio refers to the ratio of the sub-signal power of a certain sound source to the total signal power.
[0111] In some embodiments, the processor can calculate the sum of the squares of the amplitude values of the audio signal at each moment within a preset time period as the total signal power; calculate the sum of the squares of the amplitude values of the individual audio tracks corresponding to the sound-emitting subject at each moment within the preset time period as the sub-signal power of the sound-emitting subject; and calculate the ratio of the sub-signal power to the total signal power as the energy proportion of the sound-emitting subject.
[0112] The amplitude value refers to the digital sample value after the audio signal is converted into a digital signal, which can reflect the magnitude of the instantaneous voltage (or sound pressure) of the audio signal at a certain moment. The preset time period is determined based on actual needs.
[0113] Step 430: Based on the energy ratio, adjust the sound field control parameters corresponding to different audio frequency bands.
[0114] An audio frequency band refers to a frequency range that divides sound according to its frequency distribution. For example, audio frequency bands include the 80Hz to 300Hz band and the 3kHz to 8kHz band.
[0115] In some embodiments, the sound field control parameters are related to the energy proportion of each sound-producing entity. As an example only, the processor can adjust the sound field control parameters corresponding to a certain audio frequency band using the following formula (1): (1).
[0116] in, These represent the adjusted sound field control parameters for the audio signal within a specific audio frequency range. This represents the adjustment weight corresponding to the reference sound field control parameter of the i-th sound-emitting entity within the sound frequency band. This represents the reference sound field control parameters for the i-th sound-emitting entity within the sound frequency band. This refers to the total number of sound-producing entities contained in the audio signal within that audio frequency range.
[0117] In some embodiments, It is positively correlated with the energy proportion of the i-th sound-producing entity to ensure that the sound-producing entity with stronger energy in the audio signal receives a greater degree of optimization compensation.
[0118] In some embodiments, the processor can construct a frequency response vector based on the audio features of the individual audio tracks of the i-th sound-producing entity, and then retrieve and determine the frequency response vector from a preset parameter library. More information about the preset parameter library. Figure 3 The relevant description in step 330.
[0119] In some embodiments, the processor can divide the audio features of each individual audio track into identity audio features and emotional audio features; determine the speech stability of each individual audio track based on the identity audio features and emotional audio features of each individual audio track of the audio signal; and adjust the sound field control parameters based on the speech stability of each individual audio track.
[0120] Identity audio features refer to audio features that reflect the identity of a speaker, determined by the physiological structure of the speaker (such as the vocal tract, pharynx, oral cavity, etc.). In some embodiments, the processor can extract formant structure, fundamental frequency, MFCC, etc. from the audio features of a single audio track, thereby constructing the identity audio features of the speaker corresponding to that single audio track.
[0121] Emotional audio features refer to dynamic audio characteristics that reflect the psychological state or vocal intention of the speaker. For example, emotional audio features include the fundamental frequency variation of a single audio track, the rising slope of the energy envelope, and the average speech rate of the speaker.
[0122] In some embodiments, the processor can determine the fundamental frequency variation and the rising slope of the energy envelope based on the audio characteristics of a single audio track; the processor can identify non-silent speech segments through Voice Activity Detection (VAD), calculate the number of beats in the non-silent speech segments using a beat detection algorithm, and use the ratio of the number of beats to the total duration of the non-silent speech segments as the average speech rate of the speaker; and construct emotional audio features through fundamental frequency variation, the rising slope of the energy envelope, and duration characteristics.
[0123] Using the methods described above, the processor can extract the identity audio features and emotional audio features of each individual audio track in the audio signal.
[0124] Speech stability refers to an indicator that reflects the continuity and regularity of audio signals in the time and frequency domains.
[0125] In some embodiments, the processor matches identity audio features with a human voice database and emotion features with an emotion database; in response to a failure to match identity audio features with a human voice database and a failure to match emotion features with an emotion database, it determines that the identity audio features and the individual audio track are not human voices, and assigns a lower speech stability (e.g., 0) to the individual track.
[0126] The voice database and emotion database were pre-constructed by technical personnel. The voice database contains multiple identity audio features derived from the analysis of audio signals from people of different ages and genders, while the emotion database contains multiple emotion feature data derived from the analysis of voices from people of different ages and genders under different emotional states.
[0127] In some embodiments, in response to a successful match between identity audio features and a human voice database and / or a successful match between emotion features and an emotion database, the processor determines that the individual audio track is a human voice and determines speech stability based on the harmonic-to-noise ratio (HNR) and fundamental frequency stability of the audio signal.
[0128] Harmonic noise ratio (HNR) refers to the ratio of the energy of harmonic components to the energy of noise components in an audio signal. In some embodiments, the processor can segment the audio signal of the separated sound source (i.e., a single audio track) into multiple speech frames; for each speech frame, its autocorrelation function is calculated; within the fundamental frequency range, the maximum peak value of the autocorrelation function is found, the delay time corresponding to the peak value is taken as the fundamental frequency period, and the autocorrelation value corresponding to the peak value is taken as the energy of the harmonic components (i.e., harmonic energy); the autocorrelation value at zero delay is taken as the total energy of the frame signal, and the difference between the total energy and the harmonic energy is taken as the energy of the noise components (i.e., noise energy); based on the ratio of harmonic energy to noise energy, the HNR of the speech frame is determined.
[0129] Fundamental frequency stability reflects the degree of fluctuation of the fundamental frequency over time. In some embodiments, as an example only, the processor can extract the fundamental frequency of each speech frame, calculate the fundamental frequency change value between adjacent speech frames (i.e., the difference or relative rate of change between the fundamental frequency of the previous speech frame and the fundamental frequency of the next speech frame), and calculate the fundamental frequency stability using the following formula (2): (2).
[0130] in, Indicates fundamental frequency stability (0≤ ≤1); This represents the fundamental frequency variation between adjacent speech frames; This indicates the threshold for the transition. In some embodiments, the threshold is set based on actual needs.
[0131] In some embodiments, for a single audio track containing multiple speech frames, the processor can normalize the harmonic noise ratio and fundamental frequency stability of each speech frame, sum them, and use the sum as the speech stability of that speech frame; and, the processor can use the statistical value (such as the median, weighted average, or average) of the speech stability of multiple speech frames as the speech stability of the single audio track. The weights of the weighted average are determined empirically.
[0132] The first and last speech frames are not included in the calculation of the harmonic noise ratio, fundamental frequency stability, and speech stability, or the harmonic noise ratio, fundamental frequency stability, and speech stability of the first and last speech frames are taken as default values.
[0133] In some embodiments, the adjustment weights corresponding to the reference sound field control parameters of the i-th sound-emitting entity The is positively correlated with the speech stability of the single audio track corresponding to the i-th vocal subject.
[0134] As an example only, the processor can be updated using the following formula (3). : (3).
[0135] in, This represents the adjusted weights (i.e., the updated weights) corresponding to the reference sound field control parameters of the i-th sound-producing entity within a certain sound frequency range. ); This represents the speech stability of the single audio track corresponding to the i-th speaker within the audio segment. (Regarding...) See the description in formula (1).
[0136] In some embodiments, the processor can process each sound-producing entity. After normalization, it is replaced in formula (1) This allows for the adjustment of sound field control parameters.
[0137] In some embodiments of this specification, human voice has a distinct harmonic structure and high speech stability; while random noise (such as object impacts), although having a large energy peak, exhibits a discrete and irregular jump-like spectral distribution, resulting in low speech stability. If such high-energy noise is not effectively suppressed, the processor may misidentify it as a valid sound source, thereby triggering incorrect adjustments to the sound field control parameters, causing abnormal frequency response jumps in the output audio, and severely affecting the coherence of the listening experience and the stability of the sound quality.
[0138] Therefore, by determining speech stability based on identity and emotional audio characteristics, and then adjusting the sound field control parameters, the system can accurately identify non-human voice interference and weaken the gain of its corresponding frequency bands, effectively avoiding misjudgment parameter jumps caused by noise interference. Simultaneously, the system can also provide directional compensation for weak but stable human voices. This ensures that human voices at different distances and volumes achieve balanced and clear sound quality, significantly improving the auditory coherence of the audio output.
[0139] Step 440: When a switch of the sound-generating subject mark is detected, a transition parameter within a preset time window is generated, and based on the transition parameter, the sound field control parameter of the corresponding audio frequency segment is smoothly transitioned.
[0140] In some embodiments, the processor can calculate the energy percentage of each sound-producing entity in the audio signal in real time, determine the sound-producing entity with the highest energy percentage at the current moment as the target sound-producing entity, and determine that the sound-producing entity marker has switched when the target sound-producing entity changes between two adjacent moments.
[0141] A preset time window refers to a pre-defined time interval used to implement a smooth transition. In some embodiments, the preset time window is positively correlated with the absolute value of the difference in the energy ratio of two target sound-emitting entities at two adjacent moments.
[0142] As an example only, the processor calculates the preset time window using the following formula (4): (4).
[0143] in, Indicates the preset time window; Indicates the reference time window; This represents the dynamic time window. In some embodiments, the baseline time window is determined empirically, and the dynamic time window is the normalized value of the absolute value of the difference in the energy proportions of the two target sound-emitting entities.
[0144] Transition parameters refer to intermediate parameters used to smoothly transition the sound field control parameters between two sound-generating entities when the markings of the sound-generating entities change.
[0145] Smooth transition processing refers to the process of smoothly and gradually changing the sound field control parameters of two sound-producing entities to eliminate auditory abruptness.
[0146] In some embodiments, when the sound source marker is switched, the processor can generate transition parameters between the sound field control parameters of the previous target sound source and the next target sound source within a preset time window using an interpolation algorithm (such as linear interpolation, exponential interpolation, cosine interpolation, etc.), and control the filter to adjust the sound field control parameters of the frequency band corresponding to the preset time window based on the transition parameters, so as to achieve a smooth transition of the sound field control parameters of the frequency band.
[0147] In some embodiments of this specification, by separating the main sound source and dynamically identifying the contribution of each sound source in the sound field, a balanced curve that takes into account the needs of multiple parties can be calculated. This ensures that in multi-source aliasing scenarios, the speaker's voice quality is highlighted while also considering the listening experience of other participants. Furthermore, implementing smooth transition processing within a time window when switching speakers avoids abrupt noise caused by alternating voices, enabling the system to maintain excellent auditory coherence and sound quality stability in dynamic and complex interactive environments.
[0148] It is worth noting that since the speaker may appear in various emotional states during the interaction, fixing the sound field control parameters solely based on identity information cannot adapt to the frequency response changes of the same speaker in different emotional states, easily leading to auditory distortion. Furthermore, if it is impossible to accurately identify whether the current speaker is the same as the previous speaker, the parameters need to be re-initialized each time they speak, resulting in inconsistent auditory perception of the same person and affecting the continuity of the interaction.
[0149] Therefore, it is necessary to verify the consistency of the speaker's identity information to ensure that the sound field control parameters of the same speaker can be used continuously at different times to maintain the continuity of interaction; and further utilize emotional information to perform secondary gain compensation on the sound field control parameters to dynamically offset the frequency response distortion caused by emotional changes, thereby jointly ensuring the naturalness of sound quality and the stability of listening experience in both identity stability and emotional adaptation dimensions.
[0150] Figure 5 This is the third exemplary flowchart of a sound field control method according to some embodiments of this specification. Figure 5 As shown, process 500 includes the following steps. In some embodiments, process 500 may be executed by a sound field control system (such as sound field control system 200) or a processor (such as processor 120). The following description uses a processor to execute the sound field control method.
[0151] Step 510: Divide the audio features into identity audio features and emotional audio features.
[0152] For more information on identity audio features, emotional audio features, and how to determine identity audio features and emotional audio features, please see [link to relevant documentation]. Figure 4 The relevant description in step 430.
[0153] Step 520: Based on the identity audio features, identify the speaker's identifier and perform consistency verification of the speaker.
[0154] Consistency verification refers to the process of verifying whether the target speaker in a continuous speech refers to the same speaker. In some embodiments, the processor can calculate the energy proportion of each speaker in the audio signal in real time, determine the speaker with the highest energy proportion at the current moment as the target speaker, and if the target speaker at the current moment is the same speaker as the target speaker at the previous moment, the consistency verification passes.
[0155] In some embodiments, if the target sound source at the current moment is not the same as the target sound source at the previous moment, the processor can determine the corresponding sound field control parameters by means of the audio characteristics of the new target sound source. See details below. Figure 3 The corresponding content.
[0156] Step 530: Based on the emotional audio features of the speaker that have passed the consistency check, determine the emotional fluctuation information of the speaker.
[0157] Emotional fluctuation information characterizes the deviation (such as deviation tendency and deviation magnitude) of the speaker's current emotion relative to their emotion in a preset emotional state. The preset emotional state may include an emotional state labeled as "calm," etc. In some embodiments, emotional fluctuation information may include the speaker's current emotional label and fluctuation characteristics.
[0158] Emotional markers are used to reflect the emotional state of the speaker at the current moment. For example, the speaker may be in a low, excited, or calm state.
[0159] In some embodiments, the processor compares the parameters of the emotional audio features of the speaker at the current moment with the baseline emotional audio features to determine the offsets of multiple parameters of the same dimension; constructs a feature vector based on the emotional audio features and the corresponding multiple offsets; queries a first preset emotion table; and uses the sample emotion markers that meet the second preset requirements as the emotion markers of the speaker; determines fluctuation features based on the multiple offsets; and uses the emotion markers and fluctuation features as emotional fluctuation information.
[0160] The baseline emotional audio feature refers to the emotional audio feature of the speaker in a preset emotional state. In some embodiments, after system initialization or power-on, the processor can instruct the speaker to record a speech sample in a preset emotional state, analyze the corresponding emotional audio feature of the speech sample, and use the emotional audio feature as the baseline emotional audio feature of the speaker.
[0161] "Same dimension" can be understood as the parameters used for comparison in emotional audio features and baseline emotional audio features having the same meaning. For example, the fundamental frequency in emotional audio features and the fundamental frequency in baseline emotional audio features are parameters of the same dimension.
[0162] Offset refers to the difference between parameters of the same dimension in an emotional audio feature and a baseline emotional audio feature. Offset can be positive, negative, or 0.
[0163] In some embodiments, the first preset sentiment table includes multiple sample feature vectors and their corresponding sample sentiment tags.
[0164] In some embodiments, the first preset emotion table can be constructed based on historical data. For example, the processor can collect baseline speech samples from test subjects of different ages, genders, timbres, and vocal habits under preset emotional states, and then analyze them to obtain the baseline emotional audio features of the samples; then it can collect emotional speech samples from test subjects under other emotional states (i.e., sample emotional tags), analyze the emotional audio features corresponding to each emotional speech sample (i.e., sample emotional audio features), determine the offsets of the baseline emotional audio features and the sample emotional audio features in multiple parameters of the same dimension (i.e., sample offsets), and construct sample feature vectors based on the sample emotional audio features and the corresponding multiple sample offsets; and construct the first preset emotion table through multiple sample feature vectors and the corresponding sample emotional tags. In this case, the multiple sample offsets under the preset emotional state are all 0.
[0165] In some embodiments, the second preset requirement may include one of the following: the shortest vector distance between the feature vector and the sample feature vector, the maximum vector similarity, etc.
[0166] Fluctuation characteristics are quantitative indicators used to measure the intensity of the fluctuation of a speaker's emotional state at the current moment relative to its pre-set emotional state. In some embodiments, the processor can normalize multiple offsets, calculate statistical values (such as arithmetic mean or average) of the normalized offsets, and use these statistical values as fluctuation characteristics.
[0167] Step 540: Based on emotional fluctuation information, adjust the sound field control parameters for different audio frequency bands.
[0168] In some embodiments, the sound field control parameters are positively correlated with the wave characteristics. As an example only, the processor can adjust the sound field control parameters of the sound-generating entity in a certain acoustic frequency range based on the following formula (5): (5).
[0169] in, These are the adjusted sound field control parameters for the sound-producing subject in a certain audio frequency range; ΔB represents the sound field control parameters of the sound-generating entity before adjustment in a certain audio frequency range; ΔB represents the wave characteristics of the sound-generating entity.
[0170] In some embodiments, the sound field control parameters of the sound-emitting subject before adjustment in a certain audio frequency range. This can be determined by querying historical data, or by constructing a frequency response vector based on the audio characteristics of the individual audio tracks of the sound source, and then searching for it in a preset parameter library. More information about the preset parameter library can be found in the documentation. Figure 3 The relevant description in step 330.
[0171] In some embodiments, the processor can determine the formula (5) above. As in formula (1) The updated value is substituted into formula (1) to adjust the sound field control parameters for different audio frequency bands.
[0172] In some embodiments, the processor can also determine the energy proportion of different vocal subjects in the audio signal; and adjust the sound field control parameters based on the energy proportion and emotional fluctuation information. For further explanation of energy proportion and how to determine it, see [link to relevant documentation]. Figure 4 And its related descriptions.
[0173] In some embodiments, the sound field control parameters are positively correlated with the energy proportion and fluctuation characteristics of the sound-generating entity. As an example only, the processor can adjust the sound field control parameters of the sound-generating entity within a certain audio frequency range using the following formula (6): (6).
[0174] in, This represents the adjusted sound field control parameters for the i-th sound-generating entity within a certain audio frequency range; ΔB represents the wave characteristics of this sound-generating entity. (Regarding...) For further explanation, please refer to formula (1) and its related description. For further explanation, see formula (5) and its related description.
[0175] In some embodiments, the processor can determine the formula (6) above. As in formula (1) The updated value is substituted into formula (1) to adjust the sound field control parameters for different audio frequency bands.
[0176] In some embodiments of this specification, the sound field control parameters are adjusted by combining energy ratio and emotional fluctuation information to achieve dual compensation of physical identity and emotional state. This can effectively offset the frequency response distortion caused by emotional fluctuations, avoid sound quality deterioration under extreme emotions, and ensure that the audio output by the microphone in complex interactive contexts always maintains clear spatial layers and a natural and stable listening experience.
[0177] In some embodiments of this specification, consistency verification is performed using identity audio features to ensure the continuity of sound field control during interaction; secondary gain compensation is performed on the sound field control parameters using emotional fluctuation information, which can dynamically offset the auditory distortion caused by emotional changes and maintain the naturalness of sound quality and auditory stability.
[0178] It should be noted that the above descriptions of processes 300, 400, and 500 are for illustrative purposes only and do not limit the scope of this specification. Those skilled in the art can make various modifications and changes to processes 300, 400, and 500 under the guidance of this specification. However, these modifications and changes remain within the scope of this specification.
[0179] It is worth noting that if the processor cannot initiate sound field parameter adjustment in the early stages of emotional state transition before the auditory abnormality becomes obvious, it will only respond passively after the emotional change has caused obvious auditory distortion (such as high-frequency harshness, low-frequency dullness, and trembling sound). At this time, the listener has already perceived the deterioration of sound quality, and the compensation effect is greatly reduced. Moreover, the lag in parameter adjustment will also lead to sudden changes or abrupt changes in sound, destroying the naturalness and coherence of audio output, and failing to achieve active and smooth sound field optimization.
[0180] Therefore, it is necessary to accurately identify emotional fluctuations and adjust sound field control parameters in the early stages of emotional state transition (i.e., before auditory abnormalities become obvious) to achieve proactive compensation for timbre distortion under different emotional states, avoid sound quality distortion being perceived by the audience, and ensure the naturalness and consistency of audio output.
[0181] Figure 6 This is an exemplary schematic diagram illustrating the generation of emotional fluctuation information according to some embodiments of this specification.
[0182] In some embodiments, such as Figure 6 As shown, the processor can extract the duration feature 621, energy envelope feature 622, and fundamental frequency perturbation value 623 of the sound-producing subject in the audio signal 610; determine the first fluctuation amplitude 631 of the duration feature 621, the second fluctuation amplitude 632 of the energy envelope feature 622, and the third fluctuation amplitude 633 of the fundamental frequency perturbation value 623; and generate emotional fluctuation information 640 based on the first fluctuation amplitude 631, the second fluctuation amplitude 632, and the third fluctuation amplitude 633.
[0183] Duration features refer to characteristics that reflect the length or speed of a speaker's pronunciation. For example, duration features can at least include the speaker's average speech rate.
[0184] In some embodiments, the processor can identify the start and end times of the speech signal (i.e., determine the non-silent speech segment) using VAD technology; calculate the number of beats in the non-silent speech segment using a beat detection algorithm (e.g., based on energy envelope peak detection or a deep learning model); and use the ratio of the number of beats to the total duration of the non-silent speech segment as the average speech rate of the speaker (i.e., the duration feature).
[0185] Energy envelope features refer to the contour features that reflect the amplitude change of an audio signal over time, and are used to reflect the energy distribution of the audio signal. In some embodiments, energy envelope features may include at least the rate at which energy rises from zero to a peak value (i.e., the slope of the energy envelope). In some embodiments, a processor can determine the energy envelope features in non-silent speech segments by analyzing the energy envelope in the audio features of the speaker.
[0186] Jitter value refers to the frequency fluctuation of an audio signal. In some embodiments, the processor can extract the fundamental frequency sequence (i.e., the sequence of fundamental frequency distribution over time) within a non-silent speech segment, calculate the difference between multiple adjacent fundamental frequencies within the non-silent speech segment, calculate the average of the multiple differences, and use the average value as the fundamental frequency jitter value within the non-silent speech segment.
[0187] The first fluctuation amplitude refers to the degree of deviation of the duration feature from its reference duration feature. In some embodiments, the processor can use the difference between the duration feature and the reference duration feature as the first fluctuation amplitude. The reference fluctuation amplitude can be obtained based on the analysis of the audio signal of the speaker in a preset emotional state.
[0188] The second fluctuation amplitude refers to the degree of deviation of the rising slope of the energy envelope from the reference rising slope. In some embodiments, the processor can use the difference between the rising slope of the energy envelope and the reference rising slope as the second fluctuation amplitude. The reference rising slope can be obtained based on the analysis of the audio signal of the speaker in a preset emotional state.
[0189] The third fluctuation amplitude refers to the degree of deviation of the fundamental frequency perturbation value from the reference fundamental frequency perturbation value. In some embodiments, the processor can use the difference between the fundamental frequency perturbation value and the reference fundamental frequency perturbation value as the third fluctuation amplitude. The reference fundamental frequency perturbation value can be obtained based on the analysis of the audio signal of the speaker under a preset emotional state.
[0190] In some embodiments, emotional fluctuation information may include, for example, the emotional characteristics of the speaker (such as anger, sadness, etc.) and the degree to which they deviate from a normal state.
[0191] In some embodiments, the processor may construct a fluctuation feature vector based on a first fluctuation amplitude, a second fluctuation amplitude, and a third fluctuation amplitude; query a second preset sentiment table based on the fluctuation feature vector to determine the sentiment marker; calculate the statistical values (such as the arithmetic mean or average value) of the normalized first fluctuation amplitude, the second fluctuation amplitude, and the third fluctuation amplitude, and use the statistical values as the comprehensive change amplitude of the non-silent speech segment; and use the sentiment marker and the comprehensive change amplitude as sentiment fluctuation information.
[0192] In some embodiments, the second preset sentiment table includes multiple sample first fluctuation amplitudes, sample second fluctuation amplitudes, and sample third fluctuation amplitudes, as well as their corresponding sample sentiment labels. The construction method of the second preset sentiment table is similar to that of the first preset sentiment table.
[0193] Some embodiments in this specification, by analyzing multi-dimensional physical quantities such as duration characteristics, energy envelope characteristics, and fundamental frequency perturbation values, can capture microscopic acoustic anomalies that are difficult for the human ear to perceive, thereby achieving accurate identification in the early stages of the speaker's emotional state transition. By refining the emotion judgment conditions into fluctuation amplitude and emotional characteristics, it can provide a high-precision decision-making basis for the subsequent adjustment of sound field control parameters, and thus provide accurate compensation for timbre distortion under different emotional states (such as high-frequency spikes when angry, and vibrato distortion when fearful), significantly improving the accuracy of semantic expression and auditory realism in complex interactive contexts.
[0194] In some embodiments, the processor can determine the gain adjustment operation and gain adjustment magnitude for different audio frequency bands based on emotion tags and emotion fluctuation information.
[0195] Gain adjustment refers to the operation of enhancing or suppressing the energy of a certain audio frequency band in an audio signal. In some embodiments, gain adjustment includes amplification gain (i.e., enhancing the energy of the audio frequency band) and reduction gain (i.e., suppressing the energy of the audio frequency band).
[0196] In some embodiments, the processor may determine the fluctuating frequency bands in the audio signal and their corresponding gain adjustment operations based on emotion tags.
[0197] Among them, the fluctuation frequency band refers to the audio signal frequency band that produces abnormal sound perception due to emotional changes during the process of changing from a preset emotional state to the current emotional marker.
[0198] In some embodiments, the fluctuation frequency bands and gain adjustment operations corresponding to different emotion tags can be determined experimentally. For example, the processor can collect sample audio signals under different emotion tags; for a sample audio signal under a certain emotion tag, multiple users select frequency bands that sound abnormal (such as sharp or blurry sounds), and the processor will use the frequency band selected most often as the fluctuation frequency band under that audio; perform multiple different gain adjustment operations on the fluctuation frequency band, and use the gain adjustment operation selected most often by the users as the gain adjustment operation corresponding to the fluctuation frequency band; thereby establishing a mapping relationship between multiple emotion tags and fluctuation frequency bands and gain adjustment operations.
[0199] As an example, for emotions marked as excitement or anger, the processor can apply reduced gain to the 3kHz-8kHz frequency band (i.e., the fluctuating frequency band) to remove sibilance and sharpness. For emotions marked as sadness or fatigue, the processor can apply amplified gain to the 100Hz-500Hz frequency band (i.e., the fluctuating frequency band) to enhance the warmth and support of the sound.
[0200] Gain adjustment range refers to the amount of adjustment performed during a gain adjustment operation.
[0201] In some embodiments, the processor may determine the gain adjustment magnitude based on the overall change magnitude. As an example only, the processor may determine the gain adjustment magnitude based on the following formula (7): (7).
[0202] in, Indicates the frequency band of fluctuation Gain adjustment range, Indicates the benchmark adjustment range. This represents the adjustment factor. In some embodiments, It can be set based on experience. With fluctuation frequency band The combined change amplitude of the non-silent speech segment is linearly or exponentially correlated.
[0203] In some embodiments, the processor can adjust the sound field control parameters of the fluctuating frequency band based on the gain adjustment operation and gain adjustment amplitude corresponding to the fluctuating frequency band.
[0204] In some embodiments of this specification, the gain adjustment operation and the gain adjustment amplitude are dynamically determined based on emotional fluctuation information to achieve differentiated frequency band compensation. This can effectively eliminate auditory spikes or blurriness caused by emotional fluctuations and provide users with a stable and natural high-fidelity audio experience.
[0205] This specification also provides a sound field control device, including at least one processor and at least one memory; the at least one memory is used to store computer instructions; and the at least one processor is used to execute at least some of the computer instructions to implement the sound field control method described in any of the above embodiments.
[0206] This specification also provides a computer-readable storage medium that stores computer instructions. When a computer reads the computer instructions from the storage medium, the computer executes the sound field control method described in any of the above embodiments.
[0207] The basic concepts have been described above. Obviously, for those skilled in the art, the detailed disclosure above is merely illustrative and does not constitute a limitation of this specification. Although not explicitly stated herein, those skilled in the art may make various modifications, improvements, and corrections to this specification. Such modifications, improvements, and corrections are suggested in this specification and therefore remain within the spirit and scope of the exemplary embodiments described herein.
[0208] Furthermore, unless expressly stated in the claims, the order of elements and sequences, the use of numbers and letters, or other names in this specification are not intended to limit the order of the processes and methods described herein. Although various examples have been discussed in the foregoing disclosure of some embodiments of the invention that are currently considered useful, it should be understood that such details are for illustrative purposes only, and the appended claims are not limited to the disclosed embodiments; rather, the claims are intended to cover all modifications and equivalent combinations that conform to the spirit and scope of the embodiments described herein. For example, while the system components described above can be implemented using hardware devices, they can also be implemented solely using software solutions, such as installing the described system on an existing server or mobile device.
[0209] Similarly, it should be noted that, in order to simplify the description disclosed herein and thus aid in the understanding of one or more embodiments of the invention, the foregoing description of embodiments in this specification may sometimes combine multiple features into a single embodiment, drawing, or description thereof. However, this method of disclosure does not imply that the subject matter of this specification requires more features than those mentioned in the claims. In fact, the embodiments contain fewer features than all the features of a single embodiment disclosed above.
[0210] Finally, it should be understood that the embodiments in this specification are merely illustrative of the principles of the embodiments described herein. Other variations may also fall within the scope of this specification. Therefore, alternative configurations of the embodiments in this specification are intended to be illustrative rather than limiting, and should be considered consistent with the teachings of this specification. Accordingly, the embodiments in this specification are not limited to those explicitly described and illustrated herein.
Claims
1. A sound field control method, characterized in that, include: Acquire the audio signal of the speaker; Based on the audio signal, audio features characterizing the vocal attributes of the vocal subject are generated; Based on the aforementioned audio features, sound field control parameters are generated; as well as Based on the sound field control parameters, the audio device is controlled to output audio.
2. The method according to claim 1, characterized in that, The method further includes: The sound-producing entity is marked to generate a sound-producing entity mark; Determine the energy percentage of the sound-producing entity in the current audio signal; Based on the energy ratio, adjust the sound field control parameters corresponding to different audio frequency bands; and When a switch of the sound-generating entity marker is detected, transition parameters within a preset time window are generated, and based on the transition parameters, the sound field control parameters of the corresponding audio frequency band are subjected to smooth transition processing.
3. The method according to claim 1, characterized in that, Also includes: The audio features are divided into identity audio features and emotional audio features; Based on the aforementioned identity audio features, the speaker's identifier is identified, and consistency verification of the speaker is performed. Based on the emotional audio features of the vocal subject that have passed the consistency check, the emotional fluctuation information of the vocal subject is determined. as well as Based on the emotional fluctuation information, the sound field control parameters for different audio frequency bands are adjusted.
4. The method according to claim 3, characterized in that, The step of determining the emotional fluctuation information of the speaker based on the emotional audio features of the speaker that have passed the consistency check includes: Extract the duration characteristics, energy envelope characteristics, and fundamental frequency perturbation value of the sound-producing entity from the audio signal; Determine the first fluctuation amplitude of the duration feature, the second fluctuation amplitude of the energy envelope feature, and the third fluctuation amplitude of the fundamental frequency perturbation value; and The emotional fluctuation information is generated based on the first fluctuation amplitude, the second fluctuation amplitude, and the third fluctuation amplitude.
5. A sound field control system, characterized in that, include: The acquisition module is configured to acquire the audio signal of the sound-producing entity; The first generation module is configured to generate audio features that characterize the vocal attributes of the vocal subject based on the audio signal. The second generation module is configured to generate sound field control parameters based on the audio features; as well as The control module is configured to control the audio device to output audio based on the sound field control parameters.
6. The system according to claim 5, characterized in that, Also includes: The tag generation module is configured to tag the sound-emitting subject and generate a sound-emitting subject tag; An energy determination module is configured to determine the energy percentage of the sound-producing entity in the current audio signal; The first adjustment module is configured to adjust the sound field control parameters corresponding to different audio frequency bands based on the energy ratio. as well as The transition processing module is configured to generate transition parameters within a preset time window when a switch of the sound-generating subject mark is detected, and to perform smooth transition processing on the sound field control parameters of the corresponding sound frequency band based on the transition parameters.
7. The system according to claim 5, characterized in that, Also includes: The segmentation module is configured to segment the audio features into identity audio features and emotion audio features; The verification module is configured to identify the voice subject marker of the voice subject based on the identity audio features, and to perform consistency verification of the voice subject; The emotion determination module is configured to determine the emotional fluctuation information of the vocal subject based on the emotional audio features of the vocal subject that have passed the consistency check. as well as The second adjustment module is configured to adjust the sound field control parameters for different audio frequency bands based on the emotional fluctuation information.
8. The system according to claim 7, characterized in that, The emotion determination module is further configured as follows: Extract the duration characteristics, energy envelope characteristics, and fundamental frequency perturbation value of the sound-producing entity from the audio signal; Determine the first fluctuation amplitude of the duration feature, the second fluctuation amplitude of the energy envelope feature, and the third fluctuation amplitude of the fundamental frequency perturbation value; as well as The emotional fluctuation information is generated based on the first fluctuation amplitude, the second fluctuation amplitude, and the third fluctuation amplitude.
9. A sound field control device, characterized in that, The device includes at least one processor and at least one memory; The at least one memory is used to store computer instructions; and The at least one processor is configured to execute at least a portion of the computer instructions to implement the method as described in any one of claims 1 to 4.
10. A computer-readable storage medium, characterized in that, The storage medium stores computer instructions. When the computer reads the computer instructions from the storage medium, the computer executes the method as described in any one of claims 1 to 4.