Volume adjustment method and apparatus for online multi-party call, electronic device, and storage medium

WO2026174569A1PCT designated stage Publication Date: 2026-08-27BOE TECHNOLOGY GROUP CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2025/078675
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2025-02-23
Publication Date
2026-08-27

Smart Images

  • Figure CN2025078675_27082026_PF_FP_ABST
    Figure CN2025078675_27082026_PF_FP_ABST
Patent Text Reader

Abstract

A volume adjustment method and apparatus for an online multi-party call, an electronic device, and a storage medium. The volume adjustment method for an online multi-party call comprises: acquiring audio data of an online multi-party call in real time to obtain current audio data (S110), wherein the current audio data comprises audio data of a plurality of speakers; separating the audio data of the plurality of speakers to obtain an audio feature of each of the plurality of speakers (S120); on the basis of the audio feature of each speaker, performing identity authentication on the plurality of speakers to obtain audio data corresponding to the identity of each speaker (S130); and on the basis of a user instruction and the audio data corresponding to the identity of each speaker, selecting a target speaker from among the plurality of speakers, and adjusting a volume value of the target speaker (S140). The volume adjustment method for an online multi-party call can prevent interference caused by speaking of multiple speakers, improve the experience of online multi-party calls, reduce communication costs, and improve the efficiency of online multi-party calls.
Need to check novelty before this filing date? Find Prior Art

Description

Online multi-person call volume adjustment methods and devices, electronic devices and storage media Technical Field

[0001] Embodiments of this disclosure relate to a method and apparatus for adjusting volume in online multi-person calls, an electronic device, and a storage medium. Background Technology

[0002] As the world enters the information age, the rapid development of computer, communication, and multimedia technologies has led to a shift in people's needs for understanding things and exchanging information. This has evolved from paper, pen, books, and voice to more accurate, faster, and richer expressions through sound, light, and electrical signals. Driven by this demand, multimedia computer technology has combined with communication technology to gradually develop into a new emerging technology: multimedia communication technology. Based on multimedia communication technology, online multi-person calls (e.g., multi-person remote conferences, multi-person online group chats) can be initiated anytime, anywhere, conveniently and quickly, without the need for face-to-face interaction. Summary of the Invention

[0003] This disclosure provides at least one embodiment of a method for adjusting the volume of an online multi-person call, comprising: acquiring audio data of the online multi-person call in real time to obtain current audio data, wherein the current audio data includes audio data of multiple speakers; separating the audio data of the multiple speakers to obtain audio features of each speaker; verifying the identities of the multiple speakers based on the audio features of each speaker to obtain audio data corresponding to the identity of each speaker; and selecting a target speaker from the multiple speakers based on a user instruction and the audio data corresponding to the identity of each speaker, and adjusting the volume value of the target speaker.

[0004] In the online multi-person call volume adjustment method provided in at least one embodiment of this disclosure, the step of separating the audio data of the multiple speakers to obtain the audio features of each speaker includes: extracting features from the audio data of the multiple speakers to obtain the audio features of the multiple speakers; inputting the audio features of the multiple speakers into a speaker separation model, and outputting the audio features of each speaker after processing.

[0005] In the online multi-person call volume adjustment method provided in at least one embodiment of this disclosure, the step of inputting the audio features of the multiple speakers into a speaker separation model and outputting the audio features of each speaker after processing includes: encoding the audio features of the multiple speakers into a high-dimensional feature sequence; capturing local features in the high-dimensional feature sequence, wherein the local features include dependencies corresponding to each speaker; performing a separation operation on the high-dimensional feature sequence based on the local features to obtain a feature sequence corresponding to each speaker; and decoding the feature sequence corresponding to each speaker into the audio features of each speaker.

[0006] In the online multi-person call volume adjustment method provided in at least one embodiment of this disclosure, the step of verifying the identities of the multiple speakers based on the audio features of each speaker to obtain audio data corresponding to the identity of each speaker includes: inputting the audio features of each speaker into a speaker identity verification model; using the speaker identity verification model to extract the identity features of each speaker from the audio features of each speaker; verifying the identity of the corresponding speaker based on the identity features of each speaker, and outputting the audio data corresponding to the identity of each speaker.

[0007] In the online multi-person call volume adjustment method provided in at least one embodiment of this disclosure, the step of extracting the identity features of each speaker from the audio features of each speaker using the speaker identity verification model includes: preprocessing the audio features of a first speaker among the plurality of speakers to obtain preprocessed audio features of the first speaker; extracting context information from the preprocessed audio features of the first speaker to obtain a first feature map containing the context information; performing masking processing on the first feature map to remove noise from the first feature map and retain or enhance the key information of the first speaker; and extracting the identity features of the first speaker based on the masked first feature map.

[0008] In the online multi-person call volume adjustment method provided in at least one embodiment of this disclosure, the step of confirming the identity of the corresponding speaker based on the identity features of each speaker includes: performing cluster analysis on the identity features of each speaker based on information in the speaker identity database to confirm the speaker identity.

[0009] In the online multi-person call volume adjustment method provided in at least one embodiment of this disclosure, the step of performing cluster analysis on the identity features of each speaker based on information in the speaker identity database to confirm the speaker's identity includes: comparing the identity features of each speaker with the information in the speaker identity database to obtain a similarity score of the identity features of each speaker relative to the speaker identity database; and confirming that the first speaker is the first identity if the similarity score of the identity features of the first speaker among the plurality of speakers relative to the first identity in the speaker identity database is higher than a preset value.

[0010] In the online multi-person call volume adjustment method provided in at least one embodiment of this disclosure, the user instruction includes a speaker selection instruction, the audio data corresponding to the identity of each speaker includes the identity information of the corresponding speaker, and the step of selecting a target speaker from the plurality of speakers based on the user instruction and the audio data corresponding to the identity of each speaker includes: selecting a target speaker from the plurality of speakers based on the speaker selection instruction and the identity information of each speaker.

[0011] In at least one embodiment of the online multi-person call volume adjustment method provided in this disclosure, the user instruction includes a volume adjustment instruction, which is used to determine the position of the volume slider. The audio data corresponding to the identity of each speaker includes the volume information of the corresponding speaker. The step of adjusting the volume value of the target speaker based on the user instruction and the audio data corresponding to the identity of each speaker includes: adjusting the volume slider to a first position based on the volume adjustment instruction; and obtaining the current volume value of the target speaker based on the volume information of the target speaker and the first position.

[0012] In the online multi-person call volume adjustment method provided in at least one embodiment of this disclosure, the volume slider corresponds to a volume adjustment range, the volume adjustment range includes multiple volume preset values, the first position of the volume slider corresponds to a first volume preset value among the multiple volume preset values, and the step of obtaining the current volume value of the target speaker based on the volume information of the target speaker and the first position includes: determining the first volume preset value based on the first position; and obtaining the current volume value of the target speaker based on the volume information of the target speaker and the first volume preset value.

[0013] In at least one embodiment of the online multi-person call volume adjustment method provided in this disclosure, the user instruction includes a speaker selection instruction and a volume adjustment instruction. The volume adjustment instruction is used to determine the position of the volume slider. The audio data corresponding to the identity of each speaker includes the identity information of the corresponding speaker and the volume information of the corresponding speaker. The step of selecting a target speaker from the plurality of speakers based on the user instruction and the audio data corresponding to the identity of each speaker, and adjusting the volume value of the target speaker, includes: displaying the identity information of the plurality of speakers on the screen; selecting the target speaker based on the speaker selection instruction and the identity information of the plurality of speakers displayed on the screen, wherein the speaker selection instruction is generated by the user clicking on the identity information of the target speaker on the screen; adjusting the volume slider to a first position based on the volume adjustment instruction, wherein the first position corresponds to a first volume preset value in the volume adjustment range of the volume slider; and obtaining the current volume value of the target speaker based on the volume information of the target speaker and the first volume preset value.

[0014] In the online multi-person call volume adjustment method provided in at least one embodiment of this disclosure, the audio data corresponding to the identity of each speaker includes the volume information of the corresponding speaker, and the step of adjusting the volume value of the target speaker based on the user instruction and the audio data corresponding to the identity of each speaker includes: automatically adjusting the volume value of the target speaker to the target volume value based on the volume information of the target speaker.

[0015] In at least one embodiment of the online multi-person call volume adjustment method provided in this disclosure, the step of automatically adjusting the volume value of the target speaker to a target volume value based on the volume information of the target speaker includes: performing a first processing on the volume information of the target speaker by an amplifier, sampling the first-processed volume information of the target speaker to obtain a plurality of first sampling signals; performing a second processing on the plurality of first sampling signals to obtain an average power value of the plurality of first sampling signals; comparing the average power value of the plurality of first sampling signals with the target volume value to obtain a first comparison result; and automatically adjusting the gain of the amplifier based on the first comparison result to automatically adjust the volume value of the target speaker to the target volume value.

[0016] In the online multi-person call volume adjustment method provided in at least one embodiment of this disclosure, before acquiring the audio data of the online multi-person call in real time, the online multi-person call volume adjustment method further includes: acquiring a target volume value, wherein the target volume value is a volume value that the user is accustomed to or a volume value that the user customizes.

[0017] In the online multi-person call volume adjustment method provided in at least one embodiment of this disclosure, in response to the target volume value being the volume value that the user is accustomed to, the step of obtaining the target volume value includes: obtaining and storing multiple volume values ​​set by the user locally, and selecting the first volume value that the user sets the most times from the multiple volume values ​​as the target volume value.

[0018] In at least one embodiment of the online multi-person call volume adjustment method provided in this disclosure, after acquiring the audio data of the online multi-person call in real time, the online multi-person call volume adjustment method further includes: extracting the volume information of the current audio data; and automatically adjusting the current volume value of the online multi-person call to the target volume value based on the volume information of the current audio data.

[0019] In at least one embodiment of the online multi-person call volume adjustment method provided in this disclosure, the step of automatically adjusting the current volume value of the online multi-person call to the target volume value based on the volume information of the current audio data includes: performing a first processing on the volume information of the current audio data by an amplifier, sampling the volume information of the first-processed current audio data to obtain a plurality of second sampling signals; performing a second processing on the plurality of second sampling signals to obtain an average power value of the plurality of second sampling signals; comparing the average power value of the plurality of second sampling signals with the target volume value to obtain a second comparison result; and automatically adjusting the gain of the amplifier based on the second comparison result to automatically adjust the current volume value of the online multi-person call to the target volume value.

[0020] At least one embodiment of this disclosure also provides an online multi-person call volume adjustment device, comprising: an acquisition module configured to acquire audio data of the online multi-person call in real time to obtain current audio data, wherein the current audio data includes audio data of multiple speakers; a speaker separation module configured to separate the audio data of the multiple speakers to obtain audio features of each speaker; a speaker identity verification module configured to verify the identity of the multiple speakers based on the audio features of each speaker to obtain audio data corresponding to the identity of each speaker; and a speaker volume adjustment module configured to select a target speaker from the multiple speakers based on a user instruction and the audio data corresponding to the identity of each speaker, and adjust the volume value of the target speaker.

[0021] At least one embodiment of this disclosure also provides an electronic device. The electronic device includes: a processor; and a memory including one or more computer program modules; wherein the one or more computer program modules are stored in the memory and configured to be executed by the processor, the one or more computer program modules being used to implement the online multi-person call volume adjustment method provided in any embodiment of this disclosure.

[0022] At least one embodiment of this disclosure also provides a storage medium storing non-transitory computer-readable instructions that, when executed by a computer, implement the online multi-person call volume adjustment method provided in any embodiment of this disclosure. Attached Figure Description

[0023] To more clearly illustrate the technical solutions of the embodiments of this disclosure, the accompanying drawings of the embodiments will be briefly described below. Obviously, the drawings described below only relate to some embodiments of this disclosure and are not intended to limit this disclosure.

[0024] Figure 1 is an exemplary flowchart of an online multi-person call volume adjustment method provided in at least one embodiment of the present disclosure;

[0025] Figure 2 is a schematic diagram of an example of a speaker separation model provided in at least one embodiment of the present disclosure;

[0026] Figure 3 is a schematic diagram of an example of a speaker identification model provided in at least one embodiment of this disclosure;

[0027] Figure 4A is a schematic diagram illustrating an example of the relationship between the position of the volume slider and the volume level provided in at least one embodiment of this disclosure;

[0028] Figure 4B is a schematic diagram illustrating an example of the logarithmic relationship between the volume slider position and the volume level provided in at least one embodiment of this disclosure;

[0029] Figure 4C is a schematic diagram of an example of an online multi-person call volume adjustment method provided in at least one embodiment of the present disclosure;

[0030] Figure 5 is another exemplary flowchart of an online multi-person call volume adjustment method provided in at least one embodiment of the present disclosure;

[0031] Figure 6 is a schematic block diagram of an online multi-person call volume adjustment device provided in at least one embodiment of the present disclosure;

[0032] Figure 7 is a schematic block diagram of an electronic device provided in at least one embodiment of the present disclosure;

[0033] Figure 8 is a schematic block diagram of another electronic device provided in at least one embodiment of the present disclosure; and

[0034] Figure 9 is a schematic diagram of a storage medium provided in at least one embodiment of the present disclosure. Detailed Implementation

[0035] To make the objectives, technical solutions, and advantages of the embodiments of this disclosure clearer, the technical solutions of the embodiments of this disclosure will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this disclosure. All other embodiments obtained by those skilled in the art based on the described embodiments of this disclosure without creative effort are within the scope of protection of this disclosure.

[0036] Unless otherwise defined, the technical or scientific terms used in this disclosure shall have the ordinary meaning understood by one of ordinary skill in the art to which this disclosure pertains. The terms “first,” “second,” and similar terms used in this disclosure do not indicate any order, quantity, or importance, but are merely used to distinguish different components. Similarly, the terms “an,” “a,” or “the,” and similar terms do not indicate a quantity limitation, but rather indicate the presence of at least one. The terms “including,” “comprising,” or “containing,” and similar terms mean that the element or object preceding the word encompasses the elements or objects listed following the word and their equivalents, without excluding other elements or objects. The terms “connected,” “linked,” or similar terms are not limited to physical or mechanical connections, but can include electrical connections, whether direct or indirect. The terms “upper,” “lower,” “left,” and “right,” etc., are used only to indicate relative positional relationships, and these relative positional relationships may change accordingly when the absolute position of the described objects changes.

[0037] The present disclosure will now be described through several specific embodiments. To keep the following description of the embodiments of the present disclosure clear and concise, detailed descriptions of known functions and components may be omitted. When any component of the embodiments of the present disclosure appears in more than one drawing, the component is represented by the same or similar reference numerals in each drawing.

[0038] In online multi-person calls, issues arise such as some speakers being too loud while others are too soft, or inconsistent microphone volume output. Therefore, volume adjustment is a crucial aspect of online multi-person calls. In one scenario, if the call volume is too low, users cannot discern specific voice information and may even need to press their ear to the speaker and strain to listen; this could be due to the speaker being too far from the microphone or the microphone's pickup volume being too low. In another scenario, if the call volume is too high, users suffer from ear pain and have to keep their distance from the speaker; this could be due to the speaker being too close to the microphone or speaking too forcefully. In yet another scenario, fluctuating call volume, with both of the aforementioned problems occurring simultaneously within a single audio clip, creates an unpleasant and unsettling experience for the listener. Finally, when multiple people are speaking simultaneously and a user wants to hear one speaker clearly, interference from other speakers may occur.

[0039] At least one embodiment of this disclosure provides a method for adjusting the volume of an online multi-person call, comprising: acquiring audio data of an online multi-person call in real time to obtain current audio data, wherein the current audio data includes audio data of multiple speakers; separating the audio data of the multiple speakers to obtain audio features of each speaker; verifying the identities of the multiple speakers based on the audio features of each speaker to obtain audio data corresponding to the identity of each speaker; and selecting a target speaker from the multiple speakers based on user instructions and the audio data corresponding to the identity of each speaker, and adjusting the volume value of the target speaker.

[0040] At least one embodiment of this disclosure also provides an online multi-person call volume adjustment device, electronic device, and storage medium for implementing the online multi-person call volume adjustment method of the above embodiments.

[0041] The method, apparatus, electronic device, and storage medium provided in at least one embodiment of this disclosure can designate a target speaker and adjust the volume of the target speaker during online multi-person calls, thereby preventing interference from multiple people talking, improving the experience of online multi-person calls, reducing communication costs, and increasing the efficiency of online multi-person calls.

[0042] At least one embodiment of the present disclosure will now be described in detail with reference to the accompanying drawings. It should be noted that the same reference numerals will be used to refer to the same elements described in different drawings.

[0043] Figure 1 is an exemplary flowchart of an online multi-person call volume adjustment method provided in at least one embodiment of the present disclosure.

[0044] For example, as shown in Figure 1, the online multi-person call volume adjustment method may include the following steps S110 to S140.

[0045] Step S110: Acquire audio data of online multi-person calls in real time to obtain the current audio data.

[0046] For example, in step S110, the current audio data contains audio data from multiple speakers. For instance, during an online multi-person call, there may be multiple speakers speaking simultaneously at the current moment; in this case, the acquired current audio data may contain audio data from multiple speakers.

[0047] It should be noted that the current time can be any moment in an online multi-person call, and the current audio data can be audio data acquired in real time at any moment. The specific selection can be made according to actual needs, and the embodiments disclosed herein do not impose any restrictions on this.

[0048] For example, in step S110, audio data can be acquired in real time using the `pyaudio` library. Other methods can be selected for real-time acquisition as needed, and the embodiments disclosed herein do not limit this.

[0049] Step S120: Separate the audio data of multiple speakers to obtain the audio features of each speaker.

[0050] For example, in step S120, a speaker separation model can be used to separate the audio data of multiple speakers, separating the audio features of different speakers speaking simultaneously to obtain the audio features of each speaker. For example, the speaker separation model can be a trained neural network model, other types of processing models, or specific processors or other hardware modules. The specific model can be selected according to actual needs, and the embodiments of this disclosure do not limit this.

[0051] Step S130: Based on the audio features of each speaker, verify the identities of multiple speakers and obtain audio data corresponding to the identity of each speaker.

[0052] For example, in step S130, the audio features of each speaker obtained in step S120 can be input into the speaker identification model to identify the identity of the corresponding speaker and output the audio data corresponding to the identity of the corresponding speaker. For example, the speaker identification model can be a trained neural network model, other types of processing models, or specific processors or other hardware modules. The specific model can be selected according to actual needs, and the embodiments of this disclosure do not limit this.

[0053] Step S140: Based on user instructions and audio data corresponding to the identity of each speaker, select the target speaker from multiple speakers and adjust the volume of the target speaker.

[0054] For example, in step S140, the user instruction may include a speaker selection instruction. Based on the speaker selection instruction, a target speaker can be selected according to the user's needs. For example, the target speaker can be selected by identifying the identity information in each audio data output in step 130. For example, the user instruction may also include a volume adjustment instruction. Based on the volume adjustment instruction, the volume of the target speaker can be adjusted to the user's desired volume value. For example, the position of the volume slider can be determined based on the volume adjustment instruction, and the user can adjust the volume of the target speaker to the desired volume value by adjusting the volume slider. Alternatively, a user-preset or custom target volume value can be preset. After the user selects a target speaker, the system (e.g., the system provided in at least one embodiment of this disclosure may include, but is not limited to, a computer system, a neural network system, or other processor devices) can automatically adjust the volume of the target speaker to the target volume value.

[0055] In some examples, step S120 of Figure 1 may further include the following steps S121 to S122.

[0056] Step S121: Extract features from the audio data of multiple speakers to obtain the audio features of multiple speakers.

[0057] For example, in step S121, in order to analyze the audio data, feature extraction can be performed on the audio data first, so as to further analyze and process the extracted audio features of multiple speakers. For example, filter bank features (hereinafter referred to as FBank features), Mel-Frequency Cepstral Coefficients (MFCC) features, or other types of features can be extracted from the audio data of multiple speakers. Other types of features can also be selected to be extracted according to actual needs. The embodiments of this disclosure do not limit this.

[0058] For example, taking the extraction of FBank features from audio data of multiple speakers as an example, the audio data of multiple speakers is first pre-emphasized to enhance the high-frequency components of the audio signal. The reason for pre-emphasis is that when human vocal organs radiate sound waves, the air, as the carrier (or load) of the speech signal, both transmits and loses energy; and the medium, as the carrier of sound energy, experiences greater energy loss at higher frequencies, given a fixed sound source size. For example, audio signals often exhibit spectral tilt, meaning the amplitude of high-frequency components is smaller than that of low-frequency components. Pre-emphasis can balance the spectrum and increase the amplitude of high-frequency components. For specific implementation methods of pre-emphasis, please refer to the descriptions in this field, which will not be elaborated here.

[0059] For example, further, the pre-emphasized audio data can be framed and windowed. For instance, an audio signal is a continuous segment, typically several hundred milliseconds for streaming and several seconds or longer for non-streaming. Therefore, it's necessary to divide the variable-length audio data into fixed-length segments. For instance, the reason for framing is that the frequencies in the signal change over time (unstable). Some signal processing algorithms (such as Fourier transform) typically prefer a stable signal because the frequency profile is lost over time. To avoid this, the signal needs to be framed, assuming that the signal within each frame is short-term invariant. For instance, a frame can be 10-30ms, providing enough periods without drastic changes. For instance, to avoid missing signals at window boundaries, frame overlap (i.e., frames need to overlap slightly) is needed when shifting frames to prevent excessive changes in characteristics between frames. For instance, subsequent operations are performed on single frames; for example, a 25ms frame can be selected with a 10ms frame shift.

[0060] For example, after framing, each frame's signal typically needs to be windowed, meaning each frame's signal is multiplied by a smooth window function. The purpose is to smoothly attenuate the signals at both ends of the frame, reducing the intensity of sidelobes after the subsequent Fourier transform and achieving a higher quality spectrum. For example, the time difference between frames can be set to 10ms, allowing for overlap between frames; otherwise, the signal at the frame connection point would be weakened by windowing, resulting in information loss. For example, the Fourier transform is performed frame-by-frame to obtain the spectrum of each frame; generally, only the amplitude spectrum is retained, while the phase spectrum is discarded. For example, speech fluctuates continuously over a long range and lacks fixed characteristics, making it difficult to process. Therefore, each frame is substituted into a window function, with the out-of-window value set to 0. This aims to eliminate potential signal discontinuities at the ends of frames. For example, window functions can include, but are not limited to, rectangular windows, Hamming windows, and Hanning windows, and the specific choice depends on actual needs. For example, the specific implementation methods of framing and windowing can be found in descriptions within this field and will not be elaborated upon here.

[0061] For example, after framing and windowing, the resulting signal is still in the time domain. To extract Fbank features, it needs to be further converted from the time domain to the frequency domain, for example, through Fourier transform. For instance, taking Fourier transform as an example, for digital audio, the Discrete Fourier Transform can be used. After the Fourier transform, the frequency domain signal is obtained; since the energy levels vary across each frequency band, the energy spectra of different phonemes are different. For example, the energy spectrum can be further calculated by taking the square of the modulus (or taking the logarithm, squaring, etc.) of the audio signal's spectrum to obtain the spectral line energy of the audio signal.

[0062] For example, after the above steps, applying a Mel filter bank to the energy spectrum can extract FBank features. For instance, the Mel spectrum is a spectral representation method based on the characteristics of human hearing; it transforms a linear frequency axis into a non-linear Mel frequency axis, which better matches the human ear's perception of different frequencies. Logarithmic operations can compress the dynamic range of the spectrum, making the model easier to learn. Specifically, upper and lower frequency limits can be set to filter out certain unwanted or noisy frequency ranges (typically the lower limit is set to around 20Hz, and the upper limit is half the audio sampling rate), and these ranges can be converted to Mel frequencies. Then, a triangular filter bank with K channels is configured on the Mel frequency axis, where K is typically set to 40. Further, the logarithm of the Mel filtering result can be taken. Taking the logarithm is a scaling operation on the vertical axis, which can amplify energy differences at low energies; at a deeper level, this mimics the calculation steps of the cepstrum. For instance, after the above series of operations, FBank features of audio data from multiple speakers can be extracted.

[0063] It should be noted that the specific implementation methods of Fourier transform, energy spectrum calculation, Mel filter processing, etc. described above can be found in the descriptions in this field, and will not be repeated here.

[0064] Step S122: Input the audio features of multiple speakers into the speaker separation model, and output the audio features of each speaker after processing.

[0065] For example, in step S122, the specific processing procedure of the speaker separation model may include the following steps S1221 to S1224.

[0066] Step S1221: Encode the audio features of multiple speakers into a high-dimensional feature sequence;

[0067] Step S1222: Capture local features in the high-dimensional feature sequence, wherein the local features include dependencies corresponding to each speaker;

[0068] Step S1223: Based on local features, perform a separation operation on the high-dimensional feature sequence to obtain the feature sequence corresponding to each speaker;

[0069] Step S1224: Decode the feature sequence corresponding to each speaker into the audio features of each speaker.

[0070] Figure 2 is a schematic diagram of an example of a speaker separation model provided in at least one embodiment of this disclosure. For example, Figure 2 is a specific example of steps S1221 to S1224 above.

[0071] For example, as shown in Figure 2, a speaker separation model can include an encoder, a separator, and a decoder to separate the audio features of N speakers into audio features for each speaker (audio features of speaker 1, audio features of speaker 2, ..., audio features of speaker N), where N is an integer greater than 1.

[0072] For example, an encoder responsible for high-dimensional feature extraction might consist of one-dimensional convolutional layers (Conv1D) and rectified linear units (ReLUs), with the ReLUs constraining the encoder's output to non-negative values. A decoder, for example, could be a one-dimensional transposed convolutional layer using the same kernel size and stride as the encoder. A separator, for example, could be implemented using a masking network that performs a non-linear mapping from the encoder's output to N sets of masks. For example, the main component of a masking network might be a separation module, which could consist of four convolutional modules, scaling and offset operations, joint local and global single-head self-attention (SHSA), and three gating operations, responsible for processing long sequences.

[0073] For example, as shown in Figure 2, after the audio features of N speakers are input into the encoder, in step S1221, the encoder encodes the audio features of the N speakers into a high-dimensional feature sequence and outputs it to the separator.

[0074] For example, in step S1222, the separator can capture local features in the high-dimensional feature sequence, which may include dependencies corresponding to each speaker. For example, there may be a corresponding dependency between any two positions in the high-dimensional feature sequence, and this dependency is used to represent the semantic relationship between the two positions. For example, the dependency corresponding to each speaker is used to represent the semantic relationship between positions in the high-dimensional feature sequence associated with the corresponding speaker; by capturing this dependency, the features of different speakers in the high-dimensional feature sequence can be identified. For example, the self-attention mechanism in the separator can be used to capture the above-mentioned dependencies, as described in the art, and will not be repeated here; other methods can also be used to capture the above-mentioned dependencies, and the specific method can be selected according to actual needs. The embodiments of this disclosure do not limit this.

[0075] For example, in step S1223, based on the dependencies contained in the local features, the separator can perform a separation operation on the high-dimensional feature sequence to obtain the feature sequence corresponding to each speaker. For instance, taking a masking network as an example, the masking network can predict a separation mask based on local features; the separation mask is a matrix with the same dimension as the high-dimensional feature sequence, representing the contribution ratio of different speakers at each frequency and time point. For example, for an audio feature mixture of N speakers, the separator will predict N separation masks, each corresponding to one of the N speakers. For example, the mask type can include a ratio mask, etc., which obtains the spectral features of each speaker after separation by multiplying the spectral features of the mixed audio with the separation mask.

[0076] For example, in step S1224, the decoder can decode the feature sequence corresponding to each speaker obtained in step S1223 into audio features for each speaker. For example, the decoder can convert the separated spectral features back into audio waveforms in the time domain, for example, by using the inverse short-time Fourier transform (ISTFT).

[0077] It should be noted that the speaker separation model in Figure 2 is only an example. The specific structure and implementation of the encoder, separator and decoder are not limited to those described above. The speaker separation model can also be selected with other structures and implementations. The specific choice can be made according to actual needs. The embodiments of this disclosure do not limit this.

[0078] In some examples, step S130 of Figure 1 may include the following steps S131 to S133.

[0079] Step S131: Input the audio features of each speaker into the speaker identification model;

[0080] Step S132: Use the speaker identification model to extract the identity features of each speaker from the audio features of each speaker;

[0081] Step S133: Confirm the identity of the corresponding speaker based on the identity characteristics of each speaker, and output the audio data corresponding to the identity of each speaker.

[0082] For example, in step S131, the audio features of each speaker after separation can be used as input to identify the speaker using a speaker identification model.

[0083] For example, in step S132, the specific processing procedure for extracting the identity features of each speaker using the speaker separation model may include the following steps S1321 to S1324.

[0084] Step S1321: Preprocess the audio features of the first speaker among multiple speakers to obtain the preprocessed audio features of the first speaker;

[0085] Step S1322: Extract contextual information from the preprocessed audio features of the first speaker to obtain a first feature map containing contextual information;

[0086] Step S1323: Mask the first feature map to remove noise from the first feature map and retain or enhance the key information of the first speaker;

[0087] Step S1324: Extract the identity features of the first speaker based on the masked first feature map.

[0088] It should be noted that, in addition to the methods described in steps S1321 to S1324 above, the speaker separation model can also extract the identity features of each speaker in other ways. The specific method can be selected according to actual needs, and the embodiments of this disclosure do not limit this.

[0089] For example, in step S133, the identity features of each speaker obtained in step S132 can be processed in a specific way to confirm the identity of the corresponding speaker. For example, the specific processing method may include cluster analysis, or it may be to predict the probability distribution of each possible speaker identity and select the speaker identity with the highest probability as the recognition result, or other processing methods may be used to confirm the speaker identity. The specific method can be selected according to actual needs, and the embodiments of this disclosure do not limit this.

[0090] For example, taking the cluster analysis processing method as an example, step S133 can further include step S1331: based on the information in the speaker identity database, perform cluster analysis on the identity features of each speaker to confirm the speaker's identity.

[0091] For example, in one possible implementation of step S1331, the cluster analysis process may further include: comparing the identity features of each speaker with the information in the speaker identity database to obtain a similarity score of each speaker's identity features relative to the speaker identity database; and confirming that the first speaker's identity features are the first identity if the similarity score of the first speaker's identity features relative to the first identity in the speaker identity database is higher than a preset value.

[0092] It should be noted that the first speaker can be any one of multiple speakers, and the specific choice can be made according to actual needs. The embodiments disclosed herein do not impose any restrictions on this.

[0093] Figure 3 is a schematic diagram of an example of a speaker identification model provided in at least one embodiment of this disclosure. For example, Figure 3 is a specific example of the above steps S131 to S133.

[0094] For example, as shown in Figure 3, the speaker identification model may include a voice activity endpoint module, a voiceprint model, and a classifier, which are used to identify the first speaker based on the first speaker's audio features and output audio data corresponding to the first speaker's identity.

[0095] For example, as shown in Figure 3, in step S131, the audio features of the first speaker are input into the speaker identification model. For example, in step S1321, the input audio features of the first speaker can be preprocessed using a speech activity endpoint detection module to obtain preprocessed audio features of the first speaker. For example, specifically, the speech activity endpoint detection module can remove non-human voice parts from the audio features of the first speaker, and then segment the processed audio features according to a fixed window shift and window length to obtain multiple audio segments.

[0096] For example, as shown in Figure 3, the voiceprint model is further used to extract key information from these audio segments to obtain the identity information of the first speaker. For instance, the voiceprint model can be a voiceprint recognition model based on a densely connected time-delay neural network, using a residual convolutional network as the front-end module and a time-delay neural network structure as the backbone module. For example, the front-end module can be a 2D convolutional structure used to extract more local and refined time-frequency features; the backbone module can employ dense connections to reuse hierarchical features and improve computational efficiency. For example, each layer of the backbone module can embed a lightweight context-aware mask module, which extracts contextual information at different scales through multi-granularity pooling operations. The generated mask can remove irrelevant noise from the features while retaining key speaker information.

[0097] For example, specifically, in step S1322, the voiceprint model can be used to extract contextual information from the preprocessed audio features of the first speaker to obtain a first feature map containing contextual information; in step S1323, the voiceprint model can further mask the first feature map to remove noise from the first feature map and retain or enhance the key information of the first speaker; in step S1324, the voiceprint model can extract the identity features of the first speaker based on the masked first feature map.

[0098] For example, as shown in Figure 3, further, to confirm the speaker's identity, a classifier can be used to classify the processed features, identify the identity information of each speaker, and output the corresponding time information. For example, if a speaker transition point recognition model is used in conjunction with determining the location of the speaker transition point, the identity recognition will be more accurate. For example, specifically, in step S133, a classifier can be used to perform specific processing on the identity features of each speaker obtained in step S132. For example, taking cluster analysis as an example, the identity features of the first speaker can be compared with the information in the speaker identity database to obtain the similarity score of the first speaker's identity features relative to the speaker identity database; in response to the first speaker's identity features having a similarity score relative to the first identity in the speaker identity database that is higher than a preset value, the first speaker is confirmed as the first identity.

[0099] It should be noted that the speaker identification model in Figure 3 is only an example. The structure and specific implementation of the voice activity endpoint detection module, voiceprint model, classifier and speaker transition point recognition model are not limited to the above description. The speaker identification model can also be selected with other structures and implementations. The specific selection can be made according to actual needs. The embodiments of this disclosure do not limit this.

[0100] In some examples, user instructions may include speaker selection instructions, and the audio data corresponding to each speaker's identity includes the corresponding speaker's identity information. For example, step S140 in Figure 1 may further include the following step S141.

[0101] Step S141: Select the target speaker from multiple speakers based on the speaker selection instruction and the identity information of each speaker.

[0102] For example, in step S141, the user can issue a selection command based on the identity information of each speaker. For example, the identity information of each speaker can be displayed on the screen, and the user can issue a selection command by manually clicking on the speaker's identity information on the screen; or, the speaker selection command and the speaker's identity information can also be presented in other forms, which can be selected according to actual needs, and the embodiments of this disclosure do not limit this.

[0103] In some examples, the user instructions may also include volume adjustment instructions to determine the position of the volume slider; the audio data corresponding to each speaker's identity includes the volume information of the corresponding speaker. For example, step S140 in Figure 1 may further include the following steps S142 to S143.

[0104] Step S142: Based on the volume adjustment command, adjust the volume slider to the first position;

[0105] Step S143: Based on the target speaker's volume information and first position, obtain the target speaker's current volume value.

[0106] For example, the volume slider corresponds to a volume adjustment range, which includes multiple volume preset values, and the first position of the volume slider corresponds to the first volume preset value among the multiple volume preset values. For example, step S143 may further include: determining the first volume preset value based on the first position; and obtaining the current volume value of the target speaker based on the target speaker's volume information and the first volume preset value.

[0107] For example, in some examples, step S140 of Figure 1 may also include the following steps S1401 to S1404.

[0108] Step S1401: Display the identity information of multiple speakers on the screen;

[0109] Step S1402: Select the target speaker based on the speaker selection instruction and the identity information of multiple speakers displayed on the screen, wherein the speaker selection instruction is generated by the user clicking on the identity information of the target speaker on the screen;

[0110] Step S1403: Based on the volume adjustment command, adjust the volume slider to the first position, wherein the first position corresponds to the first volume preset value in the volume adjustment range of the volume slider;

[0111] Step S1404: Based on the target speaker's volume information and the first volume preset value, obtain the target speaker's current volume value.

[0112] It should be noted that the first position can be any position within the range of the volume slider, and can be selected according to actual needs. The embodiments of this disclosure do not impose any limitations on this. Furthermore, in addition to determining the position of the volume slider, the volume adjustment command can also represent the volume the user wants to adjust in other forms, and can be selected according to actual needs. The embodiments of this disclosure do not impose any limitations on this. Figure 4A is a schematic diagram of an example of the relationship between the volume slider position and the volume level provided by at least one embodiment of this disclosure. Figure 4B is a schematic diagram of an example of the logarithmic relationship between the volume slider position and the volume level provided by at least one embodiment of this disclosure. Figure 4C is a schematic diagram of an example of an online multi-person call volume adjustment method provided by at least one embodiment of this disclosure. For example, Figures 4A to 4C are specific examples of the above steps S141 to S143.

[0113] For example, the volume variation range between two sounds can be calculated using the following formula (1), in decibels:

[0114] dB=20*lg(A1 / A2) Formula (1)

[0115] Here, A1 and A2 are the amplitudes of two sounds, representing the size of each sound sample in the program. For example, when the sound sample size (i.e., the quantization depth) is 1 bit, the dynamic range is 0 because there can only be one amplitude. For example, when the sound sample size is 8 bits (i.e., one byte), the maximum amplitude is 256 times the minimum amplitude, and the volume change range calculated using formula (1) is 48 dB (i.e., dB = 20 * lg (256)); the volume change range of 48 dB is roughly the difference between a quiet room and a running lawnmower. For example, if the sound sample size is doubled to 16 bits, the resulting volume change range is 96 dB (i.e., dB = 20 * lg (65536)).

[0116] For example, in Figures 4A to 4C, taking a volume change range of 96 dB as an example, the corresponding sound amplitude range is 0 to V (where V represents the maximum value of the sound amplitude calculated based on the volume change range).

[0117] For example, as shown in the left image of Figure 4A, the position of the volume slider has a linear relationship with the sound amplitude. For example, the right image of Figure 4A shows the relationship between the perceived volume and the slider position; as shown in the right image of Figure 4A, moving the slider the same distance on the left side results in a large perceived range of volume change, while moving the slider the same distance near the maximum sound value on the right side results in a very small perceived range of volume change.

[0118] For example, as shown in the left image of Figure 4B, the volume slider's position increases logarithmically with the sound amplitude. Similarly, as shown in the right image of Figure 4B, the perceived sound change is the same regardless of the volume slider's position or the distance it is moved. It should be noted that the slider's minimum position is only close to 0, not zero, because x > 0 in the logarithmic function y = lg x.

[0119] For example, based on the above decibel formula (1), the reference sound amplitude A2 is taken as the original sound amplitude, and A1 is the adjusted sound amplitude. The adjusted sound amplitude can be expressed by the following formula (2):

[0120] A1=A2*pow(10,db / 20) Formula (2)

[0121] For example, taking the example of Figure 4C, in step S1401, the identity information of each speaker can be displayed on the screen; in steps S141 / S1402, the user manually clicks "Xiao Zhang" on the screen to issue a speaker selection instruction and selects "Xiao Zhang" as the target speaker.

[0122] For example, as shown in Figure 4C, suppose a user wants to adjust the volume of the target speaker to half the volume level. In steps S142 / S1403, the user issues a volume adjustment command to adjust the volume slider to the first position (the position at half the volume slider range shown in Figure 4C).

[0123] For example, taking an ideal volume adjustment step of 2dB as an example, for a 96dB range, dividing it into 48 parts by a 2dB step, the volume adjustment range corresponding to the volume slider can be obtained as [-96dB, -94dB, -92dB, ..., -4dB, -2dB, 0dB]. This volume adjustment range includes multiple volume preset values ​​-96dB, -94dB, -92dB, ..., -4dB, -2dB, 0dB. For example, in steps S143 / S1404, based on the first position located at half of the volume slider range, -48dB in this volume change range can be obtained as the first volume preset value; substituting this first volume preset value into the above formula (2), the current volume value A1 after adjustment by the target speaker can be calculated as A2*pow(10, -48 / 20), where the volume information of the target speaker includes the target speaker's voice amplitude A2. For example, based on the above adjustment rules, the relationship between each position within the volume slider range and the current volume value of the target speaker can be calculated, thereby realizing the adjustment of the target speaker's volume value according to user needs.

[0124] It should be noted that the volume adjustment methods in Figures 4A to 4C are only examples. The implementation methods for adjusting the volume of the target speaker are not limited to those described above. The specific method can be selected according to actual needs. The embodiments disclosed herein do not limit this.

[0125] The online multi-person call volume adjustment method provided in at least one embodiment of this disclosure can specify the target speaker according to the user's needs during online multi-person calls, and the user can adjust the volume value of the target speaker to the desired target volume value by adjusting the volume slider, thereby preventing interference from multiple people talking and improving the online multi-person call experience.

[0126] In other examples, the audio data corresponding to each speaker's identity includes the speaker's volume information. For example, step S140 in Figure 1 may also include step S145.

[0127] Step S145: Based on the target speaker's volume information, automatically adjust the target speaker's volume value to the target volume value.

[0128] For example, in step S145, the target volume value can be a volume value that the user is accustomed to or a volume value that the user has defined, or it can be selected as other volume values ​​according to actual needs. The embodiments of this disclosure do not limit this.

[0129] For example, in step S145, the system can automatically adjust the volume to the user's desired target volume level without requiring manual adjustment of the volume slider. For instance, in some examples, the volume can be automatically adjusted by adjusting the gain compensation of the acquisition device. Specifically, if the speaker's voice is too loud, the system will automatically reduce the gain of the acquisition device; conversely, it will automatically increase the gain to ensure the volume remains at a relatively stable level, and the output is processed and quantized data. During this process, the user does not need to frequently operate the device, thus avoiding unpleasant experiences caused by fluctuating sound levels and improving the comfort and user experience of multi-person online calls.

[0130] For example, step S145 may further include the following steps S1451 to S1454.

[0131] Step S1451: After the target speaker's volume information is processed by the amplifier, the volume information of the target speaker after the first processing is sampled to obtain multiple first sampled signals;

[0132] Step S1452: Perform a second processing on the multiple first sampled signals to obtain the average power value of the multiple first sampled signals;

[0133] Step S1453: Compare the average power value of multiple first sampled signals with the target volume value to obtain a first comparison result;

[0134] Step S1454: Based on the first comparison result, automatically adjust the amplifier gain to automatically adjust the target speaker's volume value to the target volume value.

[0135] For example, in step S1451, an amplifier can be used to perform a first processing (e.g., amplification operation, etc.) on the input target speaker's volume information, and then the first-processed target speaker's volume information can be sampled to obtain multiple first sampled signals.

[0136] For example, in step S1452, the second processing of the multiple first sampled signals may include: converting the multiple first sampled signals into digital signals using an analog-to-digital converter (ADC); performing digital filtering on the digital signals converted by the ADC to remove high-frequency noise and low-frequency drift; squaring the digitally filtered signals to obtain the power values ​​of the multiple first sampled signals; and performing a moving average on the squared signals to obtain the average power value of the multiple first sampled signals. For example, the second processing may also include other processing methods, which can be selected according to actual needs, and the embodiments of this disclosure do not limit this.

[0137] For example, in step S1453, the target volume value is set as a comparison threshold, and the average power value of multiple first sampled signals is compared with the set threshold to obtain a first comparison result.

[0138] For example, in step S1454, the gain of the front-end amplifier is automatically adjusted according to the first comparison result so that the intensity of the target speaker's volume information is kept within a set range, thereby adjusting the target speaker's volume value to the target volume value.

[0139] It should be noted that, in addition to the methods described above in steps S1451 to S1454, the system can also automatically adjust the volume of the target speaker in other ways. The specific method can be selected according to actual needs, and the embodiments disclosed herein do not limit this.

[0140] The online multi-person call volume adjustment method provided in at least one embodiment of this disclosure can specify a target speaker and adjust the volume value of the target speaker during an online multi-person call, thereby preventing interference from multiple people talking, improving the experience of online multi-person calls, reducing communication costs, and improving the efficiency of online multi-person calls.

[0141] In some examples, prior to step S110 in FIG1, the online multi-person call volume adjustment method provided in at least one embodiment of the present disclosure may further include the following step S150.

[0142] Step S150: Obtain the target volume value.

[0143] For example, in step S150, the target volume value is a volume value that the user is accustomed to or a user-defined volume value. Alternatively, the target volume value can also be selected as a volume value required by other users, as the embodiments of this disclosure do not impose such limitations.

[0144] For example, in response to the target volume value being a volume value that the user is accustomed to, step S150 may further include: acquiring and storing multiple volume values ​​set by the user locally, and selecting the first volume value that the user sets most frequently from the multiple volume values ​​as the target volume value.

[0145] For example, multiple volume values ​​set by the user can be multiple volume values ​​set by the user within a certain period of time. If a certain value is set the most times, then that volume value is the volume value that the user is accustomed to. For example, multiple volume values ​​stored locally can be obtained through the MMDeviceEnumerator interface in the Windows API, or they can be obtained through other methods as needed. The embodiments disclosed herein do not limit this.

[0146] For example, after step S110 in Figure 1, the online multi-person call volume adjustment method provided in at least one embodiment of this disclosure may further include the following steps S160 to S170.

[0147] Step S160: Extract the volume information of the current audio data;

[0148] Step S170: Based on the volume information of the current audio data, automatically adjust the current volume value of the online multi-person call to the target volume value.

[0149] For example, in step S160, volume information can be extracted by converting the current audio data to decibel levels, or it can be extracted in other ways. The embodiments of this disclosure do not limit this method. For example, taking the conversion to decibel levels as an example, the root mean square (RMS) value of the current audio data can be calculated, thereby converting it to decibel levels, i.e., the current volume value. For example, as shown in the following formula (3), by calculating the amplitude of all sound samples (X1, X2, ... X...) within the time series... M The sum of the squares of the samples, divided by the total number of samples M (where M is a positive integer), and then the square root is taken to obtain the RMS value of the current audio data.

[0150] For example, in step S170, the system can automatically adjust the current volume of the online multi-person call to the target volume value required by the user; the automatic adjustment method can adopt the automatic gain control (AGC) algorithm, or other adjustment methods can be selected according to actual needs, and the embodiments of this disclosure do not limit this.

[0151] For example, step S170 may further include the following steps S171 to S174.

[0152] Step S171: After the amplifier performs the first processing on the volume information of the current audio data, the volume information of the first processed current audio data is sampled to obtain multiple second sample signals;

[0153] Step S172: Perform a second processing on the multiple second sampled signals to obtain the average power value of the multiple second sampled signals;

[0154] Step S173: Compare the average power value of multiple second sampled signals with the target volume value to obtain a second comparison result;

[0155] Step S174: Based on the second comparison result, automatically adjust the amplifier gain to automatically adjust the current volume value of the online multi-person call to the target volume value.

[0156] For example, the specific implementation methods of steps S171 to S174 are basically the same as those of steps S1451 to S1454. For details, please refer to the description of steps S1451 to S1454 above, which will not be repeated here.

[0157] It should be noted that, in addition to the methods described in steps S171 to S174 above, the system can also automatically adjust the current volume value of online multi-person calls in other ways. The specific method can be selected according to actual needs, and the embodiments disclosed herein do not limit this.

[0158] The online multi-person call volume adjustment method provided in at least one embodiment of this disclosure can sense the volume of remote speakers through an algorithm, and by setting the target volume value to a volume value that the user is accustomed to or that the user has customized, the volume adjustment of online multi-person calls can be customized, thereby improving the comfort and user experience of online multi-person calls and increasing the efficiency of completing online multi-person calls.

[0159] Figure 5 is another exemplary flowchart of an online multi-person call volume adjustment method provided in at least one embodiment of the present disclosure. For example, Figure 5 is an exemplary flowchart of steps S110 to S170 described above.

[0160] For example, as shown in Figure 5, in the online multi-person call volume adjustment method provided in at least one embodiment of this disclosure, on the one hand, the volume of the online multi-person call can be automatically adjusted to the target volume value required by the user; on the other hand, the target speaker can be specified according to the user's needs, and the volume value of the target speaker can be adjusted according to the user's needs.

[0161] For example, as shown in Figure 5, firstly, in step S150, the user's preferred or user-defined target volume value is obtained and saved; furthermore, after starting an online multi-person call, in step S110, the audio data of the online multi-person call is obtained in real time to obtain the current audio data. For example, the current audio data contains audio data of multiple speakers.

[0162] For example, as shown in Figure 5, in order to automatically adjust the volume of an online multi-person call to the target volume value required by the user, in step S160, the volume information of the current audio data is extracted; in step S170, based on the volume information of the current audio data, the current volume value of the online multi-person call is automatically adjusted to the target volume value.

[0163] For example, as shown in Figure 5, in order to adjust the volume of a specified target speaker according to user needs, in step S120, the audio data of multiple speakers are separated to obtain the audio features of each speaker; in step S130, the identities of multiple speakers are confirmed based on the audio features of each speaker to obtain audio data corresponding to the identity of each speaker; in step S140, based on the user instruction and the audio data corresponding to the identity of each speaker, the target speaker is selected from the multiple speakers, and the volume of the target speaker is adjusted.

[0164] The online multi-person call volume adjustment method provided in at least one embodiment of this disclosure can, on the one hand, automatically adjust the volume of online multi-person calls according to user habits or user needs. The automatic volume adjustment in online multi-person calls makes the user experience more technological and improves the comfort of online communication. On the other hand, it can increase or decrease the volume of a specified speaker according to user needs to prevent interference from multiple people talking, thereby saving communication costs and improving communication efficiency in online multi-person calls.

[0165] Figure 6 is a schematic block diagram of an online multi-person call volume adjustment device provided in at least one embodiment of the present disclosure.

[0166] For example, as shown in Figure 6, the online multi-person call volume adjustment device 200 includes an acquisition module 210, a speaker separation module 220, a speaker identity confirmation module 230, and a speaker volume adjustment module 240.

[0167] For example, the acquisition module 210 can be configured to acquire audio data from online multi-person calls in real time to obtain the current audio data. For example, the current audio data may contain audio data from multiple speakers. That is, the acquisition module 210 can be configured to perform step S110 as shown in Figure 1.

[0168] For example, the speaker separation module 220 can be configured to separate the audio data of multiple speakers to obtain the audio features of each speaker. That is, the speaker separation module 220 can be configured to perform step S120 as shown in Figure 1.

[0169] For example, the speaker identification module 230 can be configured to identify multiple speakers based on the audio characteristics of each speaker, and obtain audio data corresponding to the identity of each speaker. That is, the speaker identification module 230 can be configured to perform step S130 as shown in Figure 1.

[0170] For example, the speaker volume adjustment module 240 can be configured to select a target speaker from multiple speakers based on user instructions and audio data corresponding to the identity of each speaker, and adjust the volume value of the target speaker. That is, the speaker volume adjustment module 240 can be configured to perform, for example, step S140 shown in FIG1.

[0171] In some examples, the speaker separation module 220 can also be configured to: extract features from the audio data of multiple speakers to obtain audio features of multiple speakers; input the audio features of multiple speakers into the speaker separation model, process and output the audio features of each speaker.

[0172] For example, the speaker separation module 220 can also be configured to: encode the audio features of multiple speakers into a high-dimensional feature sequence; capture local features in the high-dimensional feature sequence, wherein the local features include dependencies corresponding to each speaker; perform a separation operation on the high-dimensional feature sequence based on the local features to obtain a feature sequence corresponding to each speaker; and decode the feature sequence corresponding to each speaker into the audio features of each speaker.

[0173] In some examples, the speaker identification module 230 can also be configured to: input the audio features of each speaker into the speaker identification model; use the speaker identification model to extract the identity features of each speaker from the audio features of each speaker; identify the corresponding speaker based on the identity features of each speaker, and output the audio data corresponding to the identity of each speaker.

[0174] For example, the speaker identification module 230 can also be configured to: preprocess the audio features of the first speaker among multiple speakers to obtain preprocessed audio features of the first speaker; extract context information from the preprocessed audio features of the first speaker to obtain a first feature map containing context information; perform masking processing on the first feature map to remove noise in the first feature map and retain or enhance the key information of the first speaker; and extract the identity features of the first speaker based on the masked first feature map.

[0175] For example, the speaker identity verification module 230 can also be configured to perform cluster analysis on the identity features of each speaker based on information in the speaker identity database in order to verify the speaker's identity.

[0176] For example, the speaker identity verification module 230 can also be configured to: compare the identity features of each speaker with the information in the speaker identity database to obtain the similarity score of each speaker's identity features relative to the speaker identity database; and in response to the fact that the similarity score of the identity features of the first speaker among multiple speakers relative to the first identity in the speaker identity database is higher than a preset value, confirm the first speaker as the first identity.

[0177] In some examples, the user instructions include speaker selection instructions, and the audio data corresponding to each speaker's identity includes the corresponding speaker's identity information. For example, the speaker volume adjustment module 240 can also be configured to select a target speaker from multiple speakers based on the speaker selection instructions and each speaker's identity information.

[0178] In some examples, user instructions include volume adjustment instructions to determine the position of the volume slider; audio data corresponding to each speaker's identity includes the speaker's volume information. For example, the speaker volume adjustment module 240 can also be configured to: adjust the volume slider to a first position based on the volume adjustment instructions; and obtain the target speaker's current volume value based on the target speaker's volume information and the first position.

[0179] For example, the volume slider corresponds to a volume adjustment range, which includes multiple volume preset values. The first position of the volume slider corresponds to the first volume preset value among the multiple volume preset values. For example, the speaker volume adjustment module 240 can also be configured to: determine the first volume preset value based on the first position; and obtain the current volume value of the target speaker based on the target speaker's volume information and the first volume preset value.

[0180] In some examples, user instructions include speaker selection instructions and volume adjustment instructions. The volume adjustment instructions determine the position of the volume slider, and the audio data corresponding to each speaker's identity includes the speaker's identity information and the speaker's volume information. For example, the speaker volume adjustment module 240 can also be configured to: display the identity information of multiple speakers on the screen; select a target speaker based on the speaker selection instructions and the identity information of the multiple speakers displayed on the screen, wherein the speaker selection instructions are generated by the user clicking on the target speaker's identity information on the screen; adjust the volume slider to a first position based on the volume adjustment instructions, wherein the first position corresponds to a first volume preset value within the volume adjustment range of the volume slider; and obtain the current volume value of the target speaker based on the target speaker's volume information and the first volume preset value.

[0181] In some examples, the audio data corresponding to each speaker's identity includes the speaker's volume information. For example, the speaker volume adjustment module 240 can also be configured to automatically adjust the target speaker's volume value to the target volume value based on the target speaker's volume information.

[0182] For example, the speaker volume adjustment module 240 can also be configured to: perform a first processing on the target speaker's volume information by an amplifier, sample the first-processed target speaker's volume information to obtain multiple first sampled signals; perform a second processing on the multiple first sampled signals to obtain the average power value of the multiple first sampled signals; compare the average power value of the multiple first sampled signals with the target volume value to obtain a first comparison result; and automatically adjust the gain of the amplifier based on the first comparison result to automatically adjust the target speaker's volume value to the target volume value.

[0183] In some examples, before acquiring the audio data of the online multi-person call in real time, the acquisition module 210 can also be configured to acquire a target volume value. For example, the target volume value is a volume value that the user is accustomed to or a user-defined volume value.

[0184] For example, the acquisition module 210 can also be configured to: acquire and store multiple volume values ​​set by the user locally, and select the first volume value that is set most frequently by the user from the multiple volume values ​​as the target volume value.

[0185] For example, after acquiring the audio data of an online multi-person call in real time, the speaker volume adjustment module 240 can also be configured to: extract the volume information of the current audio data; and automatically adjust the current volume value of the online multi-person call to the target volume value based on the volume information of the current audio data.

[0186] For example, the speaker volume adjustment module 240 can also be configured to: perform a first processing on the volume information of the current audio data by an amplifier, sample the volume information of the first-processed current audio data to obtain multiple second sampled signals; perform a second processing on the multiple second sampled signals to obtain the average power value of the multiple second sampled signals; compare the average power value of the multiple second sampled signals with the target volume value to obtain a second comparison result; and automatically adjust the gain of the amplifier based on the second comparison result to automatically adjust the current volume value of the online multi-person call to the target volume value.

[0187] Since the details of the operation of the online multi-person call volume adjustment device 200 have already been introduced in the above description of the online multi-person call volume adjustment method shown in Figure 1, for the sake of brevity, they will not be repeated here. For relevant details, please refer to the above description of Figures 1 to 5.

[0188] It should be noted that the various modules in the online multi-person call volume adjustment device 200 shown in Figure 6 can be configured as software, hardware, firmware, or any combination thereof to perform specific functions. For example, these modules may correspond to dedicated integrated circuits, pure software code, or modules combining software and hardware. As an example, the device described with reference to Figure 6 may be a PC computer, tablet device, personal digital assistant, smartphone, web application, or other device capable of executing program instructions, but is not limited thereto.

[0189] Furthermore, although the online multi-person call volume adjustment device 200 has been divided into modules for performing corresponding processes in the description above, those skilled in the art will understand that the processes performed by each module can also be performed without any specific module division in the device or without clear boundaries between the modules. In addition, the online multi-person call volume adjustment device 200 described above with reference to FIG6 is not limited to including the modules described above, but may also include other modules (e.g., storage module, retrieval module, etc.) as needed, or the above modules may be combined.

[0190] At least one embodiment of this disclosure also provides an electronic device including a processor and a memory; the memory includes one or more computer program modules; the one or more computer program modules are stored in the memory and configured to be executed by the processor, and the one or more computer program modules include computer-executable instructions for implementing the online multi-person call volume adjustment method provided in the embodiments of this disclosure described above. For example, the processor may be a single-core processor or a multi-core processor.

[0191] Figure 7 is a schematic block diagram of an electronic device provided in at least one embodiment of the present disclosure.

[0192] For example, as shown in Figure 7, the electronic device 300 includes a processor 310 and a memory 320. For example, the memory 320 is used to store non-transitory computer-readable instructions (e.g., one or more computer program modules). The processor 310 is used to execute the non-transitory computer-readable instructions, which, when executed by the processor 310, can perform one or more steps of the online multi-person call volume adjustment method described above. The memory 320 and the processor 310 can be interconnected via a bus system and / or other forms of connection mechanism (not shown).

[0193] For example, processor 310 can be a central processing unit (CPU), graphics processing unit (GPU), general-purpose graphics processing unit (GPGPU), digital signal processor (DSP), or other processing unit with online multi-person call volume adjustment capability and / or program execution capability, such as field-programmable gate array (FPGA); for example, the central processing unit (CPU) can be an x86, RISC-V, or ARM architecture. Processor 310 can be a general-purpose processor or a special-purpose processor, and can control other components in electronic device 300 to perform desired functions.

[0194] For example, memory 320 may include any combination of one or more computer program products, which may include various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory. Volatile memory may include, for example, random access memory (RAM) and / or cache memory. Non-volatile memory may include, for example, read-only memory (ROM), hard disk, erasable programmable read-only memory (EPROM), portable optical disc read-only memory (CD-ROM), USB memory, flash memory, etc. One or more computer program modules may be stored on the computer-readable storage medium, and processor 310 may run one or more computer program modules to implement various functions of electronic device 300. Various application programs and various data, as well as various data used and / or generated by the application programs, may also be stored in the computer-readable storage medium.

[0195] It should be noted that, in the embodiments of this disclosure, the specific functions and technical effects of the electronic device 300 can be referred to the description above of the online multi-person call volume adjustment method provided by at least one embodiment of this disclosure, and will not be repeated here.

[0196] Figure 8 is a schematic block diagram of another electronic device provided in at least one embodiment of the present disclosure.

[0197] For example, as shown in FIG8, the electronic device 400 is adapted to implement the online multi-person call volume adjustment method provided in the embodiments of this disclosure. It should be noted that the electronic device 400 shown in FIG8 is merely an example and does not impose any limitation on the functionality and scope of use of the embodiments of this disclosure.

[0198] For example, as shown in FIG8, electronic device 400 may include a processing device (e.g., a central processing unit, a graphics processing unit, etc.) 41, which can perform various appropriate actions and processes according to a program stored in read-only memory (ROM) 42 or a program loaded from storage device 48 into random access memory (RAM) 43. RAM 43 also stores various programs and data required for the operation of electronic device 400. Processing device 41, ROM 42, and RAM 43 are interconnected via bus 44. Input / output (I / O) interface 45 is also connected to bus 44. Typically, the following devices can be connected to I / O interface 45: input devices 46 including, for example, touch screens, touchpads, keyboards, mice, cameras, microphones, accelerometers, gyroscopes, etc.; output devices 47 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices 48 including, for example, magnetic tapes, hard disks, etc.; and communication devices 49. Communication device 49 allows electronic device 400 to communicate wirelessly or wiredly with other electronic devices to exchange data.

[0199] Although Figure 8 shows an electronic device 400 with various devices, it should be understood that it is not required to implement or have all of the devices shown, and the electronic device 400 may alternatively implement or have more or fewer devices.

[0200] For detailed information and technical effects regarding the 400 electronic device, please refer to the above description of the online multi-person call volume adjustment method, which will not be repeated here.

[0201] Figure 9 is a schematic diagram of a storage medium provided in at least one embodiment of the present disclosure.

[0202] For example, as shown in Figure 9, storage medium 500 stores non-transitory computer-readable instructions 510. For example, when non-transitory computer-readable instructions 510 are executed by a computer, one or more steps in the online multi-person call volume adjustment method described above are performed.

[0203] For example, the storage medium 500 can be applied in the electronic device 300 shown in FIG. 7. For example, the storage medium 500 can be the memory 320 in the electronic device 300. For example, for relevant descriptions of the storage medium 500, please refer to the corresponding description of the memory 320 in the electronic device 300 shown in FIG. 7, which will not be repeated here.

[0204] The following points need to be clarified regarding this disclosure:

[0205] (1) The accompanying drawings of the embodiments of this disclosure only involve the structures involved in the embodiments of this disclosure. Other structures can be referred to the general design.

[0206] (2) Where there is no conflict, features of the same embodiment and different embodiments of this disclosure can be combined with each other.

[0207] The above are merely specific embodiments of this disclosure, but the scope of protection of this disclosure is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this disclosure should be included within the scope of protection of this disclosure. Therefore, the scope of protection of this disclosure should be determined by the scope of the claims.

Claims

1. A method for adjusting volume in online multi-person calls, comprising: The audio data of the online multi-person call is acquired in real time to obtain the current audio data, wherein the current audio data contains audio data of multiple speakers; The audio data of the multiple speakers are separated to obtain the audio features of each speaker. Based on the audio features of each speaker, the identities of the multiple speakers are confirmed to obtain audio data corresponding to the identity of each speaker; Based on user instructions and the audio data corresponding to the identity of each speaker, a target speaker is selected from the plurality of speakers, and the volume of the target speaker is adjusted.

2. The online multi-person call volume adjustment method according to claim 1, wherein, The step of separating the audio data of the multiple speakers to obtain the audio features of each speaker includes: Feature extraction is performed on the audio data of the multiple speakers to obtain the audio features of the multiple speakers; The audio features of the multiple speakers are input into the speaker separation model, and the audio features of each speaker are output after processing.

3. The online multi-person call volume adjustment method according to claim 2, wherein, The step of inputting the audio features of the multiple speakers into the speaker separation model, processing them, and outputting the audio features of each speaker includes: The audio features of the multiple speakers are encoded into a high-dimensional feature sequence; Capture local features in the high-dimensional feature sequence, wherein the local features include dependencies corresponding to each speaker; Based on the local features, the high-dimensional feature sequence is separated to obtain the feature sequence corresponding to each speaker; The feature sequence corresponding to each speaker is decoded into the audio features of each speaker.

4. The online multi-person call volume adjustment method according to any one of claims 1-3, wherein the step of verifying the identities of the multiple speakers based on the audio characteristics of each speaker to obtain audio data corresponding to the identity of each speaker includes: Input the audio features of each speaker into the speaker identification model; The speaker identity verification model is used to extract the speaker's identity features from the audio features of each speaker; Based on the identity characteristics of each speaker, the corresponding speaker identity is confirmed, and the audio data corresponding to the identity of each speaker is output.

5. The online multi-person call volume adjustment method according to claim 4, wherein extracting the identity features of each speaker from the audio features of each speaker using the speaker identity verification model includes: The audio features of the first speaker among the plurality of speakers are preprocessed to obtain the preprocessed audio features of the first speaker. Extract the contextual information from the preprocessed audio features of the first speaker to obtain a first feature map containing the contextual information; The first feature map is masked to remove noise from the first feature map and retain or enhance the key information of the first speaker; Based on the first feature map processed by the masking, the identity features of the first speaker are extracted.

6. The online multi-person call volume adjustment method according to claim 4 or 5, wherein, The process of confirming the identity of the corresponding speaker based on the identity characteristics of each speaker includes: Based on the information in the speaker identity database, cluster analysis is performed on the identity features of each speaker to confirm the speaker's identity.

7. The online multi-person call volume adjustment method according to claim 6, wherein, The step of performing cluster analysis on the identity features of each speaker based on information in the speaker identity database to confirm the speaker's identity includes: The identity features of each speaker are compared with the information in the speaker identity database to obtain a similarity score of each speaker's identity features relative to the speaker identity database. In response to the fact that the similarity score of the identity features of the first speaker among the plurality of speakers relative to the first identity in the speaker identity database is higher than a preset value, the first speaker is confirmed to be the first identity.

8. The online multi-person call volume adjustment method according to any one of claims 1-7, wherein, The user instructions include speaker selection instructions, and the audio data corresponding to the identity of each speaker includes the identity information of the corresponding speaker. The step of selecting a target speaker from the plurality of speakers based on user instructions and the audio data corresponding to the identity of each speaker includes: Based on the speaker selection instruction and the identity information of each speaker, a target speaker is selected from the plurality of speakers.

9. The online multi-person call volume adjustment method according to any one of claims 1-8, wherein, The user instructions include volume adjustment instructions, which are used to determine the position of the volume slider. The audio data corresponding to the identity of each speaker includes the volume information of the corresponding speaker. The step of adjusting the volume of the target speaker based on user instructions and the audio data corresponding to the identity of each speaker includes: Based on the volume adjustment command, adjust the volume slider to the first position; Based on the target speaker's volume information and the first position, the target speaker's current volume value is obtained.

10. The online multi-person call volume adjustment method according to claim 9, wherein, The volume slider corresponds to a volume adjustment range, which includes multiple preset volume values. The first position of the volume slider corresponds to the first preset volume value among the multiple preset volume values. The step of obtaining the current volume value of the target speaker based on the volume information of the target speaker and the first position includes: The first volume preset value is determined based on the first position; Based on the target speaker's volume information and the first preset volume value, the target speaker's current volume value is obtained.

11. The online multi-person call volume adjustment method according to any one of claims 1-10, wherein, The user instructions include speaker selection instructions and volume adjustment instructions. The volume adjustment instructions are used to determine the position of the volume slider. The audio data corresponding to each speaker's identity includes the speaker's identity information and the speaker's volume information. The step of selecting a target speaker from the plurality of speakers based on user instructions and the audio data corresponding to the identity of each speaker, and adjusting the volume of the target speaker, includes: Display the identity information of the multiple speakers on the screen; Based on the speaker selection instruction and the identity information of the plurality of speakers displayed on the screen, the target speaker is selected, wherein the speaker selection instruction is generated by the user clicking on the identity information of the target speaker on the screen; Based on the volume adjustment command, the volume slider is adjusted to a first position, wherein the first position corresponds to a first preset volume value within the volume adjustment range of the volume slider; Based on the target speaker's volume information and the first preset volume value, the target speaker's current volume value is obtained.

12. The online multi-person call volume adjustment method according to any one of claims 1-11, wherein, The audio data corresponding to the identity of each speaker includes the volume information of the corresponding speaker. The step of adjusting the volume of the target speaker based on user instructions and the audio data corresponding to the identity of each speaker includes: Based on the target speaker's volume information, the target speaker's volume value is automatically adjusted to the target volume value.

13. The online multi-person call volume adjustment method according to claim 12, wherein, The step of automatically adjusting the volume of the target speaker to the target volume value based on the target speaker's volume information includes: After the volume information of the target speaker is processed by the amplifier, the volume information of the target speaker after the first processing is sampled to obtain multiple first sampled signals; The plurality of first sampled signals are subjected to a second processing to obtain the average power value of the plurality of first sampled signals; The average power value of the plurality of first sampled signals is compared with the target volume value to obtain a first comparison result; Based on the first comparison result, the gain of the amplifier is automatically adjusted to automatically adjust the volume value of the target speaker to the target volume value.

14. The online multi-person call volume adjustment method according to any one of claims 1-13, wherein, Before acquiring the audio data of the online multi-person call in real time, the online multi-person call volume adjustment method further includes: Obtain the target volume value, wherein the target volume value is a volume value that the user is accustomed to or a volume value that the user defines.

15. The online multi-person call volume adjustment method according to claim 14, wherein, In response to the target volume value being a volume value that the user is accustomed to, obtaining the target volume value includes: Acquire and store multiple volume values ​​set by the user locally, and select the first volume value that is set most frequently by the user from the multiple volume values ​​as the target volume value.

16. The online multi-person call volume adjustment method according to claim 14 or 15, wherein, After acquiring the audio data of the online multi-person call in real time, the online multi-person call volume adjustment method further includes: Extract the volume information of the current audio data; Based on the volume information of the current audio data, the current volume value of the online multi-person call is automatically adjusted to the target volume value.

17. The online multi-person call volume adjustment method according to claim 16, wherein, The step of automatically adjusting the current volume value of the online multi-person call to the target volume value based on the volume information of the current audio data includes: After the volume information of the current audio data is processed by the amplifier, the volume information of the current audio data after the first processing is sampled to obtain multiple second sampling signals; The plurality of second sampled signals are subjected to a second processing to obtain the average power value of the plurality of second sampled signals; The average power value of the plurality of second sampled signals is compared with the target volume value to obtain a second comparison result; Based on the second comparison result, the gain of the amplifier is automatically adjusted to automatically adjust the current volume value of the online multi-person call to the target volume value.

18. An online multi-person call volume adjustment device, comprising: The acquisition module is configured to acquire the audio data of the online multi-person call in real time to obtain the current audio data, wherein the current audio data includes the audio data of multiple speakers; The speaker separation module is configured to separate the audio data of the multiple speakers to obtain the audio features of each speaker among the multiple speakers; The speaker identity verification module is configured to verify the identities of the multiple speakers based on the audio features of each speaker, and obtain audio data corresponding to the identity of each speaker. The speaker volume adjustment module is configured to select a target speaker from the plurality of speakers based on user instructions and audio data corresponding to the identity of each speaker, and adjust the volume value of the target speaker.

19. An electronic device comprising: processor; Memory, including one or more computer program modules; The one or more computer program modules are stored in the memory and configured to be executed by the processor, and the one or more computer program modules are used to implement the online multi-person call volume adjustment method according to any one of claims 1-17.

20. A storage medium storing non-transitory computer-readable instructions that, when executed by a computer, implement the online multi-person call volume adjustment method according to any one of claims 1-17.