Electronic apparatus, method performed by the same and storage medium
By acquiring and rendering audio signals from multiple locations to estimate accurate acoustic parameters, the method improves the consistency of reverberation in virtual audio rendering, addressing the issue of immersive experience discrepancies in mixed reality applications.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- SAMSUNG ELECTRONICS CO LTD
- Filing Date
- 2025-05-07
- Publication Date
- 2026-07-02
Smart Images

Figure KR2025095305_02072026_PF_FP_ABST
Abstract
Description
ELECTRONIC APPARATUS, METHOD PERFORMED BY THE SAME AND STORAGE MEDIUM
[0001] The present disclosure relates to a field of audio processing and a field of artificial intelligence, and specifically to a method performed by an electronic apparatus, the electronic apparatus, and a computer-readable storage medium.
[0002] When sound waves propagate in a space, they are reflected by objects (e.g., obstacles such as a wall, a ceiling, a floor, and so on) in the space (e.g., a room), and some of which are absorbed by the obstacles every time they are reflected. In this way, when a sound source stops making sound, sound waves in the space go through a plurality of reflections and absorption, and finally disappear, so you may feel there are several sound waves mixing and lasting for a period of time after the sound source stops making sound, this phenomenon of sound continuity which still exists after the sound source in the space stops making sound is called reverberation. Different acoustic environments have different spatial sizes and different materials of obstacles in the space, senses of the reverberation are different.
[0003] Mixed Reality (MR) spatial audio is a technology that combines virtual and real sounds and creates a shared audio space. For example, the technology may be used for all immersive and interactive applications, for example, online conferences or classes; remote orchestra rehearsals; and holographic entertainment events such as sports, games, concerts, movies, etc. When a virtual sound source and a real sound source coexist, different sound fields may affect immersive experience of the user, and it is important to eliminate the spatial perception gap between virtual and reality. A virtual audio from a virtual sound source may be rendered by using spatial audio rendering technologies so that it sounds like it is in the same sound field as the real sound, has the same reverberation, and sounds more natural. However, the rendering of the virtual audio requires the use of parameters for reflecting acoustic characteristics of the space, and if the parameters for reflecting the acoustic characteristics of the space are inaccurate, the audio rendering of the virtual audio will be affected, which ultimately results in difference in reverberation between the real sound and the virtual audio, thereby affecting the immersive experience of the user. In view of this, there is a need for a technology that can more accurately obtain the parameters for reflecting the acoustic characteristics of the space for better rendering of the virtual audio.
[0004] According to an embodiment of the disclosure, the method may include acquiring a plurality of first audio signals respectively recorded at a plurality of different locations in a space where a sound source is located. The method may include rendering, for each of the plurality of first audio signals, using a transmission response corresponding to at least one other first audio signal to obtain a rendered audio signal. The method may include estimating parameters for reflecting acoustic characteristics of the space based on the plurality of rendered audio signals. The method may include performing, based on the estimated parameters, audio rendering of a virtual audio to obtain a reverberation audio.
[0005] According to an embodiment of the disclosure, an electronic apparatus may be provided. The electronic apparatus may include memory storing one or more instructions. and at least one processor. The instructions when executed by the at least one processor individually or collectively, may cause the electronic apparatus to acquire a plurality of first audio signals respectively recorded at a plurality of different locations in a space where a sound source is located. The instructions when executed by the at least one processor individually or collectively, may cause the electronic apparatus to render, for each of the plurality of first audio signals, using a transmission response corresponding to at least one other first audio signal to obtain a rendered audio signal. The instructions when executed by the at least one processor individually or collectively, may cause the electronic apparatus to estimate parameters for reflecting acoustic characteristics of the space based on the plurality of rendered audio signals. The instructions when executed by the at least one processor individually or collectively, may cause the electronic apparatus to perform, based on the estimated parameters, audio rendering of a virtual audio to obtain a reverberation audio.
[0006] According to an embodiment of the disclosure, a computer-readable medium storing one or more instructions may be provided. The one or more instructions, when executed by at least one processor, may cause the at least one processor of an electronic apparatus to perform operation corresponding to the method.
[0007] It should be understood that the above general description and the detailed descriptions that follow are merely exemplary and explanatory and do not limit the present disclosure.
[0008] The accompanying drawings herein are incorporated into and form part of the specification, illustrate embodiments consistent with the disclosure, which are used in conjunction with the specification to explain the principles of the disclosure and do not constitute an undue limitation of the disclosure.
[0009] FIG. 1 is a schematic diagram illustrating audio ray tracing.
[0010] FIG. 2 is a flowchart illustrating a method performed by an electronic apparatus according to an embodiment of the present disclosure.
[0011] FIG. 3 is a diagram of an exemplary architectural illustrating a method performed by an electronic apparatus according to an embodiment of the present disclosure.
[0012] FIG. 4 is a schematic diagram illustrating re-rendering reverberation consistency in the case of a single sound source, according to an embodiment of the present disclosure.
[0013] FIG. 5 is a schematic diagram illustrating re-rendering reverberation consistency in the case of a plurality of sound sources according to an embodiment of the present disclosure.
[0014] FIG. 6 is a schematic diagram illustrating re-rendering reverberation consistency in the case of a single sound source and a presence of minor noise, according to an embodiment of the present disclosure.
[0015] FIG. 7 is a schematic illustrating an initialization of a material absorption coefficient.
[0016] FIG. 8 is a schematic diagram illustrating performing a ray tracing simulation to obtain and .
[0017] FIG. 9 is a schematic illustrating an example of determining a re-rendering consistency loss and a reverberation sound loss.
[0018] FIG. 10 is a schematic diagram illustrating an adaptive environment change and parameter expansion.
[0019] FIG. 11 is a schematic diagram illustrating environment change detection.
[0020] FIG. 12 is a schematic diagram illustrating determination of a loss convergence situation.
[0021] FIG. 13 is a schematic diagram illustrating determination of an expected benefit value of each plane.
[0022] FIG. 14 is a schematic diagram illustrating gradient direction difference for different collision points in the same plane.
[0023] FIG. 15 is a schematic diagram illustrating plan expansion.
[0024] FIG. 16 is a schematic diagram illustrating sound source interference and noise interference.
[0025] FIG. 17 is a schematic diagram illustrating an environment change.
[0026] FIG. 18 is a flowchart illustrating a method performed by an electronic apparatus according to another embodiment of the present disclosure.
[0027] FIG. 19 is a schematic diagram of an example scenario to which a method according to an embodiment of the present disclosure may be applied.
[0028] FIG. 20 is a block diagram illustrating an electronic apparatus according to an embodiment of the present disclosure.
[0029] FIG. 21 is a schematic diagram illustrating a structure of an electronic apparatus according to an embodiment of the present disclosure.
[0030] The following description with reference to the accompanying drawings is provided to aid in a thorough understanding of various embodiments of the present disclosure as defined by claims and equivalents thereof. This description includes various specific details to aid in understanding but should only be considered exemplary. Accordingly, those ordinary skills in the art will recognize that various changes and modifications can be made to the various embodiments described herein without departing from the scope and spirit of the present disclosure. In addition, descriptions of well-known features and structures may be omitted for the sake of clarity and brevity.
[0031] The terms and phrases used in the claims and the following description are not limited to dictionary meaning thereof, but are used only by the inventor to enable a clear and consistent understanding of the present disclosure. Accordingly, it should be apparent to those skilled in the art that, the following description of the various embodiments of the present disclosure is provided for an illustrative purpose only and is not intended to a purpose of limiting the present disclosure as defined by the appended claims and equivalents thereof.
[0032] It should be understood that, "a", "an" and "the" in a singular form may also include a plural reference, unless the context clearly indicates otherwise. Thus, for example, a reference to a "part surface" includes a reference to one or more such surfaces. When it refers to one element as being "connected" or "coupled" to another element, the one element may be directly connected or coupled to the other element, or it may refer to a connection relationship between the one element and the other element established through an intermediate element. In addition, "connected" or "coupled" as used herein may include wirelessly connected or wirelessly coupled.
[0033] The term "include" or "may include" refers to the presence of a function, operation, or component of the corresponding disclosure that may be used in the various embodiments of the present disclosure, and does not limit the presence of one or more additional functions, operations, or features. In addition, the terms "include" or "have" may be interpreted to denote certain features, figures, steps, operations, constituent elements, components, or combinations thereof, but should not be interpreted to exclude the possibility of the presence of one or more other features, figures, steps, operations, constituent elements, components, or combinations thereof.
[0034] The term "or" as used in the various embodiments of the present disclosure includes any of the listed terms and all combinations thereof. For example, "A or B" may include A, may include B, or may include both A and B. When describing a plurality of (two or more) items, the plurality of items may refer to one, more, or all of the plurality of items if a relationship among the plurality of items is not explicitly defined. For example, for the description "a parameter A comprises A1, A2, A3", it may be implemented as parameter A comprising A1, A2 or A3, or as parameter A comprising at least two of the three items of the parameter A1, A2, A3.
[0035] All terms (including technical or scientific terms) used in the present disclosure have the same meaning as understood by those skilled in the art to which the present disclosure belongs, unless defined differently. Common terms as defined in dictionaries are interpreted to have a meaning consistent with the context in the relevant technology art and should not be interpreted in an idealized or overly formalistic manner, unless expressly so defined in the present disclosure.
[0036] At least part of the functions in a device or electronic apparatus provided in the embodiments of the present disclosure may be implemented through an AI model, such as, at least one of a plurality of modules of the device or electronic apparatus may be implemented through the AI model. A function associated with AI may be performed through the non-volatile memory, the volatile memory, and the processor.
[0037] The processor may include one or more processors. At this time, the one or more processors may be a general purpose processor, such as a central processing unit (CPU), an application processor (AP), or the like, or may be a graphics-only processing unit such as a graphics processing unit (GPU), a visual processing unit (VPU), and / or an AI-dedicated processor such as a neural processing unit (NPU).
[0038] The one or more processors control processing of input data in accordance with a predefined operating rule or artificial intelligence (AI) model stored in the non-volatile memory and the volatile memory. The predefined operating rule or artificial intelligence model is provided through training or learning.
[0039] Here, being provided through learning means that, by applying a learning algorithm to a plurality of learning data, a predefined operating rule or an AI model of a desired characteristic is made. The learning may be performed in a device or electronic apparatus itself in which AI according to an embodiment is performed, and / or may be implemented through a separate server / system.
[0040] The AI model may consist of a plurality of neural network layers. Each layer has a plurality of weight values, and performs a neural network calculation by calculating between the input data of this layer (such as, a calculation result of the previous layer and / or the input data of the AI model) and the plurality of weight values of the current layer. Examples of neural networks include, but are not limited to, a convolutional neural network (CNN), a deep neural network (DNN), a recurrent neural network (RNN), a restricted Boltzmann Machine (RBM), a deep belief network (DBN), a bidirectional recurrent deep neural network (BRDNN), a generative adversarial networks (GAN), and a deep Q-network.
[0041] The learning algorithm is a method for training a predetermined target device (for example, a robot) using a plurality of learning data to cause, allow, or control the target device to make a determination or prediction. Examples of the learning algorithm include, but are not limited to, supervised learning, unsupervised learning, semi-supervised learning, or reinforcement learning.
[0042] The methods provided in the present disclosure may involve one or more of technical fields such as speech, language, image, video, or data intelligence.
[0043] Alternatively, when involving the field of speech or language, in the method according to the present disclosure executed by electronic apparatus, a speech signal, which is an analog signal, may be received via speech input devices (e.g., a microphone), and the speech part is converted into computer readable text using an automatic speech recognition (ASR) model. The user's intent of utterance may be obtained by interpreting the converted text using a natural language understanding (NLU) model. The ASR model or NLU model may be an artificial intelligence model. The artificial intelligence model may be processed by an artificial intelligence-dedicated processor designed in a hardware structure specified for artificial intelligence model processing. Language understanding is a technique for recognizing and applying / processing human language / text and includes, e.g., natural language processing, machine translation, dialog system, question answering, or speech recognition / synthesis.
[0044] Alternatively, when involving the field of image or video, in the method according to the present disclosure executed by electronic apparatus, output data may be obtained by using image data as input data for an artificial intelligence model. The method of the present disclosure may involve the field of visual understanding in the artificial intelligence technology, and the visual understanding is a technique for recognizing and processing things as does human vision and includes, e.g., object recognition, object tracking, image retrieval, human recognition, scene recognition, 3D reconstruction / localization, or image enhancement.
[0045] Alternatively, when involving the field of data intelligence processing, in the method according to the present disclosure executed by electronic apparatus, in the reasoning or predicting stage, an artificial intelligence model can be used to perform predictions by using real-time input data. Processors of the electronic apparatus may perform a pre-processing operation on the data to convert into a form appropriate for use as an input for the artificial intelligence model. Reasoning and prediction is a technique of logically reasoning and predicting by determining information and includes, e.g., knowledge-based reasoning, optimization prediction, preference-based planning, or recommendation.
[0046] In the present application, the artificial intelligence model may be obtained by training. Here, "obtained by training" means that a predefined operation rule or artificial intelligence model configured to perform a desired feature (or purpose) is obtained by training a basic artificial intelligence model with multiple pieces of training data by a training algorithm. The artificial intelligence model may include a plurality of neural network layers. Each of the plurality of neural network layers includes a plurality of weight values and performs neural network computation by computation between a result of computation by a previous layer and the plurality of weight values.
[0047] Below, the technical solutions of the embodiments of the disclosure and the technical effects produced by the technical solutions of the disclosure will be explained by describing several optional embodiments. It should be noted that, the following embodiments may be referred to, imitated or combined with each other, and the same term, similar features and similar implementation steps in different embodiments will not be described repeatedly.
[0048] As described in the background, the rendering of the virtual audio requires the use of the parameters for reflecting the acoustic characteristics of the space, and if the parameters for reflecting the acoustic characteristics of the space are inaccurate, the audio rendering of the virtual audio will be affected, which ultimately leads to difference in the reverberation between the real sound and the virtual audio, thereby affecting the immersive experience of the user. For example, the audio rendering of the virtual audio may adopt the spatial audio rendering technologies. The spatial audio rendering technologies may have various classifications, among which an audio ray tracing technology of acoustic geometry-based (also known as "ray tracing sound rendering" or "sound ray tracing") is widely used due to its realistic rendering effect, fast computation speed, and adaptability to various spaces. The audio ray tracing technologies are briefly described below for ease of understanding, but it should be understood that spatial audio rendering technologies are not limited to the audio ray tracing technology.
[0049] The audio ray tracing technology assumes that sound energy propagates in the space in the form of rays. The rays start at the sound source and are emitted in all directions. When the rays hit surfaces in a room (e.g., a wall, a floor, and a ceiling), some energy is reflected and other energy is absorbed. It is possible to simulate the propagation progress of the sound in the space and calculate the ray energy at the reception point, by tracing each ray. Sound rays are reflected when they hit a surface. The reflection is a combination of the specular reflection and the diffuse reflection. The relative strength of each reflection is determined by the specular reflection and scattering coefficients of the surface. FIG. 1 is a schematic diagram illustrating an audio ray tracing technology.
[0050] As shown in FIG. 1, in the audio ray tracing technology, the geometry information 20 of the space and the parameters for reflecting the acoustic characteristics of the space have a significant impact on a final rendering result. For example, the parameters for reflecting the acoustic characteristics of the space may be a material absorption coefficient 30 of the reflective surface, but is not limited thereto. For example, the material absorption coefficient 30 may include reflection and scattering coefficients of the reflective surface. The material absorption coefficient 30 is the measure of how much sound is absorbed (rather than reflected) when hitting a surface. The scattering coefficient is a ratio of diffusely reflected sound energy to total reflected energy. The number of material absorption coefficients 30 to be estimated is equal to the number of reflective planes in space, and each reflective surface has its own material absorption coefficient 30.
[0051] If the parameters for reflecting the acoustic characteristics of the space used for the audio rendering of the virtual audio 10 are inaccurate, the audio rendering of the virtual audio 10 will be affected, which ultimately leads to difference in the senses of reverberation between the virtual audio 10 and the real sound, thereby affecting the immersive experience of the user. In this regard, the present disclosure reduces the difference in the senses of the reverberation between the virtual audio 10 and the real sound by improving the accuracy of the obtained parameters for reflecting the acoustic characteristics of the space.
[0052] It is found by the applicant that, in the present related technologies, the difference in the senses of the reverberation between the real sound and the virtual audio 10 is due to the failure to accurately obtain the parameters for reflecting the acoustic characteristics of the space, and the failure to obtain the accurate parameters is mainly due to the existence of at least one of the following problems:
[0053] Problem 1: recorded audio has personalized acoustic characteristics, which may interfere with the estimation of the parameters for reflecting the acoustic characteristics of the space (hereinafter, also referred to as "acoustic characteristic parameters" ). For example, 1) each person speaks with a different resonance cavity, and 2) there may be indeterminate interference during speech, such as wearing a mask. All of these have some reverberation information, however, the related technologies are unable to discard the interference in recording the audio when estimating the acoustic characteristic parameters, which results in the estimated acoustic characteristic parameters actually being a mixture result of spatial and sound source reverberation conditions.
[0054] Problem 2: when environment changes (e.g., a curtain is closed), the number of reflective surfaces in the space where the source is located also changes. Incorrect definition of the number of reflective surfaces means that the acoustic characteristic parameters of the space cannot be expressed correctly and the estimated acoustic characteristic parameters are inaccurate.
[0055] In this regard, the present disclosure obtains more accurate acoustic characteristic parameters by solving Problem 1 and / or Problem 2 described above, and then performs the audio rendering of the virtual audio 10 based on the more accurate acoustic characteristic parameters to obtain the reverberation audio 40 having more consistent sense of the reverberation with the real sound.
[0056] In the following, a method performed by an electronic apparatus according to embodiments of the present disclosure is described.
[0057] FIG. 2 is a flowchart illustrating a method performed by an electronic apparatus according to an embodiment of the present disclosure. FIG. 3 is a diagram of an exemplary architectural illustrating a method performed by an electronic apparatus according to an embodiment of the present disclosure. The method performed by the electronic apparatus according to embodiments of the present disclosure will be described below with reference to FIG. 2 and in conjunction with FIG. 3.
[0058] Referring to FIG. 2, at operation S210, a plurality of first audio signals respectively recorded at a plurality of different locations in a space where a sound source is located are acquired. The first audio signal may refer to an audio signal of a sound source recorded within a physical space. The plurality of first audio signals may refer to a plurality of audio signals of a sound source. Each first audio signal may be acquired at a different location within the physical space, with the sound source being fixed at a position in the physical space. For example, the plurality of first audio signals may be acquired by setting microphones, respectively at the plurality of different locations in the space where the sound source is located to record the sound emitted by a real sound source. The first audio signal may be referred to as a real recorded audio signal.
[0059] At operation S220, each of the plurality of first audio signals is rendered by using a transmission response corresponding to at least one other first audio signal to obtain a rendered audio signal. By operation S220, audio signals that other microphones can record when the first audio signal is used as a sound source at the sound source location may be simulated. In the following, the transmission response is also referred to as an impulse response. The transmission response may represent how an audio signal is affected as the audio signal propagates through a virtual or real environment before reaching a listener. For example, the transmission response may include the cumulative effects of the environment's geometry, materials, and obstacles on the audio signal, including phenomena such as reflection, absorption, scattering and diffraction. The impulse response may describe how an acoustic environment transforms a brief sound (e.g., impulse) emitted from a sound source into a signal received at a position of the listener. The impulse response may capture the cumulative effects of the environment, such as delays, attenuations, and spectral changes.
[0060] If the parameters (e.g., material absorption coefficients 30) for reflecting the acoustic characteristics of the space where the sound source is located, used in the audio rendering process, are accurate, the plurality of rendered audio signals 311 obtained using the above operations may accurately estimate the parameters for reflecting the acoustic characteristics of the space. The material absorption coefficient refers to a value representing the proportion of incident acoustic energy that is absorbed, rather than reflected, by a given material surface.
[0061] Hereinafter, the operation of rendering, for the each first audio signal among the plurality of first audio signals 301, using the transmission response corresponding to the at least one other first audio signal to obtain the rendered audio signal is also referred to as a "re-rendering reverberation consistency 310" operation. As shown in FIG. 3, the plurality of rendered audio signals 311 may be obtained by performing the "re-rendering reverberation consistency 310" operation based on the real recorded audio signals (e.g., the first audio signal). By estimating the parameters (e.g., material absorption coefficients 30) for reflecting the acoustic characteristics of the space based on such a plurality of rendered audio signals 311, the characteristics of the sound source itself may be avoided from interfering with the estimation of the parameters, so that the estimated parameters relate only to the acoustic field of the space and not to the characteristics of the sound source itself, thereby resulting in the estimation of more accurate parameters.
[0062] Hereinafter, operation S220 is described in detail in connection with examples.
[0063] According to embodiments, operation S220 may include: rendering, for each of the plurality of first audio signals 301, using the transmission response corresponding to the at least one other first audio signal to obtain the rendered audio signal, based on a relationship between the plurality of first audio signals, wherein the relationship is a relationship between the plurality of first audio signals obtained by eliminating dry sound from each of the plurality of first audio signals.
[0064] According to embodiments, the sound source mentioned in operation S210 may be a single sound source, and the plurality of different locations may include a first location and a second location. According to an embodiment, the rendering, for the each first audio signal, using the transmission response corresponding to the at least one other first audio signal to obtain the rendered audio signal, based on the relationship between the plurality of first audio signals includes: rendering, for the first audio signal recorded at the first location, using the transmission response corresponding to the first audio signal recorded at the second location to obtain a first rendered audio signal; rendering, for the first audio signal recorded at the second location, using the transmission response corresponding to the first audio signal recorded at the first location to obtain a second rendered audio signal. Hereinafter, the rendering, for the first audio signal recorded at a certain location, using the transmission response corresponding to the first audio signal recorded at another location may also be referred to as performing virtual path rendering from the sound source to the other location on the first audio signal recorded at the certain location.
[0065] FIG. 4 is a schematic diagram illustrating re-rendering reverberation consistency in the case of a single sound source, according to an embodiment of the present disclosure.
[0066] For example, the electronic apparatus may be a mixed reality (MR) apparatus. The plurality of microphones may be set at different locations on the MR apparatus, for example, three microphones a 403, b 405, and c 407. For the case of the single sound source, two microphones may be selected from the plurality of microphones for audio recording. Assuming that the location of the sound source is P 401, two microphones a 403, b 405 set at different locations are selected. For example, the principle of selecting the microphones may be that the two microphones are as far apart as possible such that the recorded audio information includes difference in the sound propagation paths. For example, a straight line distance may be greater than a threshold distance. The threshold distance may represent a criterion for selecting two microphones. The threshold distance value may be predetermined, for example, 6 cm. Although the two microphones a 403, b 405 are selected in the example of FIG. 4, however, the present disclosure is not limited to this example, and two microphones b 405, c 407 may also be selected, or two microphones a 403, c 405 may also be selected.
[0067] In the present disclosure, an embodiment is described in which two microphones are selected from among three microphones located at different positions, and rendering is performed using a transmission response of a first audio signal recorded at the selected microphones. However, the embodiment is not limited thereto, and the method may include selecting two or more microphones from among three or more microphones located at different positions, and performing rendering using the transmission response of a first audio signal respectively recorded at the selected microphones.
[0068] The audio signal recorded by the microphone a 403 and b 405 is a signal in the time domain, and the audio signal thereof equals to a result of dry sound of the sound source being convoluted with the impulse response from P 401 to a 403 and b 405, wherein the dry sound is the sound signal of the sound source itself without any reverberation information, and the impulse response is a complete acoustic description of how a sound signal propagates from one location to another in the space, with the formula expressed as follows:
[0069]
[0070]
[0071] Wherein dry audio denotes the dry sound, denotes the convolution operation, is the impulse response of the sound signal 409 emitted by the sound source propagating in the space from location P 401 to the location where the microphone a is located (e.g., the transmission response corresponding to the audio signal 411 recorded by the microphone a 403), is the impulse response of the sound signal emitted by the sound source propagating in the space from the location P 401 to the location where the microphone b 405 is located (e.g., the transmission response corresponding to the audio signal 413 recorded by the microphone b 405), is the audio signal recorded by the microphone a 403, and is the audio signal recorded by the microphone b 405.
[0072] and in the time domain are Fourier transformed to the frequency domain, and the frequency domain signals respectively corresponding to and may be obtained:
[0073]
[0074] Wherein is the spectrum of the dry sound, and are the spectrums of the impulse response and respectively, is the frequency domain signal corresponding to , and is the frequency domain signal corresponding to .
[0075] The same dry sound contained in the audio signals recorded by the microphones a 403 and b 405 may be removed by dividing to :
[0076]
[0077] It may be obtained by cross-multiplication:
[0078] Equation 1
[0079] The above equation is the relationship between and obtained by removing the dry sound from each of the plurality of first audio signals, based on which it may be rendered for 411 using the transmission response corresponding to the audio signal 413 recorded by the microphone b 405 to obtain the first rendered audio signal 421, and it may be rendered for 413 using the transmission response corresponding to the audio signal 411 recorded by the microphone a 403 to obtain the second rendered audio signal 423, that is, the virtual path rendering from P 401 to b 405 is performed on the audio signal 411 recorded by the microphone a 403, and the virtual path rendering from P 401 to a 403 is performed on the audio signal 413 recorded by the microphone b 405, wherein is the first rendered audio signal 421 (represented as in FIG. 4) obtained by performing the virtual path rendering from P 401 to b 405 on the audio signal 411 recorded by the microphone a 403, and is the second rendered audio signal 423 (represented as in FIG. 4) obtained by performing the virtual path rendering from P 401 to a 403 on the audio signal 413 recorded by the microphone b 405. Both rendered audio signals correspond to the same multiple audio propagation paths, e.g., both including two audio propagation paths from P 401 to a 403 and from P 401 to b 405, for example, 421 corresponds to the real audio propagation path from P 401 to a 403 and the virtual audio propagation path from P 401 to b 405, and 423 corresponds to the real audio propagation path from P 401 to b 405 and the virtual audio propagation path from P 401 to a 403. If the parameters (e.g., material absorption coefficients 30) for reflecting the acoustic characteristics of the space where the sound source is located, used in the audio rendering process, are accurate, then the two rendered audio signals should remain consistent, theoretically. Thus, the parameters for reflecting the acoustic characteristics of the space may be accurately estimated based on the obtained plurality of rendered audio signals.
[0080] Alternatively, according to embodiments, the sound source referred to in operation S210 may be a plurality of sound sources, rendering, for each of the plurality of first audio signals 301, using the transmission response corresponding to the at least one other first audio signal to obtain the rendered audio signal, based on the relationship between the plurality of first audio signals may include: rendering, for the each first audio signal, using the transmission response corresponding to the at least one other first audio signal to obtain the rendered audio signal, based on a linear combination relationship between the plurality of first audio signals. The linear combination relationship is a linear combination relationship between the plurality of first audio signals obtained by algorithmically eliminating the dry sound from the sound sources.
[0081] For example, the sound source may include a first sound source and a second sound source, and the plurality of different locations may include a first location, a second location and a third location. In this case, the rendering, for the each first audio signal, using the transmission response corresponding to the at least one other first audio signal to obtain the rendered audio signal, based on the linear combination relationship between the plurality of first audio signals may include: rendering, for the first audio signal recorded at the first location, using transmission responses respectively corresponding to the first audio signal from the first sound source and the first audio signal from the second sound source recorded respectively at the second location and the third location, to obtain a first rendered audio signal and a second rendered audio signal; rendering, for the first audio signal recorded at the second location, using the transmission responses respectively corresponding to the first audio signal from the first sound source and the first audio signal from the second sound source recorded respectively at the first location and the third location, to obtain a third rendered audio signal and a fourth rendered audio signal; rendering, for the first audio signal recorded at the third location, using the transmission responses respectively corresponding to the first audio signal from the first sound source and the first audio signal from the second sound source recorded respectively at the first location and the second location, to obtain a fifth rendered audio signal and a sixth rendered audio signal. For example, the first rendered audio signal is obtained by rendering, for the first audio signal recorded at the first location, using the transmission response corresponding to the first audio signal from the first sound source recorded at the second location and the transmission response corresponding to the first audio signal from the second sound source recorded at the third location, and furthermore, the second rendered audio signal is obtained by rendering, for the first audio signal recorded at the first location, using the transmission response corresponding to the first audio signal from the first sound source recorded at the third location and the transmission response corresponding to the first audio signal from the second sound source recorded at the second location, that is, the virtual path rendering from the first sound source to the second location and from the second sound source to the third location is performed on the first audio signal recorded at the first location to obtain the first rendered audio signal, and the virtual path rendering from the first sound source to the third location and from the second sound source to the second location is performed on the first audio signal recorded at the first location to obtain the second rendered audio signal.
[0082] For example, the third rendered audio signal is obtained by rendering, for the first audio signal recorded at the second location, using the transmission response corresponding to the first audio signal from the first sound source recorded at the third location and the transmission response corresponding to the first audio signal from the second sound source recorded at the first location, and the fourth rendered audio signal is obtained by rendering, for the first audio signal recorded at the second location, using the transmission response corresponding to the first audio signal from the first sound source recorded at the first location and the transmission response corresponding to the first audio signal from the second sound source recorded at the third location. That is, the virtual path rendering from the first sound source to the third location and from the second sound source to the first location is performed on the first audio signal recorded at the second location to obtain the third rendered audio signal, and the virtual path rendering from the first sound source to the first location and from the second sound source to the third location is performed on the first real audio signal recorded at the second location to obtain the fourth rendered audio signal.
[0083] For example, the fifth rendered audio signal is obtained by rendering, for the first audio signal recorded at the third location, using the transmission response corresponding to the first audio signal from the first sound source recorded at the first location and the transmission response corresponding to the first audio signal from the second sound source recorded at the second location, and the sixth rendered audio signal is obtained by rendering, for the first audio signal recorded at the third location, using the transmission response corresponding to the first audio signal from the first sound source recorded at the second location and the transmission response corresponding to the first audio signal from the second sound source recorded at the first location. That is, the virtual path rendering from the first sound source to the first location and from the second sound source to the second location is performed on the first audio signal recorded at the third location to obtain the fifth rendered audio signal, and the virtual path rendering from the first sound source to the second location and from the second sound source to the first location is performed on the first audio signal recorded at the third location to obtain the sixth rendered audio signal.
[0084] FIG. 5 is a schematic diagram illustrating re-rendering reverberation consistency in the case of a plurality of sound sources according to an embodiment of the present disclosure.
[0085] As shown in FIG. 5, it is assumed that there are two sound sources in the space, and the locations of the two sound sources are P 501 and Q 503. In this case, for example, three microphones a 505, b 507, and c 509 may be set on the electronic apparatus to record the sounds emitted by the two sound sources, and the real audio signals (e.g., the first audio signals) recorded by the three microphones may be represented as follows:
[0086]
[0087] wherein denotes the dry sound at location P 501, denotes the dry sound at location Q 503, denotes the convolution operation, is the impulse response of ropagating in the space from the location P 501 to the location where the microphone a 505 is located (also referred to as the transmission response corresponding to the audio signal from the sound source at the location P 501 recorded by the microphone a 505), is the impulse response of propagating in the space from the location Q 503 to the location where the microphone a 505 is located (also referred to as the transmission response corresponding to the audio signal from the sound source at the location Q 503 recorded by the microphone a 505), is the impulse response of propagating in the space from the location P 501 to the location where the microphone b 507 is located (the transmission response corresponding to the audio signal from the sound source at the location P 501 recorded by the microphone b 507), is the impulse response of propagating in the space from the location Q 503 to the location where the microphone b 507 is located (also referred to as the transmission response corresponding to the audio signal from the sound source at location Q 503 recorded by the microphone b 507), is the impulse response of propagating in the space from the location P 501 to the location where the microphone c 509 is located (also referred to as the transmission response corresponding to the audio signal from the sound source at location P 501 recorded by the microphone c 509), is the impulse response of propagating in the space from the location Q 503 to the location where the microphone c 509 is located (also referred to as the transmission response corresponding to the audio signal from the sound source at location Q 503 recorded by the microphone c 509), is the audio signal recorded by the microphone a 505, is the audio signal recorded by the microphone b 507, and is the audio signal recorded by the microphone c 509.
[0088] The above signals , and are transformed to the frequency domain, respectively, which may obtain:
[0089]
[0090] Based on the existence of the relationship between the audio signals recorded by the three microphones, it may be deduced that , and have the following linear combination relationship:
[0091] Equation 2
[0092] That is:
[0093]
[0094] The rendered audio signals may be obtained by rendering, for the audio signal recorded by each microphone, using the transmission response corresponding to at least one other audio signal based on the above linear combination relationship, e.g., the plurality of rendered audio signals may be obtained by performing a plurality of times of different virtual path rendering, respectively, on the real audio signals recorded by the three microphones.
[0095] Specifically, in the above equation of the linear combination relationship, is the first rendered audio signal 511 obtained by rendering, for the audio signal recorded by the microphone a 505, using the transmission response corresponding to the audio signal from the sound source at the location P 501 recorded by the microphone b 507 and the transmission response corresponding to the audio signal from the sound source at the location Q 503 recorded by the microphone c 509, that is the first rendered audio signal 511 obtained by performing the virtual path rendering from the location P 501 to the location where the microphone b 507 is located and from the location Q 503 to the location where the microphone c 509 is located on the audio signal recorded by the microphone a (corresponding to the first term on the left side of the equation in FIG. 5).
[0096] Similarly, is the third rendered audio signal 513 obtained by performing the virtual path rendering from the location P 501 to the location where the microphone c 509 is located and from the location Q 503 to the location where the microphone a 505 is located on the audio signal recorded by the microphone b 507 (corresponding to the second term on the left side of the equation in FIG. 5), is the fifth rendered audio signal 515 obtained by performing the virtual path rendering from the location P 501 to the location where the microphone a 505 is located and from the location Q 503 to the location where the microphone b 507 is located on the real audio signal recorded by the microphone c 509 (corresponding to the third term on the left side of the equation in FIG. 5); furthermore, in the above equation, is the second rendered audio signal 512 obtained by performing the virtual path rendering from the location P 501 to the location where the microphone c 509 is located and from the location Q 503 to the location where the microphone b 507 is located on the real audio signal recorded by the microphone a 505 (corresponding to the first term on the right side of the equation in FIG. 5), is the fourth rendered audio signal 514 obtained by performing the virtual path rendering from the location P 501 to the location where the microphone a 505 is located and from the location Q 503 to the location where the microphone c 509 is located on the real audio signal recorded by the microphone c 509 (corresponding to the second term on the right side of the equation in FIG. 5), is the sixth rendered audio signal 516 obtained by performing the virtual path rendering from the location P 501 to the location where the microphone b 507 is located and from the location Q to the location where the microphone a 505 is located on the real audio signal recorded by the microphone c 509 (corresponding to the third term on the right side of the equation in FIG. 5).
[0097] As shown in FIG. 5, the audio signal recorded by each microphone is rendered twice, which results in the plurality of rendered audio signals, and the audio propagation paths respectively included in two rendered audio signal combinations of the plurality of rendered audio signals (the left side of the equation is one rendered audio signal combination, and the right side of the equation is one rendered audio signal combination) are consistent. In other words, each of the rendered audio signal combinations may include the rendered audio signals, which are obtained by rendering for the audio signals recorded at different locations, using the transmission response corresponding to the first audio signal from the sound source recorded at least one location, and in the at least two rendered audio signal combinations, the transmission responses, which are used for rendering for the audio signal recorded at the same location, may be different, and these transmission responses may collectively constitute the transmission responses corresponding to the first audio signals recorded at other locations except for that same location.
[0098] Specifically, as shown in FIG. 5, the rendered audio signal combination on the left side of the equation (including the first rendered audio signal 511, the third rendered audio signal 513, and the fifth rendered audio signal 515) and the rendered audio signal combination on the right side of the equation (including the second rendered audio signal 512, the fourth rendered audio signal 514, and the sixth rendered audio signal 516) both include a real audio propagation path from P 501 to a 505, a real audio propagation path from Q 503 to a 505, a real audio propagation path from P 501 to b 507, a real audio propagation path from Q 503 to b 507, a real audio propagation path from P 501 to c 509, a real audio propagation path from Q 503 to c 509, two virtual audio propagation paths from P 501 to a 505, two virtual audio propagation paths from Q 503 to a 505, two virtual audio propagation paths from P 501 to b 507, two virtual audio propagation paths from Q 503 to b 507, two virtual audio propagation paths from P 501 to c 509, and two virtual audio propagation paths from P 501 to c 509. If the parameters (e.g., material absorption coefficients 30) for reflecting the acoustic characteristics of the space where the sound source is located, used in the audio rendering process, are accurate, then theoretically these two rendered audio signal combinations should be consistent, e.g., the difference between them should be small. Thus, the parameters for reflecting the acoustic characteristics of the space may be accurately estimated based on the obtained plurality of rendered audio signals, for example, the parameter for reflecting the acoustic characteristics of the space may be accurately estimated based on the difference between the rendered audio signal combinations in the plurality of rendered audio signals.
[0099] Although the obtaining of the rendered audio signals is described above in the case of two sound sources and three microphone locations, however, the number of the plurality of sound sources is not limited to two, and the locations of the plurality of microphones are not limited to three, and the rendered audio signal combinations are not limited to the combinations referred to in the above examples. According to embodiments, in the case of the plurality of sound sources and the plurality of microphones, after obtaining the plurality of first audio signals recorded at the plurality of microphone locations, at least two rendered audio signal combinations may be obtained based on the relationship between the plurality of first audio signals, wherein each of the rendered audio signal combinations includes the rendered audio signals, which are obtained by rendering for the audio signals recorded at different locations, using the transmission response corresponding to the first audio signal from the sound source recorded at least one location, and in the at least two rendered audio signal combinations, the transmission responses, which are used for rendering for the audio signal recorded at the same location, are different, and these transmission responses collectively constitute the transmission responses corresponding to the first audio signals recorded at other locations except for that same location.
[0100] Alternatively, the sound source may be a single sound source and the first audio signals recorded at the plurality of different locations all comprise noise signals, and wherein difference between the noise signals is less than a threshold value. In this case, the relationship may be a relationship between the plurality of first audio signals obtained by eliminating dry sound from the sound source and eliminating the noise signals. Since the difference in noise signal interference received at a plurality of different recording locations is not significant (may be considered the same), if a relationship between the plurality of first audio signals is obtained by performing an operation on the plurality of first audio signals to eliminate the dry sound from the sound source and to eliminate the noise signals, then, based on the relationship, the rendered audio signals may be obtained by rendering, for each first audio signal, using the transmission response corresponding to at least one other first audio signal, such that not only the influence of the characteristics of the sound source itself on the subsequent parameter estimation may be eliminated, but also the interference of the noise may be eliminated. For example, alternatively, the sound source may be the single sound source, and the plurality of different locations includes a first location, a second location and a third location, in which case, the rendering, for the each first audio signal, using the transmission response corresponding to the at least one other first audio signal to obtain the rendered audio signal, based on the relationship between the plurality of first audio signals may include: rendering, for the first audio signal recorded at the first location, using the transmission responses corresponding to the first audio signals recorded respectively at the second location and the third location, to obtain a first rendered audio signal and a second rendered audio signal; rendering, for the first audio signal recorded at the second location, using the transmission responses corresponding to the first audio signals recorded respectively at the third location and the first location, to obtain a third rendered audio signal and a fourth rendered audio signal; rendering, for the first audio signal recorded at the third location, using the transmission responses corresponding to the first audio signals recorded respectively at the first location and the second location, to obtain a fifth rendered audio signal and a sixth rendered audio signal. For example, the first rendered audio signal is obtained by rendering, for the first audio signal recorded at the first location, using the transmission response corresponding to the first audio signal recorded at the second location, and the second rendered audio signal is obtained by rendering, for the first audio signal recorded at the first location, using the transmission response corresponding to the first audio signal recorded at the third location; the third rendered audio signal is obtained by rendering, for the first audio signal recorded at the second location, using the transmission response corresponding to the first audio signal recorded at the third location, the fourth rendered audio signal is obtained by rendering, for the first audio signal recorded at the second location, using the transmission response corresponding to the first audio signal recorded in the first location; the fifth rendered audio signal is obtained by rendering, for the first audio signal recorded in the third location, using the transmission response corresponding to the first audio signal recorded in the first location, and the sixth rendered audio signal is obtained by rendering, for the first audio signal recorded in the third location, using the transmission response corresponding to the first audio signal recorded at the second location. That is, the virtual path rendering from the single sound source to the second location is performed on the first audio signal recorded at the first location to obtain the first rendered audio signal, the virtual path rendering from the single sound source to the third location is performed on the first audio signal recorded at the first location to obtain the second rendered audio signal; the virtual path rendering from the single sound source to the third location is performed on the first audio signal recorded at the second location to obtain the third rendered audio signal; the virtual path rendering from the single sound source to the first location is performed on the first audio signal recorded at the second location to obtain the fourth rendered audio signal; the virtual path rendering from the single sound source to the first location is performed on the first audio signal recorded at the third location to obtain the fifth rendered audio signal, and the virtual path rendering from the single sound source to the second location is performed on the first audio signal recorded at the third location to obtain the sixth rendered audio signal.
[0101] FIG. 6 is a schematic diagram illustrating re-rendering reverberation consistency in the case of a single sound source and a presence of small difference in noise, according to an embodiment of the present disclosure.
[0102] As shown in FIG. 6, it is assumed that the location of the single sound source is P 601, when there is noise from a more distant sound source or multiple dispersed sources in the sound environment, such as the sound of a car horn on the road outside a conference room or the sound of multiple audiences in an auditorium. If the difference in the sounds reaching the plurality of microphones is small, it may be assumed that the noise signal interferences received by the plurality of microphones are the same. In this case, for example, three microphones a 603, b 605 and c 607 may be set for audio recording in the space where the sound source is located. The audio signals recorded by the three microphones are as follows:
[0103]
[0104] Wherein denotes the dry sound, denotes the convolution operation, is the impulse response of the sound signal emitted by the sound source propagating in the space from the location P 601 to the location where the microphone a 603 is located (also referred to as the transmission response corresponding to the audio signal recorded by the microphone a 603), is the impulse response of the sound signal emitted by the sound source propagating in the space from the location P 601 to the location where the microphone b 605 is located (also referred to as the transmission response corresponding to the audio signal recorded by the microphone b 605), is the impulse response of the sound signal emitted by the sound source propagating in the space from the location P 601 to the location where the microphone c 607 is located (also referred to as the transmission response corresponding to the audio signal recorded by the microphone c 607), is the noise signal, is the audio signal recorded by the microphone a 603, is the audio signal recorded by the microphone b 605, and is the audio signal recorded by the microphone c 607.
[0105] These are transformed to the frequency domain, respectively, which may obtain:
[0106]
[0107] Wherein is the spectrum of the dry sound, are the spectrums of the impulse responses respectively, is the frequency domain signal corresponding to is the frequency domain signal corresponding to is the frequency domain signal corresponding to and n is the frequency domain signal corresponding to the noise signal .
[0108] In order to eliminate the dry sound and the noise signal from the sound source, the following operation may be performed:
[0109]
[0110] The following relationship may be obtained by further cross-multiplication and simplification:
[0111] Equation 3
[0112] The above equation is the relationship between obtained by eliminating the dry sound and eliminating the noise signals. Based on the above relationship, the plurality of rendered audio signals may be obtained by rendering, for the audio signal recorded by each microphone, using the transmission responses corresponding to the audio signals recorded at the other microphones (e.g., different virtual path rendering is performed respectively on the real audio signals recorded by the three microphones, as described above). Specifically, in the above equation, denotes the first rendered audio signal 611 obtained by performing the virtual path rendering from the location P 601 to the location where the microphone b 605 is located on the audio signal recorded by the microphone a 603 (corresponding to the first term on the left side of the equation in FIG. 6), denotes the second rendered audio signal 612 obtained by performing the virtual path rendering from the location P 601 to the location where the microphone c 607 is located on the audio signal recorded by the microphone b 605 (corresponding to the second term on the left side of the equation in FIG. 6), is the third rendered audio signal 613 obtained by performing the virtual path rendering from the location P 601 to the location where the microphone a 603 is located on the audio signal recorded by the microphone c 607 (corresponding to the third term on the left side of the equation in FIG. 6); and furthermore, denotes the fourth rendered audio signal 614 obtained by performing the virtual path rendering from the location P 601 to the location where the microphone c 607 is located on the audio signal recorded by the microphone a 603 (corresponding to the first term on the right side of the equation in FIG. 6), denotes the fifth rendered audio signal 615 obtained by performing the virtual path rendering from the location P 601 to the location where the microphone a 603 is located on the audio signal recorded by the microphone b 605 (corresponding to the second term on the right side of the equation in FIG. 6), and is the sixth rendered audio signal 616 obtained by performing the virtual path rendering from the location P 601 to the location where the microphone b 605 is located on the audio signal recorded by the microphone c 607.
[0113] As shown in FIG. 6, the real audio signal recorded by each microphone is rendered twice, which results in the plurality of rendered audio signals, and the audio propagation paths respectively included in two rendered audio signal combinations of the plurality of rendered audio signals (the left side of the equation is one rendered audio signal combination, and the right side of the equation is one rendered audio signal combination) are consistent. In other words, each of the rendered audio signal combinations may include the rendered audio signals, which are obtained by rendering for the audio signals recorded at different locations, using the transmission response corresponding to the first audio signal from the sound source recorded at least one location, and in at least two rendered audio signal combinations, the transmission responses, which are used for rendering for the audio signal recorded at the same location, may be different, and these transmission responses may collectively constitute the transmission responses corresponding to the first audio signals recorded at other locations except for that same location. Specifically, as shown in FIG. 6, the rendered audio signal combination on the left side of the equation (including the first rendered audio signal 611, the second rendered audio signal 612, and the third rendered audio signal 613) and the rendered audio signal combination on the right side of the equation (including the fourth rendered audio signal 614, the fifth rendered audio signal 615, and the sixth rendered audio signal 616) both include a real audio propagation path from P 601 to a 603, a real audio propagation path from P 601 to c 607, a real audio propagation path from P 601 to b 605, a virtual audio propagation path from P 601 to a 603, a virtual audio propagation path from P 601 to b 605, and a virtual audio propagation path from P 601 to c 607. If the parameters (e.g., material absorption coefficients 30) for reflecting the acoustic characteristics of the space where the sound source is located, used in the audio rendering process, are accurate, then theoretically these two rendered audio signal combinations should be consistent, e.g., the difference between them should be small. Thus, the parameters for reflecting the acoustic characteristics of the space may be accurately estimated based on the obtained plurality of rendered audio signals, e.g., the parameters for reflecting the acoustic characteristics of the space may be accurately estimated based on the difference between the rendered audio signal combinations in the plurality of rendered audio signals.
[0114] It is to be noted that when there is large difference in the interference of the noise signals received by the plurality of microphones, it is necessary to use the rendering method shown in FIG. 5 in the case of the plurality of sound sources. For example, if the interference difference of noise signals received by a plurality of microphones exceeds a predetermined value, the electronic device may determine that the sound source generating the noise is considered to be an independent sound source. The predetermined value may be set in advance. According to an embodiment, instead of directly ignoring the difference in the noise between the plurality of sound sources, the sound sources generating the noise are considered to be independent sound sources.
[0115] After obtaining the plurality of rendered audio signals by operation S220, at operation S230, parameters for reflecting acoustic characteristics of the space are estimated based on the plurality of rendered audio signals. According to embodiments, at operation S230 may include: estimating the parameters based on difference between the plurality of rendered audio signals or difference between rendered audio signal combinations in the plurality of rendered audio signals. For example, assuming an example as illustrated in FIG. 4 above, after obtaining the first rendered audio signal 421 and the second rendered audio signal 423, the parameters for reflecting the acoustic characteristics of the space may be estimated based on the difference between the first rendered audio signal 421 and the second rendered audio signal 423. According to embodiments, first, a first loss of a model for estimating the parameters may be determined based on the difference between the plurality of rendered audio signals, and then, the parameters may be estimated based on the first loss. For example, a loss function of the model may be constructed based on the difference between the plurality of rendered audio signals, the first loss may be calculated, and the parameters may be estimated by minimizing the first loss. For example, assuming that the parameter is P, the constructed loss function may be: Taking the case of the single sound source shown in FIG. 4 as an example, then , wherein is equal to the square of the result of subtraction between the term on the left side of equation 1 and the term on the right side of equation 1. It should be noted that, however, the loss function is constructed as long as it reflects the difference between the plurality of rendered audio signals.
[0116] Similarly, in the case of a plurality of sound sources and a plurality of locations, the parameters are estimated based on the difference between the rendered audio signal combinations in the plurality of rendered audio signals. According to embodiments, the rendered audio signal combinations of the plurality of rendered audio signals may be at least two rendered audio signal combinations, as described above, in the case of the plurality of sound sources and the plurality of microphones, each of the rendered audio signal combinations includes the rendered audio signals, which are obtained by rendering for the audio signals recorded at different locations, using the transmission response corresponding to the first audio signal from the sound source recorded at least one location, and in the at least two rendered audio signal combinations, the transmission responses, which are used for rendering for the audio signal recorded at the same location, are different, and these transmission responses collectively constitute the transmission responses corresponding to the first audio signals recorded at other locations except for that same location. For example, if it is the case of the plurality of sound sources shown in FIG. 5, may be the loss function constructed to be able to reflect the difference between the terms on both sides of Equation 2. If it is the case of the single sound source and the noise difference is not significant as shown in FIG. 6, may be the loss function constructed to be able to reflect the difference between the terms on both sides of Equation 3. In the following, the first loss which is determined based on the difference between the plurality of rendered audio signals or the difference between the rendered audio signal combinations of the plurality of rendered audio signals is also referred to as the "re-rendering consistency loss", on the basis of which the more accurate parameters may be estimated.
[0117] According to embodiments, the parameters mentioned above may include, but are not limited to, a material absorption coefficient 30 of each plane included in the geometry corresponding to the space, but may be any parameter capable of reflecting the acoustic characteristics of the space.
[0118] Returning back to make the reference to FIG. 2, after estimating the parameters, at operation S240, audio rendering of the virtual audio is performed to obtain a reverberation audio based on the estimated parameters. Since the more accurate parameters can be estimated at operation S230, the reverberation audio having the consistent reverberation sense with the real sound may be obtained based on performing the audio rendering of the virtual audio based on the estimated more accurate parameters, thereby improving the immersive experience of the user.
[0119] Since humans have two ears, the perception of sound consists of common perception and differential perception of two ears, and this structure may be mimicked to further improve the accuracy of the parameter estimation. Two-ear differential perception may be simulated by recording a plurality of real audio signals at multiple different locations and the re-rendering reverberation consistency as mentioned above. Two-ear common perception may be simulated by obtaining real reverberation information corresponding to real audio signals recorded at different locations and by performing a ray-tracing simulation of real audio signals recorded at different locations to obtain their simulated reverberation information.
[0120] To this end, alternatively, the method shown in FIG. 2 further includes: obtaining real reverberation information corresponding to the plurality of first audio signals based on the plurality of first audio signals; performing a ray tracing simulation on the plurality of first audio signals to obtain simulated reverberation information corresponding to the plurality of first audio signals. In this case, operation S230 may include: estimating the parameters based on the plurality of rendered audio signals, the real reverberation information and the simulated reverberation information. Since the parameters are further estimated based on the plurality of rendered audio signals in combination with the real reverberation information and the simulated reverberation information, the accuracy of the estimated parameters may be further improved.
[0121] According to embodiments, the real reverberation information corresponding to the plurality of first audio signals may be obtained based on the plurality of first audio signals using a pre-trained neural network.
[0122] For example, the reverberation information may include first reverberation information and second reverberation information. As an example, the first reverberation information may include reverberation time information, the reverberation time information indicating a time required for a sound pressure level to be lower to a predetermined level after the sound source stops making sound. For example, the reverberation time information may indicate the time required for the sound pressure level to be lower to 60db after the sound source stops making sound, and in the following, such reverberation time information is denoted as T60. As an example, the second reverberation information may include a ratio of the direct sound energy to the early reflected sound energy (direct-to-early reflection ratio, DER). T60 describes the temporal sense, e.g., how long the reverberation time is, but does not describe energy difference in the reverberation. DER (direct-to-early reflection ratio, the ratio of the direct sound energy to the early reflected sound energy) describes how loud the reflected sound is. The energy difference in the reverberation at different frequencies makes a big difference in human perception. For example, a room with high absorption at high frequencies will have a dull reverberation feel (more content at low frequencies).
[0123] Returning to the reference to FIG. 3, for example, after obtaining the real recorded audio signal, the pre-trained neural network may be utilized to calculate its T60 302 and DER 303 information as its corresponding real reverberation information.
[0124] According to embodiments, the performing of the ray tracing simulation 350 on the plurality of first audio signals 301 to obtain the simulated reverberation information corresponding to the plurality of first audio signals 301 may include: initializing the parameters to be estimated based on geometry information of the space; performing, using the initialized parameters, the ray tracing simulation 350 on the plurality of first audio signals 301 to obtain the simulated reverberation information corresponding to the plurality of first audio signals 301.
[0125] Taking the parameter as a material absorption coefficient 30 as an example, initializing of the parameters based on the geometry information 20 of the space may include: determining the number of material absorption coefficients based on geometry corresponding to the space; determining a material type of each plane included in the geometry corresponding to the space, wherein each plane corresponds to one material absorption coefficient; determining an initial value of the material absorption coefficient of each plane based on the determined material type.
[0126] FIG. 7 is a schematic diagram illustrating an initialization of a material absorption coefficient 30. As described in FIG. 7, the geometry information 20 of the space may be, for example, 3D spatial geometry information, which may be represented using mesh facets, and may be obtained by a vision-related sensor or a 3D reconstruction algorithm. The initial meshes of the space 701 may contain thousands of facets. This increases the computational complexity because each facet has its absorption parameter. In addition, the more complex the geometry, the more difficult the ray path calculation becomes. Therefore, mesh simplification 710 may be performed using a quadric error metrics (QEM) algorithm. The complexity of the facet may depend on the computational resources of different apparatuses.
[0127] After simplifying the geometry of the space 703, the surface of an object 705 may be partitioned into blocks using SAM (Segment All Model) based on the color information of the object in the space. The material type of each block 707 may be inferred using a convolutional neural network (CNN). After inferring the material type of each block, the initial material absorption coefficient may be defined based on an absorption table. The absorption table may be determined by common material absorption coefficients recorded according to experiences. Subsequently, neighboring blocks with similar colors and initial material absorption coefficients may be fused based on the simplified meshes. The material absorption coefficient of the fused block is calculated as the average of the material absorption coefficients of the blocks that are fused, wherein the weights are defined based on the area of the block corresponding to each material absorption coefficient. In addition, the larger block containing many colors and very different material absorption coefficients may be divided. For example, clustering may be performed according to point distances and the material absorption coefficients, and each cluster is recorded as one new block. Based on the initialization process described above, the number of final initialized material absorption coefficients and the final initial material absorption coefficient of each plane in the geometry may be determined.
[0128] Returning to FIG. 3, after obtaining the initialized parameter 341, the initialized parameter 341 may be utilized to perform the ray tracing simulation 350 on each recorded real audio signal to obtain the simulated reverberation information corresponding to the real audio signal. For example, using the initialized parameter 341, the ray tracing simulation 350 is performed on each recorded real audio signal to obtain simulated and , the simulated is denoted as 351, and the simulated is denoted as 352.
[0129] As an example, the estimated parameter is the material absorption coefficient. 351 may be obtained by the following process:
[0130] 1. N rays are emitted at the source location in all directions, each ray carrying initial energy ;
[0131] 2. Ray transmissions and reflections are performed in the space using geometry information and the material absorption coefficient ( is an array, each reflecting surface in the space has one ).
[0132] 3. The histogram of the energy received at the listening location (the microphone location) is calculated by tracing the rays. When one ray hits one surface, the energy of the ray at the ray receiving location is recorded. After hitting the surface, the new ray direction is a combination of specular and random reflections. The ray continues to be traced until its propagation time exceeds a duration threshold.
[0133]
[0134] is the energy received at time t, at the listening location, and consists of N rays. Each ray is reflected M times (it is different every time, depending on the path).
[0135] 4. The output is the energy decay curve, e.g., the curve of energy change at the listening location over a period of time. The energy decay follows an exponential curve, which is a linear decay in a dB space. The slope of this decay line in the dB space is the function of , wherein , as shown in FIG. 8.
[0136] The slope is , in the case of obtaining the slope, may be obtained based on .
[0137] As shown in FIG. 8, the direct sound 811 is the sound made by a source 801 that reaches the listener 803 directly without reflection, the early reflected sound 813 (the early reflection) is the sound made by a source 801 that reaches the listener 803 after a single reflection, and late reflected sound (also referred to as the reverberation 815) is the portion after the early reflected sound 813. DER is the ratio of the direct sound energy to the early sound energy. For example, the simulated DER (i.e., ) may be obtained by using the following formula:
[0138]
[0139] Wherein is the initial energy carried by the sound source, is the energy loss coefficient of the direct sound propagating in air, which is related to the propagation distance, and is the sum of the n times of early reflected sound energy.
[0140] Return to referring to FIG. 3, after obtaining the plurality of rendered audio signals 311, the real reverberation information and the simulated reverberation information, the parameters may be estimated based on the plurality of rendered audio signals, the real reverberation information and the simulated reverberation information. According to embodiments, the estimating the parameters based on the plurality of rendered audio signals, the real reverberation information and the simulated reverberation information may include: determining a first loss of a model for estimating the parameters based on difference between the plurality of rendered audio signals or difference between rendered audio signal combinations in the plurality of rendered audio signals; determining a second loss of the model based on difference between the real reverberation information and the simulated reverberation information; estimating the parameters based on the first loss and the second loss.
[0141] According to embodiments, the determining of the second loss based on the difference between the real reverberation information and the simulated reverberation information may include: determining a third loss based on difference between real reverberation time information and simulated reverberation time information corresponding to the plurality of first audio signals; determining a fourth loss based on a difference between a real ratio of direct sound energy to early reflected sound energy and a simulated ratio of direct sound energy to early reflected sound energy corresponding to the plurality of first audio signals; determining the second loss based on the third loss and the fourth loss. In the following, the second loss may also be referred to as the "reverberation sound loss".
[0142] In the following, the process of determining the re-rendering consistency loss and the reverberation sound loss mentioned above is described taking recording the audio signal with two microphones (microphone a and microphone b) in the case of the single sound source as an example.
[0143] FIG. 9 is a schematic diagram illustrating an example of determining a re-rendering consistency loss and a reverberation sound loss.
[0144] As shown in FIG. 9, after acquiring the real audio signal recorded at the microphone a (labeled as a "audio signal a 901") and the real audio signal recorded at the microphone b (labeled as a "audio signal b 902"), the rendered audio signal ab 911 (also represented as ) and the rendered audio signal ba 912 (also represented as ) may be obtained by performing different virtual path rendering on the "audio signal a 901" and the "audio signal b 902". For example, by performing the re-rendering operation based on the frequency domain analysis, e.g., the re-rendering consistency operation referred to above, may be obtained, in which case a re-rendering consistency loss function may be constructed based on the difference between the rendered audio signal ab 911 and the rendered audio signal ba 912 to calculate the re-rendering consistency loss 913, for example, the re-rendering consistency loss = ,e.g., re-rendering consistency loss = , but is not limited thereto.
[0145] As shown in FIG. 9, real T60 302 and DER 303 may be obtained using a pre-trained CNN, and simulated 351 and 352 may be obtained based on the ray tracing model. For example, the ray tracing simulation may be performed on each recorded real audio signal based on the source location and the different recording locations (e.g., the microphone locations) using initialized parameters to obtain the simulated reverberation information corresponding to that real audio signal. Subsequently, the reverberation sound loss function may be constructed based on the difference between T60 and and the difference between DER and to calculate the reverberation sound loss, e.g., reverberation sound loss = , but not limited thereto. The loss function of the final constructed model may be: , wherein is an example of the first loss mentioned above, is an example of the third loss mentioned above, is an example of the fourth loss mentioned above, and the third loss and the fourth loss are not limited thereto. The second loss (e.g., the reverberation sound loss) may be determined based on the third loss and the fourth loss, for example, as shown above, the third loss and the fourth loss may be summed to determine the second loss, e.g., will be determined as the second loss, however, the second loss is not limited thereto.
[0146] After obtaining the re-rendering consistency loss and the reverberation sound loss, the parameters may be estimated based on the re-rendering consistency loss and the reverberation sound loss. For example, by calculating a derivative of the loss function with respect to the parameter , the parameters are optimized along the direction of gradient descending of the derivative. During the determination process, different optimization methods may be adopted, for example, the stochastic gradient descent, the Bayesian optimization algorithm, etc., and the present disclosure is not limited thereto. After estimating the parameters, the audio rendering is performed on the virtual audio to obtain the reverberation audio based on the estimated parameters. Since the parameters are estimated based on the re-rendering consistency loss and the reverberation sound loss, and thus, more accurate parameters may be obtained.
[0147] After completing the parameter optimization, the parameters of adjacent planes included in the geometry corresponding to the space may be detected, and if they are close, it is considered that one plane may be used to represent these planes with similar parameters, and combining them may reduce the number of surfaces, and thereby reducing the number of the parameters, which makes it possible to reduce the complexity of the parameter estimation. To this end, alternatively, the method shown in FIG. 2 may further include: determining, based on the estimated the parameters, adjacent planes in geometry corresponding to the space that are capable of being merged; merging the adjacent planes to reduce the number of the parameters. For example, based on the estimated parameters of each plane included in the geometry corresponding to the space, adjacent planes, the difference between whose parameters is less than a predetermined threshold, included in the geometry corresponding to the space may be determined, and then, the identified adjacent planes are merged to reduce the number of the parameters. According to embodiments, the parameters may include a material absorption coefficient of each plane comprised in geometry corresponding to the space. For example, after completing the parameter optimization, the material absorption coefficients of the adjacent planes may be detected and, if they are close, they are merged to reduce the number of surfaces, thereby reducing the number of material absorption coefficients.
[0148] Through the above operations, the parameters can be accurately estimated, based on which, the present disclosure also proposes to update the parameters describing the acoustic characteristics of the space in real time by continuously acquiring the real sound signals. Alternatively, the method shown in FIG. 2 may further include: acquiring a plurality of second audio signals of the sound source recorded in real time at the plurality of locations after the plurality of first audio signals are recorded; updating the parameters based on the plurality of second audio signals.
[0149] The second audio signal may refer to an audio signal of a sound source recorded after the first audio signal is recorded. The plurality of second audio signals may refer to audio signals of the same sound source. Each of the plurality of second audio signals may be acquired at a different location within the physical space, with the sound source being fixed at a position in the physical space. For example, the plurality of second audio signals may be acquired by recording, in real time, sound emitted from the sound source using microphones respectively located at the plurality of different positions, after the plurality of first audio signals have been recorded.
[0150] For example, the material absorption coefficients are the characteristic of reflective surfaces in a space, and their complexity (e.g., the number of reflective surfaces to be estimated) is related to the geometry information of the space, but is not exactly the same, and when initializing the material absorption coefficients, the number of the material absorption coefficients is initialized by the visual information, but there are some limitations to this method, and it is difficult to model the material absorption coefficients of the appropriate level of complexity only by the visual information. Too many reflective surfaces will affect the speed of the algorithm and cause a waste of computational resources, and too few reflective surfaces will not be able to correctly express the acoustic characteristics of the space and render the sound correctly. In addition, the number of material absorption coefficients changes when the environment changes. For the two reasons described above, the present disclosure proposes to perform environment detection and parameter expansion to rationally vary the number of the parameters (e.g., the material absorption coefficients) in the space to more accurately express the acoustic characteristics of the space.
[0151] Alternatively, the updating of the parameters based on the plurality of second audio signals comprises: detecting, based on the plurality of second audio signals, whether an environmental change occurs in the space; updating the parameters in the case of detecting that the environmental change occurs. For example, in the case of detecting that the environmental change occurs, the number of the parameters may be expanded and the expanded number of the parameters may be estimated.
[0152] As shown in FIG. 10, the environment change detection 330 may be performed based on the real recorded audio signal 1001, and if it is detected that environment change occurs, it is determined that the parameters need to be expanded. After performing the parameter expansion 340, the estimation of the expanded number of the parameters may continue to be performed. For example, returning to reference to FIG. 3, after detecting the environmental change and performing the parameter expansion 340, the ray tracing simulation 350 may be performed to obtain the simulated reverberation information based on the expanded parameters, and the estimation of P may continue to be performed based on the plurality of rendered audio signals 311 obtained by the re-rendering reverberation consistency 310 operation, the real reverberation information obtained by the neural network, and the simulated reverberation information. The estimation of the expanded number of the parameters is performed in the same manner as the estimation of the parameters described above, and will not be repeated here.
[0153] In the following, an example of detecting the environmental change is described. According to embodiments, the detecting, based on the plurality of second audio signals, whether the environmental change occurs in the space comprises: acquiring real reverberation information respectively corresponding to the first audio signal and the second audio signal recorded at the same location; detecting whether the environmental change occurs in the space based on difference between the real reverberation information respectively corresponding to the first audio signal and the second audio signal recorded at the same location. For example, the real reverberation information herein may be the reverberation time information, e.g., T60, or may also be DER, but is not limited thereto. For example, the electronic apparatus may capture real audio in the environment in real time, and for the first recorded audio, it may be used together with the initialized parameters to estimate the correct material absorption coefficient , and meanwhile, during the operation of the apparatus, for each new recording of the audio signal, its T60 may be calculated, and this information may be saved as a T60 log. When the environment changes, the reverberation in the space also changes, so the T60 of the newly recorded audio may be different from the T60 recorded in the log, and when the difference in the T60 fluctuation is greater than threshold time difference, the environment may be considered to be changed. The time difference may represent a criterion for determining the environment change. The time difference value may be predetermined, for example, 0.2s.
[0154] FIG. 11 is a schematic diagram illustrating environment change detection. As shown in FIG. 11, when the curtain is closed, the reverberation is smaller than before, a sharp drop in T60 is detected at this time, and if the drop exceeds a threshold, it may be considered that the environment change occurs. As described above, the number of the parameters may be expanded in the case that the environment change is detected.
[0155] In addition to determining whether the number of the parameters needs to be expanded according to the detection of the environmental change, according to embodiments, it is also possible to determine whether the number of the parameters needs to be expanded according to the convergence situation of the estimated loss of the model. As mentioned above, the first loss may be determined, or the first loss and the second loss may be determined, in which case, the estimated loss of the model used to estimate the parameters may be determined based on the first loss, or may be determined based on the first loss and the second loss. Alternatively, the method described in FIG. 2 may further include: determining a convergence situation of an estimated loss of the model, wherein the estimated loss is determined based on the first loss, or is determined based on the first loss and the second loss; determining whether the number of the parameters needs to be expanded based on the convergence situation; if it is determined that the number of the parameters needs to be expanded, expanding the number of the parameters.
[0156] FIG. 12 is a schematic diagram illustrating determination of a loss convergence situation. During the parameter estimation or optimization process, a loss convergence situation of the model may be determined, and corresponding operations may be performed according to the loss convergence situation. For example, as shown in FIG. 12, if the loss drops sharply, it is considered that the model does not converge, and the parameters continue to be updated with new recorded real audio signals. If the loss curve enters a smoothing phase and the values remain at a low level (e.g., the loss value is less than 1), it is considered that the model converges and no further parameter update is required. If the loss curve enters the smoothing phase, but the value is very high (e.g., the loss value is higher than 1), it is considered that the model does not converge, but the loss can no longer be decreased by updating the parameters, and at this time, it is considered necessary to perform the parameter expansion. The state c indicates the parameter expansion is required. When the state b occurs, it indicates that the current parameters have been updated and a merge check is performed. The merge check here refers to the merging of adjacent planes with similar parameters as mentioned above.
[0157] In the following, the description of how to expand the number of the parameters is continued.
[0158] According to embodiments, the expanding of the number of the parameters may include: determining a plane needs to be split among planes comprised in geometry corresponding to the space; performing a ray tracing simulation on propagation process of the plurality of second audio signals to determine position and gradient direction information of sound rays corresponding to the plurality of second audio signals colliding with the planes comprised in the geometry; splitting the plane needs to be split to expand the number of the planes comprised in the geometry, based on the position and the gradient direction information; determining the expanded number of the parameters based on the expanded number of the planes.
[0159] For example, determining an expected benefit value of each plane based on sound related information and empirical information corresponding to the each plane comprised in the geometry corresponding to the space, wherein the sound related information comprises information relating to colliding of the sound rays with the each plane, and the empirical information comprises historical splitting information and / or historical earning information for the each plane; determining the plane needs to be split based on the expected benefit value of the each plane.
[0160] FIG. 13 is a schematic diagram illustrating determination of an expected benefit value of each plane 1303. As shown in FIG. 13, relevant features for determining the expected benefit value may be determined according to the sound related information and the empirical information, the relevant features may include sound related features 1301 and experience related features 1302. For example, the sound related features 1301 may include (1) gradient direction difference of various collision points on the plane, and (2) difference in reach energy of various collision points on the plane, but are not limited thereto. The experience related features 1302 may include (1) whether each plane has been split and (2) historical earning, but are not limited thereto.
[0161] According to an embodiment, the related features of each plane may be calculated and integrated to determine the expected benefit value of each plane. For example, the benefit value of the related features of each plane 1303 is within (0,1), and the values of the above terms may be summed based on predefined weights to obtain the expected benefit value of each plane. For example, the plane with greater gradient direction difference, greater reach energy, has been split, or has higher historical earning will be assigned the higher benefit value.
[0162] FIG. 14 is a schematic diagram illustrating gradient direction difference for different collision points in the same plane. FIG. 14 illustrates the gradient direction of each ray on the plane, the same material tends to have the same gradient direction, so the gradient direction difference may be used to determine the feature of the plane that needs to be split. If the average of the gradient directions of all the rays is smaller than a threshold gradient direction, the plane may be considered stable and does not need to be split and may be assigned a small benefit value. The threshold gradient direction may represent a criterion for determining whether the plane is stable.
[0163] After determining the plane that needs to be split, the plane that needs to be split may be split to expand the number of planes included in the geometry, and thus to expand the number of the parameters, based on the position and the gradient direction information of colliding of the sound rays with the planes included in the geometry corresponding to the space. FIG. 15 is a schematic diagram illustrating plane expansion. As shown in FIG. 15, given the features of the collision points (e.g., the position and the gradient direction), they are divided into two parts using a line, and then the two sub-planes 1501, 1502 are subdivided until the number of points in the sub-planes is less than a certain threshold or the gradient direction difference of the points in the sub-planes is small.
[0164] Returning back to make the reference to FIG. 2, after estimating the parameters, at operation S240, audio rendering of a virtual audio is performed to obtain a reverberation audio based on the estimated parameters. As shown in FIG. 3, the audio rendering may be performed on the virtual audio based on the estimated parameters and geometry information of the space where the sound source is located using the ray tracing sound rendering technology. Here, the virtual audio refers to not a real recording audio signal in the space, for example, the virtual audio may be an audio signal of a remote participant played by a conferencing apparatus.
[0165] In the ray tracing process for audio rendering, a plurality of rays may be emitted from a sound source in various directions in order to simulate the propagation of acoustic energy. Each ray may represent a portion of the acoustic energy emitted from the sound source, and the ray may travel through a virtual or physical space. During propagation, the ray may interact with one or more surfaces, such as walls, ceilings, or obstacles present in the environment. When the ray contacts a surface, the ray may be reflected according to the angle of incidence and the acoustic reflection characteristics of the surface, and a portion of the acoustic energy of the ray may be absorbed in accordance with the material absorption coefficient of the surface.
[0166] The propagation path of each ray may be tracked as the ray undergoes successive reflections within the environment. When the ray reaches a receiver or a predefined reception region, information associated with the ray, such as the arrival time, the propagation path length, and the remaining acoustic energy, may be recorded. Based on a collection of such ray propagation events, an impulse response may be constructed to represent the acoustic transmission characteristics between the sound source and the receiver. The impulse response may then be used to generate a rendered audio signal by convolving the impulse response with a dry audio signal. Those of skill in the art are aware of the specific operations of the ray tracing audio rendering technology, and it will not be repeated here.
[0167] For example, after estimating the correct material absorption coefficient , may be utilized for the audio rendering to make the virtual audio having a reverberation sense. However, according to the real-life experience, too much reverberation may affect the clarity of the sound. For example, in an empty auditorium, the speech of a person who is far away may not be heard clearly due to the reverberation. In this regard, alternatively, the method shown in FIG. 2 may further include: determining clarity of the reverberation audio and adjusting the parameters based on the clarity; performing, based on the adjusted parameters, the audio rendering of the virtual audio to obtain the reverberation audio. By this operation, a more appropriate reverberation audio may be obtained, which ensures that the reverberation audio obtained by final rendering does not affect the clarity of the sound, thereby improving the listening experience of the user. Returning to refer to FIG. 3, as shown in FIG. 3, after obtaining the reverberation audio, determination of the reverberation clarity 360 may be performed, and the adjustment of the reverberation clarity 370 may be realized by adjusting the parameters according to the determination results.
[0168] For example, the clarity of the reverberation audio may be determined based on the C50 of the reverberation audio. The C50 is an objective measurement standard of speech intelligibility or clarity, which is used to describe a ratio of the direct sound to the reverberation sound before and after 50 milliseconds. The larger the C50 is, the higher the speech intelligibility is. The C50 is in unit of dB, which is related to the clarity. If the reverberation is too large and causes a decrease in the clarity (e.g., C50 is less than the threshold), the amplitude of the reverberation may be controlled. For example, the reverberation may be controlled by adjusting the parameters. For example, the value of the material absorption coefficient may be increased based on the empirical range of the material absorption coefficient. For example, the value of the material absorption coefficient may be increased if the reverberation is large and thus causes the decrease in the clarity.
[0169] The method performed by an electronic apparatus according to embodiments of the present disclosure has been described in connection with examples, above, and the method according to embodiments of the present disclosure allows for more accurate estimation of the parameters for reflecting the acoustic characteristics of the space where the sound source is located, avoids the influence of the characteristics of the sound source on the parameter estimation, and makes it possible to render the virtual audio using the accurately estimated parameters, such that the rendered virtual sound and the real sound have the consistent senses when heard by the user, as if they come from the same sound environment, which improves the immersive sense of the user. For example, as shown in FIG. 16, when a person at the sound source is wearing a mask or there is some noise in the environment, since the method according to embodiments of the present disclosure renders the virtual audio using accurately estimated parameters of the acoustic characteristics, the reverberation audio obtained by performing the audio rendering of the virtual audio is not affected by these interferences, and the rendered virtual sound of other people may be made to sound as if it comes from the same sound environment as the real sound.
[0170] In addition, the method according to embodiments of the present disclosure may adaptively update the parameters according to the changes in the environment, enabling more accurate parameters to be obtained, such that the virtual audio may be rendered using the accurately estimated parameters, which may result in the rendered virtual sound and the real sound having the consistent senses when heard by the user. For example, as shown in FIG. 17, in a remote conference, the voice of a remote participant may be rendered to a fixed location in an actual conference room using the method according to embodiments of the present disclosure, and when the environment changes, for example, when the curtain in the figure is closed, the method according to embodiments of the present disclosure may detect that the change occurs in the environment and update the parameters in this regard, for example, the number of the parameters is expanded, and then new estimation of the parameters are performed.
[0171] It should be noted that, although in the above, in the description of the method of FIG. 2, the operation of updating the parameters according to the environment change is regarded as an operation further included in the method of FIG. 2, however, the operation of updating the parameters according to the environment change may also be applied separately without being regarded as an operation further included in the method of FIG. 2.
[0172] FIG. 18 is a flowchart illustrating a method performed by an electronic apparatus according to another embodiment of the present disclosure.
[0173] Referring to FIG. 18, at operation S1810, an audio signal recorded in real time in a space where a sound source is located is acquired. According to embodiments, the audio signals include, for example, the first audio signal recorded at a first time point and the second audio signal recorded at a second time point after the first time point. According to embodiments, at least one location in the space may be set to record the audio signals, for example, the audio signals may be recorded at the plurality of different locations at each time point, as described above. That is, the audio signal recorded at each time point may be one or more audio signals.
[0174] At operation S1820, whether an environmental change occurs in the space is detected based on the audio signal. According to embodiments, operation S1820 may include: acquiring real reverberation information respectively corresponding to the first audio signal and the second audio signal recorded at the same location; detecting whether the environmental change occurs in the space based on difference between the real reverberation information respectively corresponding to the first audio signal and the second audio signal recorded at the same location. In the above, how to detect whether the environmental change occurs has been described and will not be repeated here.
[0175] At operation S1830, parameters for reflecting acoustic characteristics of the space are updated in the case of detecting that the environmental change occurs. According to embodiments, operation S1830 may include: expanding the number of the parameters and estimating the expanded number of the parameters, in the case of detecting the environmental change occurs. How to expand the number of the parameters in case of detecting the environment change occurs has been described above, therefore, it will not be repeated here, and the relevant content may be found in the corresponding description above.
[0176] At operation S1840, audio rendering of a virtual audio is performed to obtain a reverberation audio based on the updated parameters.
[0177] According to the method shown in FIG. 18, the more accurate parameters may be obtained due to the ability to adaptively update the parameters according to the environment change, such that the virtual audio may be rendered using the accurately estimated parameters, which may result in the rendered virtual sound and the real sound having the consistent sense when heard by the user.
[0178] Alternatively, the method shown in FIG. 18 may also include: determining clarity of the reverberation audio and adjusting the parameters based on the clarity; performing, based on the adjusted parameters, the audio rendering of the virtual audio to obtain the reverberation audio. This operation has been described above and will not be repeated here.
[0179] According to embodiments, the parameters may include a material absorption coefficient of each plane comprised in geometry corresponding to the space, but are not limited thereto.
[0180] The methods performed by an electronic apparatus according to embodiments of the present disclosure have been described, above. The methods according to embodiments of the present disclosure may be applied in various scenarios. The following is a brief description of an example scenario to which the methods according to embodiments of the present disclosure may be applied.
[0181] FIG. 19 is a schematic diagram of an example scenario to which a method according to embodiments of the present disclosure may be applied.
[0182] In the example scenario, the material absorption coefficient may be used, for example, to render the virtual audio when a remote participant is speaking. During the conference, may be continuously updated based on the real audio signal recorded in real time to obtain more accurate . When the virtual audio of the remote participant needs to be rendered, the latest may be used. The reality participant may use MR glasses, and microphones may be set at a plurality of different locations for recording the audio signals on the MR glasses. The application of the method according to embodiments of the present disclosure is illustrated in the following by taking one reality participant using the MR glasses as an example.
[0183] As shown in 1901 of FIG. 19, the remote participant Kevin and the reality participant Gloria / Bob want to have a video conference. As shown in 1902 of FIG. 19, all participants are seated and Gloria puts on and activates the MR glasses. There is no sound in the room. The MR apparatus initializes the number of material absorption coefficients in the room based on the visual camera information and obtains the material absorption coefficient of each plane in the geometry corresponding to the conference room. Next, as shown in 1903 of FIG. 19, Bob begins to speak, and the glasses of Gloria record the real audio of Bob. The recorded real audio of Bob is used to optimize the material absorption coefficients to obtain more accurate . In 1904 of FIG. 19, the determination of loss convergence situation may be performed, if the loss curve enters a smoothing phase but the value is very high, the model does not converge, but the loss can no longer be decreased by the parameter updating, it is necessary to perform the parameter expansion, and the recorded audio is used to optimize the new . Next, in 1905 of FIG. 19, Bob speaks again and the glasses of Gloria record the audio of Bob. The environment change detection may be performed, and if the environment does not change, it indicates that displayed does not need to be updated. As shown in 1906 of FIG. 19, the remote participant Kevin begins to speak. The MR glasses of Gloria perform the audio rendering on the virtual audio from Kevin using the latest . Subsequently, as shown in 1907 of FIG. 19, the curtain in the conference room is closed. Bob speaks again. The MR glasses of Gloria record the audio of Bob. The environment change detection may be performed, and if the change in the environment is detected, needs to be updated, for example, the number of is expanded. As shown in 1908 of FIG. 19, the parameter expansion is performed, and then the newly recorded audio is used to optimize the newly expanded . Subsequently, as shown in 1909 of FIG. 19, Kevin begins to speak again, at this time, the MR glasses of Gloria may render the virtual audio from Kevin using the latest .
[0184] In the above example scenario, since the method according to embodiments of the present disclosure is used, it makes it possible to render the virtual audio from the remote participant Kevin throughout the video conference using the more accurate parameter , thus realizing as if the voice of Kevin is coming from the same space as the voice of the reality participants.
[0185] In the above, the example application scenario of the method according to embodiments of the present disclosure has been briefly described, however, the method according to embodiments of the present disclosure is not limited to be applied in the above example scenario, but may be applied in any scenario where mixed reality is required.
[0186] In the following, the electronic apparatus according to embodiments of the present disclosure is briefly described. FIG. 20 is a block diagram illustrating an electronic apparatus according to an embodiment of the present disclosure. Referring to FIG. 20, the electronic apparatus 2000 may include a memory 2010 and a processor 2020, wherein the processor 2020 is coupled to the memory 2010 and configured to perform any of the methods described above.
[0187] According to the embodiments of the present disclosure, there is also provided a computer program product comprising a computer program / instructions that, when executed by a processor, implements / implement any of the methods described above.
[0188] In embodiments of the present disclosure, there is also provided an electronic apparatus that includes at least one processor, and alternatively, further includes at least one transceiver and / or at least one memory coupled to the at least one processor, wherein, the at least one processor is configured to perform the operations of the method provided in any alternative embodiment of the present disclosure.
[0189] FIG. 21 illustrates a schematic diagram of a structure of an electronic apparatus applicable to an exemplary embodiment of the present application. As shown in FIG. 21, the electronic apparatus 4000 shown in FIG. 21 includes: a processor 4001 and a memory 4003. Wherein the processor 4001 and the memory 4003 are coupled, e.g., through a bus 4002. Alternatively, the electronic apparatus 4000 may further include a transceiver 4004 which may be used for data interaction between the electronic apparatus and other electronic apparatuses, such as transmitting of data and / or receiving of data. It should be noted that, each of the processor 4001, the memory 4003, and the transceiver 4004 is not limited to one in a practice application, and the structure of the electronic apparatus 4000 does not constitute a limitation of the embodiments of the present disclosure. Alternatively, the electronic apparatus may be the first network node, the second network node, or the third network node.
[0190] The processor 4001 may be a Central Processing Unit (CPU), general purpose processor, Digital Signal Processor (DSP), Application Specific Integrated Circuit (ASIC), Field Programmable Gate Array (FPGA) or other programmable logic device, transistor logic device, hardware part, or any combination thereof. It may implement or perform various exemplary logic boxes, modules, and circuits described in conjunction with the disclosed contents of the present disclosure. The processor 4001 may also be a combination that implements computing functions, such as a combination containing one or more microprocessors, a combination of a DSP and a microprocessor, and the like.
[0191] The bus 4002 may include a pathway to transfer information between the above components. The bus 4002 may be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, and the like. The bus 4002 may be classed as an address bus, a data bus, a control bus, and the like. For ease of representation, only one bold line is shown in FIG. 21, but it does not mean that there is only one bus or one type of bus.
[0192] The memory 4003 may be a Read Only Memory (ROM) or other types of static storage apparatuses that can store static information and instructions, a Random Access Memory (RAM) or other types of dynamic storage apparatuses that can store information and instructions, may be an Electrically Erasable Programmable Read Only Memory (EEPROM), Compact Disc Read Only Memory (CD-ROM) or other optical disc storages, an optical disc storage (including a compressed disc, laser disc, optical disc, digital universal disc, Blu-ray disc, etc.), a disk storage medium, other magnetic storage apparatuses, or any other medium that can be used to carry or store computer programs and can be read by a computer, it is not limited herein.
[0193] The memory 4003 is used to store computer programs or executable instructions for performing the embodiments of the present disclosure, and is controlled for execution by the processor 4001. The processor 4001 is used to execute the computer programs or executable instructions stored in the memory 4003 to implement the operations shown in the preceding method of the embodiments.
[0194] An embodiment of the present disclosure provides a computer readable storage medium storing computer programs or instructions, the computer programs or instructions, when being executed by at least one processor may perform or implement the operations in the preceding method of the embodiments and corresponding contents.
[0195] An embodiment of the present disclosure provides a computer program product including computer programs, the computer programs, when being executed by a processor, may implement the operations shown in the preceding method of the embodiments and corresponding contents.
[0196] According to an embodiment, the rendering, for the each first audio signal, using the transmission response corresponding to the at least one other first audio signal to obtain the rendered audio signal comprises: rendering, for the each first audio signal, using the transmission response corresponding to the at least one other first audio signal to obtain the rendered audio signal, based on a relationship between the plurality of first audio signals, wherein the relationship is a relationship between the plurality of first audio signals obtained by eliminating dry sound from the sound source.
[0197] According to an embodiment, the sound source is a single sound source and the plurality of different locations comprises a first location and a second location, wherein the rendering, for the each first audio signal, using the transmission response corresponding to the at least one other first audio signal to obtain the rendered audio signal, based on the relationship between the plurality of first audio signals comprises: rendering, for the first audio signal recorded at the first location, using the transmission response corresponding to the first audio signal recorded at the second location to obtain a first rendered audio signal; rendering, for the first audio signal recorded at the second location, using the transmission response corresponding to the first audio signal recorded at the first location to obtain a second rendered audio signal.
[0198] According to an embodiment, the sound source comprises a plurality of sound sources, wherein rendering, for the each first audio signal, using the transmission response corresponding to the at least one other first audio signal to obtain the rendered audio signal, based on the relationship between the plurality of first audio signals comprises: rendering, for the each first audio signal, using the transmission response corresponding to the at least one other first audio signal to obtain the rendered audio signal, based on a linear combination relationship between the plurality of first audio signals.
[0199] According to an embodiment, the sound source comprises a first sound source and a second sound source, and the plurality of different locations comprises a first location, a second location, and a third location, wherein the rendering, for the each first audio signal, using the transmission response corresponding to the at least one other first audio signal to obtain the rendered audio signal, based on the linear combination relationship between the plurality of first audio signals comprises: rendering, for the first audio signal recorded at the first location, using transmission responses respectively corresponding to the first audio signal from the first sound source and the first audio signal from the second sound source recorded respectively at the second location and the third location, to obtain a first rendered audio signal and a second rendered audio signal; rendering, for the first audio signal recorded at the second location, using the transmission responses respectively corresponding to the first audio signal from the first sound source and the first audio signal from the second sound source recorded respectively at the first location and the third location, to obtain a third rendered audio signal and a fourth rendered audio signal; rendering, for the first audio signal recorded at the third location, using the transmission responses respectively corresponding to the first audio signal from the first sound source and the first audio signal from the second sound source respectively recorded at the first location and the second location, to obtain a fifth rendered audio signal and a sixth rendered audio signal.
[0200] According to an embodiment, the sound source is a single sound source, and the first audio signals recorded at the plurality of different locations all comprise noise signals, and wherein difference between the noise signals is less than a threshold value, wherein the relationship is a relationship between the plurality of first audio signals obtained by eliminating dry sound from the sound source and eliminating the noise signals.
[0201] According to an embodiment, the sound source is the single sound source and the plurality of different locations comprises a first location, a second location and a third location, wherein the rendering, for the each first audio signal, using the transmission response corresponding to the at least one other first audio signal to obtain the rendered audio signal, based on the relationship between the plurality of first audio signals comprises: rendering, for the first audio signal recorded at the first location, using the transmission responses corresponding to the first audio signals recorded respectively at the second location and the third location, to obtain a first rendered audio signal and a second rendered audio signal; rendering, for the first audio signal recorded at the second location, using the transmission responses corresponding to the first audio signals recorded respectively at the third location and the first location, to obtain a third rendered audio signal and a fourth rendered audio signal; rendering, for the first audio signal recorded at the third location, using the transmission responses corresponding to the first audio signals recorded respectively at the first location and the second location, to obtain a fifth rendered audio signal and a sixth rendered audio signal.
[0202] According to an embodiment, the estimating of the parameters for reflecting the acoustic characteristics of the space based on the plurality of rendered audio signals comprises: estimating the parameters based on difference between the plurality of rendered audio signals or difference between rendered audio signal combinations in the plurality of rendered audio signals.
[0203] According to an embodiment, the estimating of the parameters based on the difference between the plurality of rendered audio signals or the difference between the rendered audio signal combinations in the plurality of rendered audio signals comprises: determining a first loss of a model for estimating the parameters based on the difference between the plurality of rendered audio signals or the difference between the rendered audio signal combinations in the plurality of rendered audio signals; estimating the parameters based on the first loss.
[0204] According to an embodiment, the method further comprises: obtaining real reverberation information corresponding to the plurality of first audio signals based on the plurality of first audio signals; performing a ray tracing simulation on the plurality of first audio signals to obtain simulated reverberation information corresponding to the plurality of first audio signals; wherein the estimating of the parameters for reflecting the acoustic characteristics of the space based on the plurality of rendered audio signals comprises: estimating the parameters based on the plurality of rendered audio signals, the real reverberation information and the simulated reverberation information.
[0205] According to an embodiment, the estimating the parameters based on the plurality of rendered audio signals, the real reverberation information and the simulated reverberation information comprises: determining a first loss of a model for estimating the parameters based on difference between the plurality of rendered audio signals or difference between rendered audio signal combinations in the plurality of rendered audio signals; determining a second loss of the model based on difference between the real reverberation information and the simulated reverberation information; estimating the parameters based on the first loss and the second loss.
[0206] According to an embodiment, the reverberation information comprises first reverberation information and second reverberation information, wherein the first reverberation information comprises reverberation time information, the reverberation time information indicating a time required for a sound pressure level to be lower to a predetermined level after the sound source stops making sound; wherein the second reverberation information comprises a ratio of direct sound energy to early reflected sound energy.
[0207] According to an embodiment, the determining of the second loss based on the difference between the real reverberation information and the simulated reverberation information comprises: determining a third loss based on difference between real reverberation time information and simulated reverberation time information corresponding to the plurality of first audio signals; determining a fourth loss based on a difference between a real ratio of direct sound energy to early reflected sound energy and a simulated ratio of direct sound energy to early reflected sound energy corresponding to the plurality of first audio signals; determining the second loss based on the third loss and the fourth loss.
[0208] According to an embodiment, the performing of the ray tracing simulation on the plurality of first audio signals to obtain the simulated reverberation information corresponding to the plurality of first audio signals comprises: initializing the parameters to be estimated based on geometry information of the space; performing, using the initialized parameters, the ray tracing simulation on the plurality of first audio signals to obtain the simulated reverberation information corresponding to the plurality of first audio signals.
[0209] According to an embodiment, the method further comprises: determining, based on the estimated the parameters, adjacent planes in geometry corresponding to the space that are capable of being merged; merging the adjacent planes to reduce the number of the parameters.
[0210] According to an embodiment, the parameters comprise a material absorption coefficient of each plane comprised in geometry corresponding to the space.
[0211] According to an embodiment, the method further comprises: acquiring a plurality of second audio signals of the sound source recorded in real time at the plurality of locations after the plurality of first audio signals are recorded; updating the parameters based on the plurality of second audio signals.
[0212] According to an embodiment, the updating of the parameters based on the plurality of second audio signals comprises: detecting, based on the plurality of second audio signals, whether an environmental change occurs in the space; updating the parameters in the case of detecting that the environmental change occurs.
[0213] According to an embodiment, the detecting, based on the plurality of second audio signals, whether the environmental change occurs in the space comprises: acquiring real reverberation information respectively corresponding to the first audio signal and the second audio signal recorded at the same location; detecting whether the environmental change occurs in the space based on difference between the real reverberation information respectively corresponding to the first audio signal and the second audio signal recorded at the same location.
[0214] According to an embodiment, the updating of the parameters in the case of detecting that the environmental change occurs comprises: expanding the number of the parameters and estimating the expanded number of the parameters, in the case of detecting the environmental change occurs.
[0215] According to an embodiment, the expanding of the number of the parameters, comprises: determining a plane needs to be split among planes comprised in geometry corresponding to the space; performing a ray tracing simulation on propagation process of the plurality of second audio signals to determine position and gradient direction information of sound rays corresponding to the plurality of second audio signals colliding with the planes comprised in the geometry; splitting the plane needs to be split to expand the number of the planes comprised in the geometry, based on the position and the gradient direction information; determining the expanded number of the parameters based on the expanded number of the planes.
[0216] According to an embodiment, the determining of the plane needs to be split among the planes comprised in the geometry corresponding to the space comprises: determining an expected benefit value of each plane based on sound related information and empirical information corresponding to the each plane comprised in the geometry corresponding to the space, wherein the sound related information comprises information relating to colliding of the sound rays with the each plane, and the empirical information comprises historical splitting information and / or historical earning information for the each plane; determining the plane needs to be split based on the expected benefit value of the each plane.
[0217] According to an embodiment, the method further comprises: determining a convergence situation of a estimated loss of the model, wherein the estimated loss is determined based on the first loss, or is determined based on the first loss and the second loss; determining whether the number of the parameters needs to be expanded based on the convergence situation; if it is determined that the number of the parameters needs to be expanded, expanding the number of the parameters.
[0218] According to an embodiment, the method further comprises: determining clarity of the reverberation audio and adjusting the parameters based on the clarity; performing, based on the adjusted parameters, the audio rendering of the virtual audio to obtain the reverberation audio.
[0219] According to an embodiment of the present disclosure, it is provided a method performed by an electronic apparatus, comprising: acquiring an audio signal recorded in real time in a space where a sound source is located; detecting whether an environmental change occurs in the space based on the audio signal; updating parameters for reflecting acoustic characteristics of the space in the case of detecting that the environmental change occurs; and performing, based on the updated parameters, audio rendering of a virtual audio to obtain a reverberation audio.
[0220] According to an embodiment, the audio signal comprises a first audio signal recorded at a first time point and a second audio signal recorded at a second time point after the first time point, wherein detecting whether the environmental change occurs in the space based on the audio signal comprises: acquiring real reverberation information respectively corresponding to the first audio signal and the second audio signal recorded at the same location; detecting whether the environmental change occurs in the space based on difference between the real reverberation information respectively corresponding to the first audio signal and the second audio signal recorded at the same location.
[0221] According to an embodiment, the updating of the parameters for reflecting the acoustic characteristics of the space in the case of detecting that the environmental change occurs, comprises: expanding the number of the parameters and estimating the expanded number of the parameters, in the case of detecting the environmental change occurs.
[0222] According to an embodiment, the expanding of the number of the parameters, comprises: determining a plane needs to be split among planes comprised in geometry corresponding to the space; performing a ray tracing simulation on propagation process of the plurality of second audio signals to determine position and gradient direction information of sound rays corresponding to the plurality of second audio signals colliding with the planes comprised in the geometry; splitting the plane needs to be split to expand the number of the planes comprised in the geometry based on the position and the gradient direction information; determining the expanded number of the parameters based on the expanded number of the planes.
[0223] According to an embodiment, the determining of the plane needs to be split among the planes comprised in the geometry corresponding to the space comprises: determining an expected benefit value of each plane based on sound related information and empirical information corresponding to the each plane comprised in the geometry corresponding to the space, wherein the sound related information comprises information relating to colliding of the sound rays with the each plane, and the empirical information comprises historical splitting information and / or historical earning information for the each plane; determining the plane needs to be split based on the expected benefit value of the each plane.
[0224] According to an embodiment, the method further comprises: determining clarity of the reverberation audio and adjusting the parameters based on the clarity; performing, based on the adjusted parameters, the audio rendering of the virtual audio to obtain the reverberation audio.
[0225] According to an embodiment, the parameters comprise a material absorption coefficient of each plane comprised in geometry corresponding to the space.
[0226] According to an embodiment of the present disclosure, it is provided an electronic apparatus, comprising: a memory; and a processor coupled to the memory and configured to perform the methods described above.
[0227] According to an embodiment of the present disclosure, it is provided a computer-readable storage medium storing instructions which, when executed by at least one processor, cause the at least one processor to perform the methods described above.
[0228] According to an embodiment of the present disclosure, it is provided a computer program product comprising computer program / instructions, the computer program / instructions when executed by a processor implements / implement the methods described above.
[0229] According to a technical solution provided by embodiments of the present disclosure, by acquiring a plurality of first audio signals recorded at a plurality of different locations in a space where a sound source is located, rendering, for first audio signal, using a transmission response corresponding to at least one other first audio signal to obtain a rendered audio signal, and estimating parameters for reflecting acoustic characteristics of the space based on a plurality of rendered audio signals, it is possible to avoid the characteristics of the sound source itself from interfering with the estimation of the parameters such that estimated parameters relate only to the acoustic field of the space and not to the characteristics of the sound source itself, thereby resulting that more accurate parameters are estimated, and in the case that more accurate parameters are estimated, the audio rendering of the virtual audio is performed based on the estimated more accurate parameters, which may obtain the reverberation audio having a consistent sense of the reverberation with the real sound, thereby improving the immersive experience of the user.
[0230] According to the technical solution provided by embodiments of the present disclosure, by acquiring an audio signal recorded in real time in a space where a sound source is located, detecting whether an environmental change occurs in the space based on the audio signal, and updating parameters for reflecting acoustic characteristics of the space in the case of detecting that the environmental change occurs, it is possible to adaptively update the parameters based on the environmental change in a timely manner to obtain more accurate parameters, and in the case of obtaining the more accurate parameters, the audio rendering of a virtual audio is performed based on the updated parameters to obtain a reverberation audio, which may make the obtained reverberation audio capable of adapting to the environmental change, so as to have a consistent sense of the reverberation with the real sound, thereby improving the immersive experience of the user.
[0231] According to an embodiment of the disclosure, the method may include acquiring a plurality of first audio signals respectively recorded at a plurality of different locations in a space where a sound source is located. The method may include rendering, for each of the plurality of first audio signals, using a transmission response corresponding to at least one other first audio signal to obtain a rendered audio signal. The method may include estimating parameters for reflecting acoustic characteristics of the space based on the plurality of rendered audio signals. The method may include performing, based on the estimated parameters, audio rendering of a virtual audio to obtain a reverberation audio.
[0232] According to an embodiment of the disclosure, the sound source may be a single sound source. The plurality of different locations may comprise a first location and a second location. The method may include rendering, for the first audio signal recorded at the first location, using the transmission response corresponding to the first audio signal recorded at the second location to obtain a first rendered audio signal. The method may include rendering, for the first audio signal recorded at the second location, using the transmission response corresponding to the first audio signal recorded at the first location to obtain a second rendered audio signal.
[0233] According to an embodiment of the disclosure, the sound source may comprise a plurality of sound sources. The method may include rendering, for each of the plurality of first audio signals, using the transmission response corresponding to the at least one other first audio signal to obtain the rendered audio signal, based on a linear combination relationship between the plurality of first audio signals.
[0234] According to an embodiment of the disclosure, the sound source may comprise a first sound source and a second sound source. The plurality of different locations may comprise a first location, a second location, and a third location. The method may include rendering, for the first audio signal recorded at the first location, using transmission responses respectively corresponding to the first audio signal from the first sound source and the first audio signal from the second sound source recorded respectively at the second location and the third location, to obtain a first rendered audio signal and a second rendered audio signal. The method may include rendering, for the first audio signal recorded at the second location, using the transmission responses respectively corresponding to the first audio signal from the first sound source and the first audio signal from the second sound source recorded respectively at the first location and the third location, to obtain a third rendered audio signal and a fourth rendered audio signal. The method may include rendering, for the first audio signal recorded at the third location, using the transmission responses respectively corresponding to the first audio signal from the first sound source and the first audio signal from the second sound source respectively recorded at the first location and the second location, to obtain a fifth rendered audio signal and a sixth rendered audio signal.
[0235] According to an embodiment of the disclosure, the method may include estimating the parameters based on difference between the plurality of rendered audio signals or difference between rendered audio signal combinations in the plurality of rendered audio signals.
[0236] According to an embodiment of the disclosure, the method may include obtaining real reverberation information corresponding to the plurality of first audio signals based on the plurality of first audio signals. The method may include performing a ray tracing simulation on the plurality of first audio signals to obtain simulated reverberation information corresponding to the plurality of first audio signals. The method may include estimating the parameters based on the plurality of rendered audio signals, the real reverberation information and the simulated reverberation information.
[0237] According to an embodiment of the disclosure, each of the real reverberation information and the simulated reverberation information may comprise first reverberation information and second reverberation information. The first reverberation information may comprise reverberation time information, the reverberation time information indicating a time required for a sound pressure level to be lower to a predetermined level after the sound source stops making sound. The second reverberation information may comprise a ratio of direct sound energy to early reflected sound energy.
[0238] According to an embodiment of the disclosure, the method may include initializing the parameters to be estimated based on geometry information of the space. The method may include performing, using the initialized parameters, the ray tracing simulation on the plurality of first audio signals to obtain the simulated reverberation information corresponding to the plurality of first audio signals.
[0239] According to an embodiment of the disclosure, the parameters may comprise a material absorption coefficient of each plane comprised in geometry corresponding to the space.
[0240] According to an embodiment of the disclosure, the method may include acquiring a plurality of second audio signals of the sound source recorded in real time at the plurality of locations after the plurality of first audio signals are recorded. The method may include updating the parameters based on the plurality of second audio signals.
[0241] According to an embodiment of the disclosure, the method may include detecting, based on the plurality of second audio signals, whether an environmental change occurs in the space. The method may include updating the parameters based on detecting that the environmental change occurs.
[0242] According to an embodiment of the disclosure, the method may include determining a plane needs to be split among planes comprised in geometry corresponding to the space. The method may include performing a ray tracing simulation on propagation process of the plurality of second audio signals to determine position and gradient direction information of sound rays corresponding to the plurality of second audio signals colliding with the planes comprised in the geometry. The method may include splitting the plane needs to be split to expand the number of the planes comprised in the geometry, based on the position and the gradient direction information. The method may include determining the expanded number of the parameters based on the expanded number of the planes.
[0243] According to an embodiment of the disclosure, the method may include determining an expected benefit value of each plane based on sound related information and empirical information corresponding to the each plane comprised in the geometry corresponding to the space, wherein the sound related information comprises information relating to colliding of the sound rays with the each plane, and the empirical information comprises at least one of historical splitting information or historical earning information for the each plane. The method may include determining the plane needs to be split based on the expected benefit value of the each plane.
[0244] According to an embodiment of the disclosure, an electronic apparatus may be provided. The electronic apparatus may include memory storing one or more instructions. and at least one processor. The instructions when executed by the at least one processor individually or collectively, may cause the electronic apparatus to acquire a plurality of first audio signals respectively recorded at a plurality of different locations in a space where a sound source is located. The instructions when executed by the at least one processor individually or collectively, may cause the electronic apparatus to render, for each of the plurality of first audio signals, using a transmission response corresponding to at least one other first audio signal to obtain a rendered audio signal. The instructions when executed by the at least one processor individually or collectively, may cause the electronic apparatus to estimate parameters for reflecting acoustic characteristics of the space based on the plurality of rendered audio signals. The instructions when executed by the at least one processor individually or collectively, may cause the electronic apparatus to perform, based on the estimated parameters, audio rendering of a virtual audio to obtain a reverberation audio.
[0245] According to an embodiment of the disclosure, a computer-readable medium storing one or more instructions may be provided. The one or more instructions, when executed by at least one processor, may cause the at least one processor of an electronic apparatus to perform operation corresponding to the method.
[0246] The terms "first", "second", "third", "fourth", "1", "2" and the like (if exists) in the specification and claims of the present disclosure and the above drawings are used to distinguish similar objects, and need not be used to describe a specific order or sequence. It should be understood that, data used as such may be interchanged in appropriate situations, so that the embodiments of the present disclosure described here may be implemented in an order other than the illustration or text description.
[0247] It should be understood that, although each operation is indicated by an arrow in the flowcharts of the embodiments of the present disclosure, an implementation order of these operations is not limited to an order indicated by the arrows. Unless explicitly stated herein, in some implementation scenarios of the embodiments of the present disclosure, the implementation operations in the flowcharts may be executed in other orders according to requirements. In addition, some or all of the operations in each flowchart may include a plurality of sub operations or stages, based on an actual implementation scenario. Some or all of these sub operations or stages may be executed at the same time, and each sub operation or stage in these sub operations or stages may also be executed at different times. In scenarios with different execution times, an execution order of these sub operations or stages may be flexibly configured according to a requirement, which is not limited by the embodiment of the present disclosure.
[0248] The above text and accompanying drawings are provided as examples only to assist readers in understanding the present disclosure. They are not intended and should not be interpreted as limiting the scope of the present disclosure in any way. Although certain embodiments and examples have been provided, based on the content disclosed herein, it is apparent to those skilled in the art that, changes can be made to the illustrated embodiments and examples without departing from the scope of the present disclosure, and other similar implementation methods based on the technical concepts of the present disclosure also belongs to a protection scope of the embodiments of the present disclosure.
Claims
1.A method performed by an electronic apparatus comprising:acquiring a plurality of first audio signals 301 respectively recorded at a plurality of different locations in a space where a sound source is located;rendering, for each of the plurality of first audio signals 301, using a transmission response corresponding to at least one other first audio signal to obtain a rendered audio signal 311;estimating parameters for reflecting acoustic characteristics of the space based on the plurality of rendered audio signals 311; andperforming, based on the estimated parameters, audio rendering 100 of a virtual audio 10 to obtain a reverberation audio 40.2.The method of claim 1, wherein the sound source is a single sound source and the plurality of different locations comprises a first location and a second location, wherein the rendering, for each of the plurality of first audio signals 301, using the transmission response corresponding to the at least one other first audio signal to obtain the rendered audio signal 311 comprises:rendering, for the first audio signal 411 recorded at the first location, using the transmission response corresponding to the first audio signal 413 recorded at the second location to obtain a first rendered audio signal 421; andrendering, for the first audio signal 413 recorded at the second location, using the transmission response corresponding to the first audio signal 411 recorded at the first location to obtain a second rendered audio signal 423.3.The method of claim 1, wherein the sound source comprises a plurality of sound sources, wherein rendering, for each of the plurality of first audio signals 301, using the transmission response corresponding to the at least one other first audio signal to obtain the rendered audio signal comprises:rendering, for each of the plurality of first audio signals 301, using the transmission response corresponding to the at least one other first audio signal to obtain the rendered audio signal, based on a linear combination relationship between the plurality of first audio signals.4.The method of claim 3, wherein the sound source comprises a first sound source and a second sound source, and the plurality of different locations comprises a first location, a second location, and a third location, andwherein the rendering, for each of the plurality of first audio signals 301, using the transmission response corresponding to the at least one other first audio signal to obtain the rendered audio signal comprises:rendering, for the first audio signal recorded at the first location 505, using transmission responses respectively corresponding to the first audio signal from the first sound source and the first audio signal from the second sound source recorded respectively at the second location 507 and the third location 509, to obtain a first rendered audio signal 511 and a second rendered audio signal 512;rendering, for the first audio signal recorded at the second location 507, using the transmission responses respectively corresponding to the first audio signal from the first sound source and the first audio signal from the second sound source recorded respectively at the first location 505 and the third location 509, to obtain a third rendered audio signal 513 and a fourth rendered audio signal 514; andrendering, for the first audio signal recorded at the third location 509, using the transmission responses respectively corresponding to the first audio signal from the first sound source and the first audio signal from the second sound source respectively recorded at the first location 505 and the second location 507, to obtain a fifth rendered audio signal 515 and a sixth rendered audio signal 516.5.The method any one of claims 1 to 4, wherein the estimating of the parameters for reflecting the acoustic characteristics of the space based on the plurality of rendered audio signals comprises:estimating the parameters based on difference between the plurality of rendered audio signals or difference between rendered audio signal combinations in the plurality of rendered audio signals.6.The method any one of claims 1 to 5, further comprising:obtaining real reverberation information corresponding to the plurality of first audio signals 301 based on the plurality of first audio signals 301; andperforming a ray tracing simulation 350 on the plurality of first audio signals 301 to obtain simulated reverberation information corresponding to the plurality of first audio signals 301;wherein the estimating of the parameters for reflecting the acoustic characteristics of the space based on the plurality of rendered audio signals comprises:estimating the parameters based on the plurality of rendered audio signals 311, the real reverberation information and the simulated reverberation information.7.The method of claim 6, wherein each of the real reverberation information and the simulated reverberation information comprises first reverberation information and second reverberation information,wherein the first reverberation information comprises reverberation time information, the reverberation time information indicating a time required for a sound pressure level to be lower to a predetermined level after the sound source stops making sound; andwherein the second reverberation information comprises a ratio of direct sound energy to early reflected sound energy.8.The method any one of claim 6 and 7, wherein the performing of the ray tracing simulation on the plurality of first audio signals to obtain the simulated reverberation information corresponding to the plurality of first audio signals comprises:initializing the parameters to be estimated based on geometry information 20 of the space; andperforming, using the initialized parameters 341, the ray tracing simulation 350 on the plurality of first audio signals 301 to obtain the simulated reverberation information corresponding to the plurality of first audio signals 301.9.The method any one of claims 1 to 8, wherein the parameters comprise a material absorption coefficient 30 of each plane comprised in geometry corresponding to the space.10.The method of any one of claims 1 to 9, further comprising:acquiring a plurality of second audio signals of the sound source recorded in real time at the plurality of locations after the plurality of first audio signals are recorded; andupdating the parameters based on the plurality of second audio signals.11.The method of claim 10, wherein the updating of the parameters based on the plurality of second audio signals comprises:detecting, based on the plurality of second audio signals, whether an environmental change occurs in the space; andupdating the parameters based on detecting that the environmental change occurs.12.The method of claim 11, wherein the updating of the parameters in the case of detecting that the environmental change occurs comprises:determining a plane needs to be split among planes comprised in geometry corresponding to the space;performing a ray tracing simulation 350 on propagation process of the plurality of second audio signals to determine position and gradient direction information of sound rays corresponding to the plurality of second audio signals colliding with the planes comprised in the geometry;splitting the plane needs to be split to expand the number of the planes comprised in the geometry, based on the position and the gradient direction information; anddetermining the expanded number of the parameters based on the expanded number of the planes.13.The method of claim 12, wherein the determining of the plane needs to be split among the planes comprised in the geometry corresponding to the space comprises:determining an expected benefit value of each plane 1303 based on sound related information and empirical information corresponding to the each plane comprised in the geometry corresponding to the space, wherein the sound related information comprises information relating to colliding of the sound rays with the each plane, and the empirical information comprises at least one of historical splitting information or historical earning information for the each plane; anddetermining the plane needs to be split based on the expected benefit value of the each plane 1303.14.An electronic apparatus comprising:at least one processor including processing circuitry,memory comprising one or more storage media storing instructions that, when executed by the at least one processor individually or collectively, cause the electronic device to:acquire a plurality of first audio signals 301 respectively recorded at a plurality of different locations in a space where a sound source is located;render, for each of the plurality of first audio signals 301, using a transmission response corresponding to at least one other first audio signal to obtain a rendered audio signal 311;estimate parameters for reflecting acoustic characteristics of the space based on the plurality of rendered audio signals 311; andperform, based on the estimated parameters, audio rendering 100 of a virtual audio 10 to obtain a reverberation audio 40.15.A computer-readable storage medium storing instructions which, when executed by at least one processor, cause an electronic apparatus to perform the method any one of claims 1 to 13.