Audio signal playback method, device and electronic device
By separating and generating target direct audio signals and reverberating audio signals from the recorded audio signals, the problem of increasing costs of high-performance equipment in the prior art is solved, and the effect of accurately restoring the sound field on ordinary equipment is achieved.
Patent Information
- Application Number
- CN202111122077.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-09-24
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2041-09-24
AI Technical Summary
In the prior art, in order to enhance the playback effect of the audio signal, high-performance playback equipment is often required when playing back and recording audio signals, resulting in an increase in equipment manufacturing cost.
By separating the recorded audio signals of each sound source from the first audio signal, the real-time orientation of the sound source relative to the user's head is determined, the target direct audio signal and the target reverberation audio signal are generated, and the sound field is fused and played in order to restore the sound field.
It realizes the accurate restoration of the sound field formed by the sound source without relying on high-performance hardware devices, and improves the playback effect of the audio signal.
Smart Images

Figure CN113889140B_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present disclosure relate to the field of computer technologies, and in particular, to an audio signal playing method, apparatus, and electronic device. Background Art
[0002] In practical applications, after an audio signal is recorded, a user often needs to play back the recorded audio signal. When playing back the recorded audio signal, various means can be used to enhance the playing effect of the audio signal, thereby improving the user's experience.
[0003] In related methods, a dedicated playback device is used to play back the recorded audio signal to enhance the playing effect of the audio signal. However, this method often has high hardware requirements for the playback device, and therefore, may increase the manufacturing cost of the device. Summary of the Invention
[0004] This disclosure section is provided to introduce concepts in a brief form, which will be described in detail in the subsequent detailed implementation section. This disclosure section is not intended to identify the key features or essential features of the claimed technical solution, nor is it intended to limit the scope of the claimed technical solution.
[0005] Embodiments of the present disclosure provide an audio signal playing method, apparatus, and electronic device, which can relatively accurately restore the sound field formed by at least one of the above sound sources.
[0006] In a first aspect, embodiments of the present disclosure provide an audio signal playing method, the method including: separating, from a first audio signal, recorded audio signals corresponding to each sound source among at least one sound source; determining, based on the first audio signal, the real-time azimuth of each sound source among the at least one sound source relative to the user's head; for each of the sound sources, generating a target direct audio signal corresponding to the sound source and a target reverberant audio signal corresponding to the sound source according to the real-time azimuth of the sound source and the recorded audio signal corresponding to the sound source; playing a second audio signal generated by fusing the target direct audio signals and target reverberant audio signals corresponding to each of the sound sources.
[0007] In a second aspect, embodiments of the present disclosure provide an audio signal playback device, which includes: a separation unit configured to separate, from a first audio signal, recorded audio signals corresponding to respective sound sources among at least one sound source; a determination unit configured to determine, based on the first audio signal, real-time azimuths of respective sound sources among the at least one sound source relative to a user's head; a generation unit configured to, for each of the respective sound sources, generate a target direct audio signal corresponding to the sound source and generate a target reverberation audio signal corresponding to the sound source according to the real-time azimuth of the sound source and the recorded audio signal corresponding to the sound source; and a playback unit configured to play a second audio signal generated by fusing the target direct audio signals and the target reverberation audio signals corresponding to the respective sound sources.
[0008] In a third aspect, embodiments of the present disclosure provide an electronic device, including: one or more processors; and a storage device configured to store one or more programs, which when executed by the one or more processors, cause the one or more processors to implement the audio signal playback method as described in the first aspect.
[0009] In a fourth aspect, embodiments of the present disclosure provide a computer-readable medium, on which a computer program is stored, and when the program is executed by a processor, the steps of the audio signal playback method as described in the first aspect are implemented.
[0010] The audio signal playback method, device, and electronic device provided by the embodiments of the present disclosure extract a direct audio signal corresponding to a sound source and a reverberation audio signal corresponding to the sound source according to the real-time azimuth of the sound source relative to the user's head. Thus, by taking into account the movement of the sound source, the target direct audio signal and the target reverberation audio signal corresponding to the sound source are extracted more accurately. Further, by playing the second audio signal, the sound field formed by the at least one sound source can be restored more accurately. BRIEF DESCRIPTION OF THE DRAWINGS
[0011] In combination with the accompanying drawings and with reference to the following specific embodiments, the above and other features, advantages, and aspects of the embodiments of the present disclosure will become more apparent. Throughout the drawings, the same or similar reference numerals denote the same or similar elements. It should be understood that the drawings are schematic and the original elements and elements are not necessarily drawn to scale.
[0012] Figure 1 is a flowchart of some embodiments of the audio signal playback method of the present disclosure;
[0013] Figure 2 is a flowchart of generating a target direct audio signal in some embodiments of the audio signal playback method of the present disclosure;
[0014] Figure 3 is a flowchart of generating a target reverberation audio signal in some embodiments of the audio signal playback method of the present disclosure;
[0015] Figure 4 are schematic structural diagrams of some embodiments of the audio signal playback device of the present disclosure;
[0016] Figure 5 is an exemplary system architecture to which the audio signal playback method of the present disclosure can be applied in some embodiments;
[0017] Figure 6 is a schematic diagram of the basic structure of an electronic device provided according to some embodiments of the present disclosure. Detailed implementation manners
[0018] Embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although some embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. On the contrary, these embodiments are provided to more thoroughly and completely understand the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are only for exemplary purposes and are not used to limit the protection scope of the present disclosure.
[0019] It should be understood that the steps recited in the method embodiments of the present disclosure can be executed in a different order and / or in parallel. In addition, the method embodiments may include additional steps and / or omit the steps shown. The scope of the present disclosure is not limited in this regard.
[0020] As used herein, the term "including" and its variants are open-ended, that is, "including but not limited to". The term "based on" is "at least partially based on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". The relevant definitions of other terms will be given in the following description.
[0021] It should be noted that the concepts such as "first" and "second" mentioned in the present disclosure are only used to distinguish different devices, modules or units, and are not used to limit the order or interdependence relationship of the functions performed by these devices, modules or units.
[0022] It should be noted that the modifications of "one" and "multiple" mentioned in the present disclosure are illustrative rather than restrictive. Those skilled in the art should understand that unless otherwise clearly specified in the context, it should be understood as "one or more".
[0023] The names of the messages or information exchanged between multiple devices in the embodiments of the present disclosure are only for illustrative purposes and are not used to limit the scope of these messages or information.
[0024] Please refer toFigure 1 , which shows the flow of some embodiments of the audio signal playback method according to the present disclosure. As Figure 1 shown, the audio signal playback method includes the following steps:
[0025] Step 101, separate the recorded audio signals corresponding to each sound source in at least one sound source from the first audio signal.
[0026] The first audio signal may be a recorded audio signal. Among them, the first audio signal contains the recorded audio signals corresponding to each sound source in the above at least one sound source. It can be understood that the recorded audio signal corresponding to a sound source may be an audio signal recorded for the sound generated by the sound source.
[0027] Optionally, the first audio signal is an audio signal recorded using a microphone array. At this time, the first audio signal is formed by audio signals recorded from multiple directions. The microphone array may be disposed on a terminal device or on a recording device (such as a recording pen) outside the terminal device.
[0028] In some scenarios, the execution subject of the audio signal playback method may use various audio signal separation algorithms to process the first audio signal, thereby separating the recorded audio signals corresponding to each sound source in at least one sound source from the first audio signal. For example, the audio signal separation algorithm may include, but is not limited to, the IVA (Independent Vector Analysis) algorithm, the MVDR (Minimum Variance Distortionless Response) algorithm, etc.
[0029] Step 102, based on the first audio signal, determine the real-time azimuth of each sound source in the above at least one sound source relative to the user's head.
[0030] During the process of recording the first audio signal, the sound source may move. Therefore, the azimuth of the sound source relative to the user's head may change. For example, the azimuth of the sound source relative to the user's head may be directly in front, directly behind, front left, back left, front right, back right, directly above, etc.
[0031] In some scenarios, the above execution subject may input the first audio signal into an azimuth recognition model to obtain the real-time azimuth of each sound source relative to the user's head output by the azimuth recognition model. Among them, the azimuth recognition model may be a neural network model that recognizes the real-time azimuth of each sound source relative to the user's head from an audio signal.
[0032] Step 103: For each of the above sound sources, based on the real-time orientation of the sound source and the recorded audio signal corresponding to the sound source, generate a target direct audio signal corresponding to the sound source and a target reverberant audio signal corresponding to the sound source.
[0033] The sound propagated from the sound source to the user's ear includes direct sound and reverberant sound. Among them, the direct sound can be the sound that directly propagates to the user's ear without reflection. The reverberant sound can be the sound that propagates to the user's ear after reflection.
[0034] It can be understood that the recorded audio signal is formed by at least one of the following: the direct audio signal corresponding to the direct sound propagated to the user's ear, and the reverberant audio signal corresponding to the reverberant sound propagated to the user's ear.
[0035] The target direct audio signal can be the direct audio signal extracted from the recorded audio signal. The target reverberant audio signal can be the reverberant audio signal extracted from the recorded audio signal.
[0036] In some scenarios, the above execution entity can input the real-time orientation of the sound source and the recorded audio signal corresponding to the sound source into the first extraction model to obtain the target direct audio signal output by the first extraction model. Among them, the first extraction model can be a neural network model for extracting the direct audio signal corresponding to the sound source. Similarly, the above execution entity can input the real-time orientation of the sound source and the recorded audio signal corresponding to the sound source into the second extraction model to obtain the target reverberant audio signal output by the second extraction model. Among them, the second extraction model can be a neural network model for extracting the reverberant audio signal corresponding to the sound source.
[0037] It can be understood that if the orientation of the sound source relative to the user's head changes, the direct sound and reverberant sound propagated from the sound source to the user's ear will also change. Therefore, based on the real-time orientation of the sound, the direct audio signal and reverberant audio signal corresponding to the sound source can be extracted more accurately.
[0038] Step 104: Play the second audio signal generated by fusing the target direct audio signals and target reverberant audio signals corresponding to the above various sound sources.
[0039] The second audio signal can include a left-channel audio signal and a right-channel audio signal.
[0040] In some scenarios, the above execution entity can fuse the target direct audio signals and target reverberant audio signals corresponding to each sound source into a second audio signal. Further, the above execution entity can play the second audio signal.
[0041] It should be noted that the above execution entity can play the second audio signal through a speaker or through headphones.
[0042] It can be understood that by playing the second audio signal, the sound field formed by the at least one sound source can be restored.
[0043] In this embodiment, according to the real-time orientation of the sound source relative to the user's head, the direct audio signal corresponding to the sound source and the reverberant audio signal corresponding to the sound source are extracted. Thus, by taking into account the movement of the sound source, the direct audio signal and the reverberant audio signal corresponding to the sound source are extracted more accurately. Further, by playing the second audio signal, the sound field formed by the at least one sound source can be restored more accurately.
[0044] In some embodiments, the above-mentioned execution entity can determine the real-time orientation of each of the above-mentioned sound sources relative to the user's head in the following manner.
[0045] First step, based on the first audio signal, determine the movement trajectory of each of the at least one sound source.
[0046] The movement trajectory may include the positions of the sound source at at least one moment.
[0047] In some scenarios, the above-mentioned execution entity can input the first audio signal into a position recognition model to obtain the positions of each sound source at at least one moment output by the position recognition model. Among them, the position recognition model can be a neural network model for recognizing the positions of sound sources at at least one moment. Further, for each of the above-mentioned sound sources, the above-mentioned execution entity can determine the movement trajectory of the sound source according to the positions of the sound source at at least one moment.
[0048] Second step, for each of the above-mentioned sound sources, determine the real-time position of the sound source from the movement trajectory of the sound source, and based on the real-time position of the sound source and the real-time attitude data of the user's head, determine the real-time orientation of the sound source relative to the user's head.
[0049] The real-time attitude data of the user's head can be data representing the attitude of the user's head collected in real time. The above-mentioned real-time attitude data may include the pitch angle and azimuth angle of the user's head.
[0050] In some scenarios, attitude detection sensors such as accelerometers, angular velocity sensors, and gyroscopes are provided on the earphones communicatively connected to the terminal device. The earphones can send the acceleration, angular velocity, and magnetic induction intensity collected by the attitude detection sensors to the terminal device. Further, the above-mentioned execution entity can determine the pitch angle and azimuth angle of the user's head according to the acceleration, angular velocity, and magnetic induction intensity sent by the earphones.
[0051] It can be understood that the movement of the sound source or the change in the posture of the user's head may cause the azimuth of the sound source relative to the user's head to change. Therefore, based on the real-time position of the sound source and the real-time posture data of the user's head, the azimuth of the sound source relative to the user's head can be accurately determined in real time.
[0052] In some embodiments, the above-mentioned execution entity can determine the movement trajectories of the respective sound sources in the following manner.
[0053] Specifically, a sound source localization algorithm and a sound source tracking algorithm are used to process the first audio signal to determine the movement trajectories of the respective sound sources among the at least one sound source.
[0054] The sound source localization algorithm is used to locate the real-time position of the sound source. For example, the sound source localization algorithm may include, but is not limited to, the GCC (Generalized Cross Correlation) algorithm, the GCC-PHAT (Generalized Cross Correlation-Phase Transform) algorithm, etc.
[0055] The sound source tracking algorithm is used to determine the movement trajectory of the sound source by tracking the real-time position of the sound source.
[0056] It can be understood that through the sound source localization algorithm and the sound source tracking algorithm, the movement trajectory of the sound source can be determined quickly and accurately. Further, it is possible to quickly and accurately restore the sound field formed by the at least one sound source.
[0057] In some embodiments, the above-mentioned execution entity may generate a target direct audio signal corresponding to the sound source according to the Figure 2 flow shown, and this flow includes step 201.
[0058] Step 201: For each of the above-mentioned sound sources, perform a first processing step. Among them, the first processing step includes steps 2011 to 2012.
[0059] Step 2011: Select a first convolution function corresponding to the real-time azimuth of the sound source.
[0060] The first convolution function is used to extract the target direct audio signal corresponding to the sound source from the audio signal. Optionally, the first convolution function is the HRTF (Head Related Transfer Function).
[0061] Corresponding first convolution functions are set for each azimuth of the sound source relative to the user's head. The above-mentioned execution entity can select the first convolution function corresponding to the real-time azimuth of the sound source from the set first convolution functions.
[0062] In step 2012, based on the recorded audio signal corresponding to the sound source and the selected first convolution function, a convolution audio signal is obtained through convolution, and a target direct audio signal corresponding to the sound source is generated.
[0063] The convolution audio signal may be the convolution result of the recorded audio signal and the first convolution function.
[0064] In some scenarios, the above-mentioned execution entity may use the obtained convolution audio signal as the target direct audio signal corresponding to the sound source.
[0065] It can be understood that the direct sounds transmitted to the user's ears from sound sources located in different directions are different. Therefore, on the premise of considering the movement of the sound source, the first convolution function is used to accurately extract the target direct audio signal corresponding to the sound source from the recorded audio signal corresponding to the sound source.
[0066] In some embodiments, the above-mentioned execution entity may execute step 2012 in the following manner.
[0067] Specifically, based on the actual distance between the sound source and the user's head, the convolution audio signal is corrected to generate the target direct audio signal corresponding to the sound source.
[0068] During the playback of the audio signal, the sound source may move, resulting in a change in its actual distance from the user's head. The first convolution function may determine the convolution audio signal based on a preset distance between the sound source and the user's head. Therefore, there may be an error between the convolution audio signal obtained through the first convolution function and the target direct audio signal.
[0069] It can be understood that correcting the convolution audio signal based on the movement of the sound source can reduce the error of the finally obtained target direct audio signal.
[0070] In some embodiments, the above-mentioned execution entity may Figure 3 generate a target reverberation audio signal corresponding to the sound source according to the shown process, and this process includes step 301.
[0071] In step 301, for each of the above-mentioned sound sources, a second processing step is executed. The second processing step includes steps 3011 to 3013.
[0072] In step 3011, the recorded audio signal corresponding to the sound source is encoded into a surround audio signal through a predetermined audio encoding method.
[0073] The predetermined audio encoding method may be an audio encoding method for encoding a recorded audio signal into a surround audio signal. The surround audio signal generated through the predetermined audio encoding method includes audio signals of a target number of channels.
[0074] Optionally, the predetermined audio encoding method is the Ambisonic encoding method. In some scenarios, the surround audio signal generated by the Ambisonic encoding method may include audio signals of 4 channels.
[0075] Step 3012: Decode the surround audio signal corresponding to the sound source into a target surround audio signal suitable for speaker playback through the audio decoding method corresponding to the speaker.
[0076] In practical applications, a speaker has a corresponding audio decoding method.
[0077] Step 3013: Generate a target reverberant audio signal corresponding to the sound source by convolving the target surround audio signal corresponding to the sound source with a second convolution function corresponding to the speaker.
[0078] The second convolution function is used to extract the target reverberant audio signal corresponding to the sound source from the audio signal. Optionally, the second convolution function is an RIR (Room Impulse Response) function.
[0079] In practical applications, different speakers often have different performances. Therefore, by setting corresponding second convolution functions for different speakers, target reverberant audio signals matching the performances of the speakers can be extracted.
[0080] It can be understood that by combining the predetermined audio encoding method and the second convolution function, when extracting the target reverberant audio signal, not only the performance of the speaker can be taken into account, but also the sound surround feeling given to the user by the finally extracted target reverberant audio signal can be enhanced. Thus, a target reverberant audio signal with high accuracy and strong sound surround effect for the user can be extracted from the recorded audio signal. Further, by playing the second audio signal, the feeling of the user being in a real sound field can be enhanced.
[0081] For further reference Figure 4 , as an implementation of the methods shown in the above figures, the present disclosure provides some embodiments of an audio signal playback device. The device embodiments correspond to Figure 1 the method embodiments shown, and the device can be specifically applied to various electronic devices.
[0082] As Figure 4As shown in the figure, the audio signal playback device of this embodiment includes: a separation unit 401, a determination unit 402, a generation unit 403, and a playback unit 404. The separation unit 401 is configured to separate the recorded audio signals corresponding to each sound source in at least one sound source from the first audio signal; the determination unit 402 is configured to determine the real-time azimuth of each sound source in the at least one sound source relative to the user's head based on the first audio signal; the generation unit 403 is configured to generate a target direct audio signal corresponding to each sound source and a target reverberation audio signal corresponding to each sound source according to the real-time azimuth of the sound source and the recorded audio signal corresponding to the sound source; the playback unit 404 is configured to play a second audio signal generated by fusing the target direct audio signals and target reverberation audio signals corresponding to each sound source.
[0083] In this embodiment, for the specific processing of the separation unit 401, determination unit 402, generation unit 403, and playback unit 404 of the audio signal playback device and the technical effects brought by them, reference can be made to Figure 1 the relevant descriptions of steps 101, 102, 103, and 104 in the corresponding embodiments, which will not be elaborated here.
[0084] In some embodiments, the determination unit 402 is further configured to, for each of the above sound sources, determine the real-time position of the sound source from the movement trajectory of the sound source, and determine the real-time azimuth of the sound source relative to the user's head based on the real-time position of the sound source and the real-time attitude data of the user's head.
[0085] In some embodiments, the determination unit 402 is further configured to process the first audio signal using a sound source localization algorithm and a sound source tracking algorithm to determine the movement trajectories of each of the at least one sound source, where the sound source localization algorithm is used to locate the real-time position of the sound source, and the sound source tracking algorithm is used to determine the movement trajectory of the sound source by tracking the real-time position of the sound source.
[0086] In some embodiments, the generation unit 403 is further configured to, for each of the above sound sources, perform a first processing step: select a first convolution function corresponding to the real-time azimuth of the sound source, where the first convolution function is used to extract the target direct audio signal corresponding to the sound source from the audio signal; generate the target direct audio signal corresponding to the sound source based on the convolution audio signal obtained by convolving the recorded audio signal corresponding to the sound source with the selected first convolution function.
[0087] In some embodiments, the generation unit 403 is further configured to correct the convolution audio signal based on the actual distance between the sound source and the user's head to generate the target direct audio signal corresponding to the sound source.
[0088] In some embodiments, the generating unit 403 is further configured to, for each of the above sound sources, perform a second processing step: encoding the recorded audio signal corresponding to the sound source into a surround audio signal by a predetermined audio encoding method, where the surround audio signal generated by the predetermined audio encoding method includes audio signals of a target number of channels; decoding the surround audio signal corresponding to the sound source into a target surround audio signal suitable for speaker playback by an audio decoding method corresponding to the speaker; generating a target reverberant audio signal corresponding to the sound source by convolving the target surround audio signal corresponding to the sound source with a second convolution function corresponding to the speaker, where the second convolution function is used to extract the target reverberant audio signal corresponding to the sound source from the audio signal.
[0089] In some embodiments, the first audio signal is an audio signal recorded using a microphone array.
[0090] Further reference Figure 5 , Figure 5 shows an exemplary system architecture to which the audio signal playback method according to some embodiments of the present disclosure can be applied.
[0091] As Figure 5 shown, the system architecture may include terminal devices 501, 502, and headphones 503, 504. Among them, the terminal device and the headphones can establish a communication connection through Bluetooth, a headphone cable, etc.
[0092] Various applications (such as audio signal processing applications, audio and video playback applications, etc.) can be installed on the terminal devices 501, 502.
[0093] In some scenarios, the terminal devices 501, 502 can separate the recorded audio signals corresponding to each sound source in at least one sound source from the first audio signal; the terminal devices 501, 502 can determine the real-time azimuth of each sound source in the at least one sound source relative to the user's head based on the first audio signal; for each of the above sound sources, the terminal devices 501, 502 can generate a target direct audio signal corresponding to the sound source and generate a target reverberant audio signal corresponding to the sound source according to the real-time azimuth of the sound source and the recorded audio signal corresponding to the sound source; the terminal devices 501, 502 can play a second audio signal generated by fusing the target direct audio signals and target reverberant audio signals corresponding to each of the above sound sources through the headphones 503, 504.
[0094] In some scenarios, the terminal devices 501, 502 can play the second audio signal through the speakers provided thereon. At this time, Figure 5 the system architecture shown does not include the headphones 503, 504.
[0095] The terminal devices 501 and 502 can be hardware or software. When the terminal devices 501 and 502 are hardware, they can be various electronic devices with a display screen and supporting information interaction, including but not limited to smartphones, tablets, laptop computers, desktop computers, and so on. When the terminal devices 501 and 502 are software, they can be installed in the above-listed electronic devices, and can be implemented as multiple software or software modules, or can be implemented as a single software or software module, without specific limitation here.
[0096] It should be noted that the audio signal playback method provided by the embodiments of the present disclosure can be executed by a terminal device. Correspondingly, the audio signal playback device can be provided in the terminal device.
[0097] It should be understood that Figure 5 the number of terminal devices and headphones in
[0098] is merely illustrative. According to the implementation requirements, there can be any number of terminal devices and headphones. Figure 6 is merely illustrative. According to the implementation requirements, there can be any number of terminal devices and headphones. Figure 5 The terminal devices in some embodiments of the present disclosure can include but are not limited to mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Tablet Computers), PMPs (Portable Multimedia Players), in-vehicle terminals (such as in-vehicle navigation terminals), etc., and fixed terminals such as digital TVs, desktop computers, etc. Figure 6 The electronic device shown in Figure 6 is merely an example and should not impose any limitation on the functions and usage scope of the embodiments of the present disclosure.
[0099] As Figure 6 shown, the electronic device can include a processing device (such as a central processing unit, a graphics processing unit, etc.) 601, which can perform various appropriate actions and processes according to the program stored in the read-only memory (ROM) 602 or the program loaded from the storage device 608 into the random access memory (RAM) 603. In the RAM 603, various programs and data required for the operation of the electronic device are also stored. The processing device 601, the ROM 602, and the RAM 603 are connected to each other through a bus 604. The input / output (I / O) interface 605 is also connected to the bus 604.
[0100] Typically, the following devices can be connected to the I / O interface 605: input devices 606 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; output devices 607 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; storage devices 608 including, for example, magnetic tape, a hard disk, etc.; and a communication device 609. The communication device 609 can allow the electronic device to communicate with other devices wirelessly or wiredly to exchange data. Although Figure 6 an electronic device having various devices is shown, it should be understood that it is not required to implement or have all the shown devices. Instead, more or fewer devices can be implemented or had. Figure 6 Each block shown in
[0101] can represent one device or, as needed, multiple devices.
[0102] It should be noted that the computer-readable medium described in some embodiments of the present disclosure may be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. A computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of the computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In some embodiments of the present disclosure, the computer-readable storage medium may be any tangible medium that contains or stores a program, which can be used by or in conjunction with an instruction execution system, apparatus, or device. In some embodiments of the present disclosure, a computer-readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The computer-readable signal medium may also be any computer-readable medium other than the computer-readable storage medium, which can send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium may be transmitted using any appropriate medium, including but not limited to: wires, optical cables, RF (radio frequency), etc., or any suitable combination of the above.
[0103] In some embodiments, the client and the server can communicate using any currently known or future-developed network protocol such as HTTP (HyperText Transfer Protocol), and can be interconnected with digital data communication in any form or medium (e.g., a communication network). Examples of communication networks include local area networks ("LANs"), wide area networks ("WANs"), the Internet (e.g., the Internet), and end-to-end networks (e.g., ad hoc end-to-end networks), as well as any currently known or future-developed networks.
[0104] The above computer-readable medium may be included in the above electronic device or exist separately without being assembled into the electronic device. The above computer-readable medium carries one or more programs, and when the one or more programs are executed by the electronic device, the electronic device is caused to: separate, from a first audio signal, the recorded audio signals corresponding to each sound source in at least one sound source; determine, based on the first audio signal, the real-time azimuth of each sound source in the at least one sound source relative to the user's head; for each of the sound sources, generate a target direct audio signal corresponding to the sound source and a target reverberant audio signal corresponding to the sound source according to the real-time azimuth of the sound source and the recorded audio signal corresponding to the sound source; and play a second audio signal generated by fusing the target direct audio signals and the target reverberant audio signals corresponding to each of the sound sources.
[0105] Computer program code for performing the operations of some embodiments of the present disclosure may be written in one or more programming languages or combinations thereof. The programming languages include, but are not limited to, object-oriented programming languages such as Java, Smalltalk, C++, and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code may execute entirely on the user's computer, partially on the user's computer, execute as a stand-alone software package, execute partially on the user's computer and partially on a remote computer, or execute entirely on the remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., through the Internet using an Internet service provider).
[0106] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowchart or block diagram may represent a module, a program segment, or a part of code that contains one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and combinations of blocks in the block diagram and / or flowchart, may be implemented by a dedicated hardware-based system for performing the specified functions or operations, or may be implemented by a combination of dedicated hardware and computer instructions.
[0107] The units described in some embodiments of the present disclosure can be implemented in software or in hardware. Among them, the names of these units do not, in some cases, constitute a limitation on the unit itself. For example, the determination unit can also be described as a unit that "determines the real-time azimuth of each sound source in the above at least one sound source relative to the user's head based on the first audio signal".
[0108] The functions described above herein can be performed, at least in part, by one or more hardware logic components. For example, without limitation, exemplary types of hardware logic components that can be used include: Field Programmable Gate Arrays (FPGAs), Application Specific Integrated Circuits (ASICs), Application Specific Standard Products (ASSPs), Systems on Chip (SOCs), Complex Programmable Logic Devices (CPLDs), and so on.
[0109] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media would include electrical connections based on one or more wires, portable computer disks, hard disks, Random Access Memory (RAM), Read Only Memory (ROM), Erasable Programmable Read Only Memory (EPROM or flash memory), optical fibers, portable compact disc read only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0110] The above description is only some preferred embodiments of the present disclosure and an explanation of the applied technical principles. Those skilled in the art should understand that the scope of disclosure involved in the embodiments of the present disclosure is not limited to the technical solutions formed by the specific combination of the above technical features, and should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the above disclosure concept. For example, technical solutions formed by mutually replacing the above features with (but not limited to) technical features having similar functions disclosed in the present disclosure.
[0111] In addition, although the operations are depicted in a particular order, this should not be construed as requiring that the operations be performed in the particular order shown or in sequential order. In certain circumstances, multitasking and parallel processing may be advantageous. Similarly, although several specific implementation details are included in the foregoing description, these should not be construed as limiting the scope of the present disclosure. Certain features that are described in the context of separate embodiments may also be implemented in combination in a single embodiment. Conversely, various features that are described in the context of a single embodiment may also be implemented separately or in any suitable subcombination in multiple embodiments.
[0112] Although the subject matter has been described in language specific to structural features and / or methodological acts, it is to be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or acts described above. Rather, the specific features and acts described above are merely example forms of implementing the claims.
Claims
1. An audio signal playback method, characterized in that, Comprising: Separating, from a first audio signal, the recorded audio signals corresponding to respective sound sources in at least one sound source; Based on the first audio signal, determining the real-time azimuth of respective sound sources in the at least one sound source relative to a user's head; For each of the sound sources, generating a target direct audio signal corresponding to the sound source and a target reverberant audio signal corresponding to the sound source according to the real-time azimuth of the sound source and the recorded audio signal corresponding to the sound source; Playing a second audio signal generated by fusing the target direct audio signals and the target reverberant audio signals corresponding to respective sound sources; Wherein, generating the target direct audio signal corresponding to the sound source includes: For each of the sound sources, performing a first processing step: Selecting a first convolution function corresponding to the real-time azimuth of the sound source, wherein the first convolution function is used to extract the target direct audio signal corresponding to the sound source from an audio signal; Generating the target direct audio signal corresponding to the sound source based on a convolution audio signal obtained by convolving the recorded audio signal corresponding to the sound source with the selected first convolution function.
2. The method according to claim 1, wherein The determining, based on the first audio signal, the real-time azimuth of respective sound sources in the at least one sound source relative to a user's head includes: Based on the first audio signal, determining the movement trajectories of respective sound sources in the at least one sound source; For each of the sound sources, determining the real-time position of the sound source from the movement trajectory of the sound source, and based on the real-time position of the sound source and the real-time attitude data of the user's head, determining the real-time azimuth of the sound source relative to the user's head.
3. The method according to claim 2, wherein The determining, based on the first audio signal, the movement trajectories of respective sound sources in the at least one sound source includes: Processing the first audio signal using a sound source localization algorithm and a sound source tracking algorithm to determine the movement trajectories of respective sound sources in the at least one sound source, wherein the sound source localization algorithm is used to locate the real-time position of a sound source, and the sound source tracking algorithm is used to determine the movement trajectory of a sound source by tracking the real-time position of the sound source.
4. The method according to claim 1, characterized in that, The generating the target direct audio signal corresponding to the sound source based on a convolution audio signal obtained by convolving the recorded audio signal corresponding to the sound source with the selected first convolution function includes: Based on the actual distance between the sound source and the user's head, correcting the convolution audio signal to generate the target direct audio signal corresponding to the sound source.
5. The method according to claim 1, wherein The generating the target reverberant audio signal corresponding to the sound source includes: For each of the sound sources, performing a second processing step: Encoding the recorded audio signal corresponding to the sound source into a surround audio signal by a predetermined audio encoding method, wherein the surround audio signal generated by the predetermined audio encoding method includes audio signals of a target number of channels; Decoding the surround audio signal corresponding to the sound source into a target surround audio signal suitable for speaker playback by an audio decoding method corresponding to the speaker; Generating the target reverberant audio signal corresponding to the sound source by convolving the target surround audio signal corresponding to the sound source with a second convolution function corresponding to the speaker, wherein the second convolution function is used to extract the target reverberant audio signal corresponding to the sound source from an audio signal.
6. The method according to any one of claims 1 to 4, characterized in that, The first audio signal is an audio signal recorded using a microphone array.
7. An audio signal playback device, characterized in that, Comprising: A separation unit, configured to separate, from a first audio signal, recorded audio signals corresponding to respective sound sources among at least one sound source; A determination unit, configured to determine, based on the first audio signal, real-time azimuths of respective sound sources among the at least one sound source relative to a user's head; A generation unit, configured to, for each of the respective sound sources, generate a target direct audio signal corresponding to the sound source and generate a target reverberation audio signal corresponding to the sound source according to the real-time azimuth of the sound source and the recorded audio signal corresponding to the sound source; A playback unit, configured to play a second audio signal generated by fusing the target direct audio signals and the target reverberation audio signals corresponding to the respective sound sources; Wherein, generating the target direct audio signal corresponding to the sound source includes: For each of the respective sound sources, performing a first processing step: Selecting a first convolution function corresponding to the real-time azimuth of the sound source, wherein the first convolution function is used to extract the target direct audio signal corresponding to the sound source from an audio signal; Generating the target direct audio signal corresponding to the sound source based on a convolution audio signal obtained by convolving the recorded audio signal corresponding to the sound source with the selected first convolution function.
8. An electronic device, characterized in that, Comprising: One or more processors; A storage device, configured to store one or more programs, When the one or more programs are executed by the one or more processors, enabling the one or more processors to implement the method according to any one of claims 1-6.
9. A computer-readable medium having a computer program stored thereon, characterized in that, When the program is executed by a processor, it implements the method according to any one of claims 1-6.
Citation Information
Patent Citations
System for generating immersive audio utilizing visual cues
US20170070835A1
System, apparatus and method for consistent acoustic scene reproduction based on adaptive functions
US20170078819A1
Audio system for dynamic determination of personalized acoustic transfer functions
US20190394564A1