Audio generation method and apparatus, electronic device, and storage medium

By combining deep learning-based audio track separation and audio timing detection with music structure analysis, three-dimensional spatial metadata is generated, solving the technical challenge of converting stereo audio into immersive 3D audio, achieving more realistic three-dimensional audio effects, and improving the user experience.

CN114827886BActive Publication Date: 2026-03-31BEIJING DAJIA INTERNET INFORMATION TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-04-26
Publication Date
2026-03-31

AI Technical Summary

Technical Problem

Existing technologies struggle to effectively convert stereo audio into immersive 3D audio, resulting in monotonous 3D audio effects and an inability to accurately separate audio tracks from individual sound sources such as human voices and musical instruments.

Method used

The deep learning-based audio track separation system separates audio track signals from multiple sound sources and generates three-dimensional spatial metadata based on user location information and feature information. Combined with audio cadence detection and music structure analysis, each audio track signal is given a unique three-dimensional spatial trajectory. Finally, a three-dimensional audio signal is generated through spatial audio rendering technology.

Benefits of technology

It enhances the immersive experience of 3D audio, improves the user's auditory experience, and generates more realistic 3D audio effects.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114827886B_ABST
    Figure CN114827886B_ABST
Patent Text Reader

Abstract

The present disclosure provides an audio generation method, device, electronic equipment and storage medium. The audio generation method can include: obtaining a to-be-processed audio signal; obtaining a plurality of track signals corresponding to a plurality of sound sources by performing track separation on the to-be-processed audio signal; determining user orientation information in a three-dimensional space, and generating three-dimensional space metadata corresponding to each track signal in the plurality of track signals; and generating a three-dimensional audio signal corresponding to the to-be-processed audio signal based on the user orientation information, the plurality of separated track signals, and the three-dimensional space metadata corresponding to each track signal. The present disclosure can generate a three-dimensional audio that is closer to a real three-dimensional audio, increase the immersion of the three-dimensional audio, and improve the user experience.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of audio processing technology, and in particular to an audio generation method, an audio generation device, an electronic device, a storage medium, and a program product. Background Technology

[0002] 3D audio refers to audio signals that are distributed in three-dimensional space and played dynamically. For example, as music is played, human voices move from far to near in three-dimensional space, and instruments such as drums and bass bounce back and forth in different places in three-dimensional space.

[0003] However, most mainstream audio formats are currently either dual-channel stereo or single-channel, while multi-channel and 3D audio content is relatively scarce. Therefore, automatically converting traditional stereo audio into 3D audio is an important direction for the development of 3D audio. Summary of the Invention

[0004] This disclosure provides an audio generation method, an audio generation apparatus, an electronic device, and a storage medium to at least solve the aforementioned problems.

[0005] According to a first aspect of the present disclosure, an audio generation method is provided. The method may include: acquiring an audio signal to be processed; obtaining multiple audio track signals for multiple sound sources by performing track separation on the audio signal to be processed; determining user location information in three-dimensional space and generating three-dimensional spatial metadata corresponding to each of the multiple audio track signals, wherein the three-dimensional spatial metadata includes location information and sound width information of the corresponding sound source in three-dimensional space; and generating a three-dimensional audio signal corresponding to the audio signal to be processed based on the user location information, the separated multiple audio track signals, and the three-dimensional spatial metadata corresponding to each audio track signal.

[0006] As an example, generating three-dimensional spatial metadata corresponding to each of the plurality of audio track signals may include: obtaining feature information of the audio signal to be processed, wherein the feature information includes at least one of tempo information and structural information of the audio signal to be processed, wherein the structural information includes type information of each audio segment of the audio signal to be processed; and generating three-dimensional spatial metadata corresponding to each audio track signal based on the feature information.

[0007] As an example, generating three-dimensional spatial metadata corresponding to each audio track signal based on the feature information may include: for the human voice signal in the plurality of audio track signals, determining the position adjustment information of the sound source corresponding to the human voice signal relative to the user's orientation information in three-dimensional space according to the type of each audio segment in the structure information, and determining the three-dimensional spatial metadata of the human voice signal based on the position adjustment information.

[0008] As an example, generating three-dimensional spatial metadata corresponding to each of the audio track signals based on the feature information may include: for the instrument signals in the plurality of audio track signals, determining the movement information of the sound source corresponding to the instrument signal in three-dimensional space according to the beat information and the type of each audio segment in the structure information, and determining the three-dimensional spatial metadata of the instrument signal based on the movement information.

[0009] As an example, generating three-dimensional spatial metadata corresponding to each of the plurality of audio track signals may include: determining a preset template corresponding to each audio track signal from a plurality of preset templates, wherein the preset template includes at least one of the following: the movement trajectory information, movement speed information, and sound width change information of the corresponding sound source in three-dimensional space; and generating three-dimensional spatial metadata for each audio track signal using the determined preset template.

[0010] As an example, generating three-dimensional spatial metadata corresponding to each of the plurality of audio track signals may include: obtaining user-inputted setting information, wherein the setting information includes at least one of the movement trajectory, movement speed, and sound width variation value of each sound source corresponding to the plurality of audio track signals in three-dimensional space; and generating three-dimensional spatial metadata for each audio track signal based on the setting information.

[0011] As an example, generating a three-dimensional audio signal corresponding to the audio signal to be processed based on the user's location information, the separated multiple audio track signals, and the three-dimensional spatial metadata corresponding to each audio track signal may include: identifying the type of playback device used to play the three-dimensional audio signal; obtaining a rendering strategy corresponding to the type of playback device; and generating a three-dimensional audio signal corresponding to the audio signal to be processed based on the user's location information, the separated multiple audio track signals, and the three-dimensional spatial metadata corresponding to each audio track signal through the rendering strategy.

[0012] As an example, generating a three-dimensional audio signal corresponding to the audio signal to be processed through the rendering strategy may include: when the playback device is an in-ear playback device, generating a three-dimensional audio signal corresponding to the audio track signal based on the orientation information corresponding to each audio frame of the audio track signal and the user's orientation information for each audio track signal; and adjusting the sound width of the three-dimensional audio signal corresponding to the audio track signal based on the sound width information of each audio frame.

[0013] As an example, generating a three-dimensional audio signal corresponding to the audio signal to be processed through the rendering strategy may include: when the playback device is an external playback device, for each audio track signal, rendering the audio track signal based on the directional information of the sound source corresponding to the audio track signal and the directional information of multiple speakers to generate a three-dimensional audio signal corresponding to the audio track signal; adjusting the sound width of the three-dimensional audio signal corresponding to the audio track signal based on the sound width information of the sound source.

[0014] According to a second aspect of the present disclosure, an audio generation apparatus is provided. The apparatus may include: an acquisition module configured to acquire an audio signal to be processed; an audio track separation module configured to obtain multiple audio track signals for multiple sound sources by performing audio track separation on the audio signal to be processed; a metadata generation module configured to determine user location information in three-dimensional space and generate three-dimensional space metadata corresponding to each of the multiple audio track signals, wherein the three-dimensional space metadata includes location information and sound width information of the corresponding sound source in three-dimensional space; and a rendering module configured to generate a three-dimensional audio signal corresponding to the audio signal to be processed based on the user location information, the separated multiple audio track signals, and the three-dimensional space metadata corresponding to each audio track signal.

[0015] As an example, the metadata generation module can be configured to: acquire feature information of the audio signal to be processed, wherein the feature information includes at least one of tempo information and structural information of the audio signal to be processed, wherein the structural information includes type information of each audio segment of the audio signal to be processed; and generate three-dimensional spatial metadata corresponding to each audio track signal based on the feature information.

[0016] As an example, the metadata generation module can be configured to: for the human voice signal in the plurality of audio track signals, determine the position adjustment information of the sound source corresponding to the human voice signal relative to the user's orientation information in three-dimensional space according to the type of each audio segment in the structure information, and determine the three-dimensional space metadata of the human voice signal based on the position adjustment information.

[0017] As an example, the metadata generation module can be configured to: for the instrument signal in the plurality of audio track signals, determine the movement information of the sound source corresponding to the instrument signal in three-dimensional space according to the beat information and the type of each audio segment in the structure information, and determine the three-dimensional space metadata of the instrument signal based on the movement information.

[0018] As an example, the metadata generation module can be configured to: determine a preset template corresponding to each audio track signal from a plurality of preset templates, wherein the preset template includes at least one of the movement trajectory information, movement speed information and sound width change information of the corresponding sound source in three-dimensional space; and generate three-dimensional space metadata for each audio track signal using the determined preset template.

[0019] As an example, the metadata generation module can be configured to: obtain user-inputted setting information, wherein the setting information includes at least one of the movement trajectory, movement speed, and sound width variation value of each sound source corresponding to the plurality of audio track signals in three-dimensional space; and generate three-dimensional space metadata for each audio track signal based on the setting information.

[0020] As an example, the rendering module can be configured to: identify the type of playback device used to play the three-dimensional audio signal; obtain a rendering strategy corresponding to the type of the playback device; and generate a three-dimensional audio signal corresponding to the audio signal to be processed based on the user's location information, the separated multiple audio track signals, and the three-dimensional spatial metadata corresponding to each audio track signal.

[0021] As an example, the rendering module can be configured to: when the playback device is an in-ear playback device, generate a three-dimensional audio signal corresponding to the audio track signal based on the orientation information corresponding to each audio frame of the audio track signal and the user's orientation information for each audio track signal; and adjust the sound width of the three-dimensional audio signal corresponding to the audio track signal based on the sound width information of each audio frame.

[0022] As an example, the rendering module can be configured to: when the playback device is an external playback device, for each audio track signal, render the audio track signal based on the directional information of the sound source corresponding to the audio track signal and the directional information of multiple speakers to generate a three-dimensional audio signal corresponding to the audio track signal; and adjust the sound width of the three-dimensional audio signal corresponding to the audio track signal based on the sound width information of the sound source.

[0023] According to a third aspect of the present disclosure, an electronic device is provided, the electronic device may include: at least one processor; at least one memory storing computer-executable instructions, wherein the computer-executable instructions, when executed by the at least one processor, cause the at least one processor to perform the audio generation method as described above.

[0024] According to a fourth aspect of the present disclosure, a computer-readable storage medium is provided that stores instructions which, when executed by at least one processor, cause the at least one processor to perform the audio generation method as described above.

[0025] According to a fifth aspect of the present disclosure, a computer program product is provided, wherein instructions in the computer program product are executed by at least one processor in an electronic device to perform the audio generation method as described above.

[0026] The technical solutions provided by the embodiments of this disclosure have at least the following beneficial effects:

[0027] By separating the audio tracks of multiple individual sound sources from the audio to be processed and assigning these sound sources various different three-dimensional spatial trajectories, the audio to be processed is converted into three-dimensional audio, thereby obtaining three-dimensional audio that is closer to true three-dimensional audio, increasing the immersiveness of the three-dimensional audio and improving the user experience.

[0028] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and are not intended to limit this disclosure. Attached Figure Description

[0029] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this disclosure and, together with the description, serve to explain the principles of this disclosure, and are not intended to unduly limit this disclosure.

[0030] Figure 1 This is a flowchart of an audio generation method according to an embodiment of the present disclosure;

[0031] Figure 2 and Figure 3 A schematic diagram of a virtual three-dimensional space according to an embodiment of the present disclosure is shown;

[0032] Figure 4 This is a schematic flowchart of an audio generation method according to another embodiment of the present disclosure;

[0033] Figure 5 This is a schematic diagram of the structure of an audio generation device according to an embodiment of the present disclosure;

[0034] Figure 6 This is a block diagram of an audio generation apparatus according to an embodiment of the present disclosure;

[0035] Figure 7 This is a block diagram of an electronic device according to an embodiment of the present disclosure.

[0036] Throughout the accompanying drawings, it should be noted that the same reference numerals are used to denote the same or similar elements, features, and structures. Detailed Implementation

[0037] To enable those skilled in the art to better understand the technical solutions of this disclosure, the technical solutions in the embodiments of this disclosure will be clearly and completely described below with reference to the accompanying drawings.

[0038] The following description, provided with reference to the accompanying drawings, is intended to aid in a full understanding of embodiments of the present disclosure as defined by the claims and their equivalents. Various specific details are included to aid understanding, but these details are to be considered exemplary only. Therefore, those skilled in the art will recognize that various changes and modifications may be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Furthermore, for clarity and brevity, descriptions of well-known functions and structures are omitted.

[0039] The terms and words used in the following description and claims are not limited to their literal meaning, but are intended solely by the inventors to achieve a clear and consistent understanding of this disclosure. Therefore, it will be apparent to those skilled in the art that the following description of various embodiments of this disclosure is provided for illustrative purposes only and is not intended to limit the purpose of this disclosure as defined by the claims and their equivalents.

[0040] It should be noted that the terms "first," "second," etc., used in the specification, claims, and accompanying drawings of this disclosure are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this disclosure described herein can be implemented in orders other than those illustrated or described herein. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this disclosure. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this disclosure as detailed in the appended claims.

[0041] In related technologies, traditional signal processing schemes are used to separate the direct sound and background sound from stereo signals, and then different processing is applied to the direct sound and background sound to generate 3D audio signals. However, the audio track separation algorithm based on traditional signal processing has limited effectiveness and cannot specifically separate audio tracks from individual sound sources such as vocals and various instruments. The immersive experience of the generated 3D audio is significantly different from that of true 3D audio. Alternatively, based on audio timing detection results, the audio signal can be rotated to a certain extent in three-dimensional space to construct a 3D audio effect. However, this approach does not perform targeted processing for each sound element based on the audio track separation algorithm. For example, when the drum sound rotates, the vocal signal must also rotate, resulting in a monotonous 3D audio effect.

[0042] This disclosure combines a series of technologies such as audio sync detection, audio track separation, and music structure analysis to assign different three-dimensional spatial trajectories to audio tracks such as vocals, drums, and bass, converting traditional stereo audio into 3D audio signals, and finally generating immersive 3D audio content through audio space rendering technology.

[0043] In the following, the methods and apparatus of this disclosure will be described in detail with reference to the accompanying drawings, according to various embodiments of this disclosure.

[0044] Figure 1 This is a flowchart of an audio generation method according to an embodiment of the present disclosure. Figure 1 The audio generation method can be used to convert traditional stereo audio, single-channel audio, etc., into three-dimensional audio.

[0045] The audio generation method according to this disclosure can be executed by any electronic device. The electronic device can be at least one of a smartphone, tablet computer, laptop computer, and desktop computer. The electronic device may have a target application installed for implementing the three-dimensional audio generation method of this disclosure.

[0046] Reference Figure 1 In step S101, the audio signal to be processed is acquired. Here, the audio signal to be processed can be, for example, a stereo music signal, a single-channel music signal, etc.

[0047] In step S102, multiple audio track signals for multiple sound sources are obtained by performing audio track separation on the audio signal to be processed.

[0048] A deep learning-based audio track separation system can be used to separate audio tracks from the audio signal being processed. Deep learning-based audio track separation technology can separate vocals and individual instrument signals with higher fidelity. For example, a deep learning-based audio track separation system can consist of three parts: an encoder, a separation module, and a decoder. First, a short-time Fourier transform is performed on the audio signal to obtain its spectral data. This spectral data is then input into the encoder to obtain the encoded features of the audio signal. The separation module further extracts the target audio track features from the encoded features. Finally, based on the target audio track features, the decoder obtains the target masking matrix corresponding to the target audio track signal. After multiplying the spectral data with the target masking matrix, a short-time inverse Fourier transform is performed on the result to obtain the target audio track signal. Here, the target audio track signal may include one or more audio track signals. For example, after performing the above process on a traditional stereo music signal, multiple audio track signals such as vocals, drums, bass, and other instruments can be obtained.

[0049] In step S103, the user's orientation information in three-dimensional space is determined, and three-dimensional spatial metadata for each of the separated audio track signals is generated. The user orientation information may include at least a portion of the virtual user's three-dimensional position coordinates, direction, speed, and movement path at each point in time or during each period in three-dimensional space. For example, the user's orientation information may be fixed at the center position in three-dimensional space. The three-dimensional spatial metadata for each sound source can be determined by referring to the user orientation information. The three-dimensional spatial metadata of an audio track signal can be used to indicate how the corresponding sound source changes in three-dimensional space, such as how the sound source moves in three-dimensional space, or the change in the sound width of the sound source in three-dimensional space.

[0050] The three-dimensional spatial metadata may include the location information and sound width information of the corresponding sound source in three-dimensional space. Here, the location information of the sound source may include part or all of the three-dimensional position coordinates, direction, velocity, and movement path of the corresponding sound source at each time point or each time period in three-dimensional space.

[0051] Figure 2 and Figure 3 A schematic diagram of a virtual three-dimensional space according to an embodiment of the present disclosure is shown.

[0052] Suppose there exists a virtual three-dimensional space with three coordinate axes: X-axis, Y-axis, and Z-axis, where the Z-axis represents the height. Each coordinate axis is set to a range of [0,1]. Figure 2 As shown.

[0053] The user / listener's center position in three-dimensional space, such as (0.5, 0.5, 0.5). Figure 3The white dots in the image can represent various sound sources, such as singing voices and musical instruments. The coordinates of the dots represent the three-dimensional spatial location of the sound sources, and the size of the dots represents the width of the sound. However, the above example is merely illustrative, and this disclosure is not limited thereto. Furthermore, the orientation of the sound sources in the three-dimensional spatial metadata can be relative to the user's orientation.

[0054] According to embodiments of this disclosure, three-dimensional spatial metadata for each audio track signal included in the audio signal to be processed can be generated based on feature information of the audio signal to be processed. The feature information may include at least one of tempo information and structural information of the audio signal to be processed.

[0055] Beat information can be obtained by performing audio beat detection on the audio signal being processed. For example, audio beat detection can be implemented using a deep learning-based model, which can mainly include a feature extraction module, a deep learning-based probability prediction module, and a global beat position estimation module. Deep learning-based beat detection can accurately determine the beat position of the audio signal.

[0056] Feature extraction typically uses frequency domain features. For example, the Mel spectrum and its first-order difference of the audio signal to be processed can be used as input features to the feature extraction module, and then the extracted features are input to the probability prediction module. The probability prediction module can be implemented using deep networks such as CRNN to learn the local and temporal features of the audio signal to be processed. Through the probability prediction module, the probability of whether each frame of audio data is a beat point can be calculated. Finally, based on the predicted probability, the globally optimal beat position is obtained using the global beat position estimation module. The global beat position estimation module can be implemented using a dynamic programming algorithm. The generated beat positions can include both normal beats and repeats. The above audio beat detection process is merely exemplary, and this disclosure is not limited thereto.

[0057] The structural information of an audio signal can be obtained by performing audio structure analysis. This structural information can include the type and timing information of each audio segment. For example, audio structure analysis technology refers to segmenting an audio signal into different types of segments, such as an overture, verse, chorus, and transition. The audio structure analysis process mainly includes several steps, such as segmentation, clustering, and identification. Taking a music signal as an example, the input music signal is first processed by frame segmentation, and the spectral features of the speech frames (such as Mel-frequency cepstral coefficients) are extracted. By calculating the feature correlation between frames, a correlation matrix of the music signal can be obtained. Based on the correlation matrix, the speech signal can be segmented into multiple segments, and these segments can be clustered according to the correlation between them. After the segmentation and clustering process, for the input music signal, a music structure such as ABC and the corresponding time points of the segments can be obtained. Finally, based on the repetition frequency of each segment and acoustic features such as volume and brightness, the verse, chorus, etc., in the music signal can be identified. The above-described audio structure analysis process is merely exemplary, and this disclosure is not limited thereto.

[0058] After obtaining the tempo information and structural information of each audio segment of the audio signal to be processed, corresponding three-dimensional spatial metadata can be generated for different audio track signals.

[0059] For the human voice signal in the separated multiple audio track signals, the position adjustment information of the sound source corresponding to the human voice signal relative to the user's orientation information in three-dimensional space can be determined according to the type of each audio segment in the structural information, and the three-dimensional spatial metadata of the human voice signal can be determined based on the position adjustment information.

[0060] As an example, for a human voice signal in an audio track signal, the sound source corresponding to the human voice signal can be set to move towards the user's location in three-dimensional space during an audio segment of the first preset type, and the information related to the movement can be used as the three-dimensional space metadata of the human voice signal.

[0061] Taking music signals as an example, for vocal signals, the distance between the sound source corresponding to the vocal signal and the listener can be gradually increased in three-dimensional space during the verse. For example, the sound source in... Figure 2 In the three-dimensional space, the coordinates change from (0.5, 0, 0.5) to (0.5, 0.3, 0.5), gradually approaching the listener, thereby increasing the sense of immersion in the music.

[0062] As another example, for a human voice signal in an audio track signal, during an audio segment of the second preset type, the height coordinate of the sound source corresponding to the human voice signal in three-dimensional space can be increased to a predetermined height and the sound width of the sound source can be increased to a predetermined sound width, and the information related to the increase to the predetermined height and the predetermined sound width can be used as three-dimensional space metadata of the human voice signal.

[0063] Taking music signals as an example, for vocal signals, during the transition between the verse and chorus, the height and width of the sound source in three-dimensional space can be gradually increased. This allows the sound source to reach a certain height and width in the chorus, thus enhancing the impact of the chorus. For example, the sound source's position coordinates are... Figure 2 In three-dimensional space, the height value changes from (0.5, 0, 0.5) to (0.5, 0, 1), and the width value changes from 0.05 to 0.5. The changes in height and width values ​​can follow preset linear or non-linear functions.

[0064] For the instrument signals in the separated multiple audio tracks, the movement information of the sound source corresponding to the instrument signal in three-dimensional space can be determined according to the type of each audio segment in the beat information and structure information, and the three-dimensional spatial metadata of the instrument signal can be determined based on the movement information.

[0065] As an example, for instrument signals in an audio track, the sound source corresponding to the instrument signal can be set to move periodically along a predetermined trajectory in three-dimensional space based on the beat information, and the information related to the movement can be used as the three-dimensional spatial metadata of the instrument signal. Taking music signals as an example, for sound sources corresponding to instruments such as drums and bass, they can change periodically along a specific trajectory in three-dimensional space according to the beat. For example, the instrument sound source can be made to rotate along a specific trajectory in three-dimensional space, and on the downbeat, the sound source can be positioned in the center of the listener's head to further enhance the listener's auditory experience.

[0066] As another example, for instrument signals in an audio track signal, the rotational speed of the sound source corresponding to the instrument signal in three-dimensional space can be increased to a predetermined rotational speed during a second preset type of audio segment, and the information related to the increase to the predetermined rotational speed can be used as the three-dimensional spatial metadata of the instrument signal. Taking music signals as an example, for sound sources corresponding to instruments such as drums and bass, the spatial rotational speed of the instrument sound source in the chorus can be increased according to the characteristics of the verse and chorus, thereby enhancing the dynamism of the chorus.

[0067] According to another embodiment of this disclosure, a preset template corresponding to each separated audio track signal can be determined from a plurality of preset templates, and three-dimensional spatial metadata for each audio track signal can be generated using the determined preset templates.

[0068] Each preset template may include at least one of the following: the sound source's trajectory information in three-dimensional space, its speed information, and its sound width variation information. The preset template can pre-define the sound source's changes in three-dimensional space, such as how the sound source moves and how its sound width changes. After separating the audio signal to be processed into multiple audio tracks, a preset template can be assigned to each separated audio track, causing the audio track signal to change according to the information in the corresponding template. For example, based on the characteristics of each audio track signal, a preset template matching the audio track signal can be determined from multiple preset templates.

[0069] As an example, preset templates may include templates that gradually move the distance between the vocal source and the listener from far to near in the verse, gradually increase the height of the vocal source in three-dimensional coordinates and the overall width of the sound in the transition between the verse and chorus, periodically change the instrumental source in three-dimensional space along a specific trajectory according to the beat, and increase the spatial rotation speed of the instrumental source in the chorus. However, the above examples are merely illustrative, and this disclosure is not limited thereto. By applying corresponding preset templates to different audio track signals, three-dimensional spatial metadata that better matches the attributes of each audio track signal is generated, enabling the generation of more realistic three-dimensional audio signals.

[0070] According to another embodiment of this disclosure, user-inputted setting information can be obtained. The setting information may include at least one of the following: the movement trajectory, movement speed, and sound width variation value of each sound source corresponding to multiple audio track signals in three-dimensional space. Then, three-dimensional spatial metadata for each audio track signal is generated based on the setting information.

[0071] As an example, three-dimensional spatial metadata can be generated for each separated audio track signal based on user input. User input can be used to set at least one of the following: the movement trajectory, movement speed, and sound width variation value of each sound source corresponding to multiple audio track signals in three-dimensional space. Users can customize the variations for each audio track signal according to their understanding of the audio content to be processed.

[0072] Taking music signals as an example, users can customize the movement trajectory of the bass sound source in three-dimensional space during the chorus, based on their understanding of the music content. When generating a three-dimensional music signal, the user-defined movement trajectory can be applied to the bass sound source. However, the above example is merely illustrative, and this disclosure is not limited thereto. Users can set the changes of each sound source in three-dimensional space according to their preferences and understanding of the music content to obtain their desired music effects, thus satisfying user needs and improving the user experience.

[0073] According to another example of this disclosure, three-dimensional spatial metadata for multiple separated audio track signals can be generated using a preset template based on the beat information and various audio segments of the audio signal to be processed. In other words, the sound source corresponding to each audio track signal can be modified in three-dimensional space according to the preset template at the corresponding beat position and segment portion.

[0074] In step S104, based on the user's location information, the separated multiple audio track signals, and the three-dimensional spatial metadata corresponding to each audio track signal, a three-dimensional audio signal corresponding to the audio signal to be processed is generated. The generated three-dimensional audio signal may include the separated audio track signals and the three-dimensional spatial metadata corresponding to each audio track signal, and the final 3D audio effect can be obtained through spatial audio rendering technology.

[0075] According to embodiments of this disclosure, when generating a three-dimensional audio signal, the type of playback device used to play the three-dimensional audio signal can be considered, and different methods can be used to render the three-dimensional audio signal for different playback devices.

[0076] Specifically, the type of playback device used to play the 3D audio signal can be identified, the rendering strategy corresponding to the type of playback device can be obtained, and then, based on the user's location information, the separated multiple audio track signals and the 3D spatial metadata corresponding to each audio track signal, the 3D audio signal corresponding to the audio signal to be processed can be generated through the obtained rendering strategy.

[0077] For in-ear playback devices, for each separated audio track signal, a three-dimensional audio signal corresponding to the audio track signal can be generated based on the orientation information corresponding to each audio frame of the audio track signal and the user's orientation information. Furthermore, the sound width of the three-dimensional audio signal corresponding to the audio track signal can be adjusted based on the sound width information of each audio frame. For example, based on the Head Relevance Transfer Function (HRTF), each audio frame of the audio track signal can be convolved with the HRTF corresponding to its three-dimensional coordinates to obtain the corresponding binaural audio signal.

[0078] For external playback devices, for each audio track signal, rendering can be performed based on the directional information of the sound source corresponding to the audio track signal and the directional information of multiple speakers to generate a three-dimensional audio signal corresponding to the audio track signal. Furthermore, the sound width of the three-dimensional audio signal corresponding to the audio track signal can be adjusted based on the sound width information of the sound source. For example, based on Vector Basis Amplitude Phase Shift (VBAP) technology, the unit direction vector of sound can be represented as a linear combination of the unit direction vectors of the multiple speakers closest to the sound source direction. The gain factor of each speaker can be calculated, thereby rendering a 3D music effect for multiple speakers.

[0079] Taking into account the type of playback device, a three-dimensional audio signal can be generated to improve the playback effect of each type of playback device.

[0080] The 3D audio generation method proposed in this disclosure can automatically convert traditional two-channel stereo music into 3D music, increasing the immersive experience of the music.

[0081] Figure 4 This is a schematic flowchart of an audio generation method according to another embodiment of the present disclosure. Figure 4 The example described is the conversion of stereo / mono music into 3D music. However, Figure 4 The system shown can also be used to convert any form of audio into 3D audio.

[0082] The input single-channel or stereo music is processed by a track separation module to extract various pre-defined audio track signals, such as audio signals from vocals, drums, and bass. Simultaneously, the input single-channel or stereo music undergoes a beat detection module to extract the music's beat information, and an audio structure analysis module extracts the structural information of musical segments such as verses and choruses. Based on this beat and structure information, and according to custom templates or user-edited data, a 3D metadata generation module determines the 3D spatial metadata for each audio track signal. The spatial audio rendering module then uses spatial audio rendering technology to obtain the final 3D music signal based on the separated audio track signals and their corresponding 3D spatial metadata.

[0083] The audio track separation module can be implemented using deep learning, and may include an encoder, decoder, and splitter. The input time-domain music signal is processed by an STFT module to obtain the corresponding music spectrum signal. After passing through an encoder consisting of multiple convolutional layers, a splitter further extracts the audio track features. Finally, the decoder obtains the target masking matrix corresponding to the target audio track signal. Multiplying the music spectrum signal by the target masking matrix and then passing it through an ISTFT module yields the target audio track signal, such as vocals, drums, bass, and other instruments.

[0084] The audio beat detection module can be implemented using deep learning, and may include a feature extraction module, a deep model-based probability prediction module, and a global beat position estimation module. First, feature extraction typically uses frequency domain features; in one implementation, the Mel spectrum and its first difference are often used as input features. The probability prediction module usually uses deep networks such as CRNNs to learn local and temporal features, calculating the probability of each frame of audio data being a beat. Finally, based on the predicted probability, the global beat position estimation module uses a dynamic programming algorithm to obtain the globally optimal beat position. The generated beat position can include both normal beats and repeats.

[0085] The music structure analysis module uses algorithms to segment music signals into different segments, such as overtures, verses, choruses, and transitions. The music structure analysis process mainly includes segmentation, clustering, and identification. First, the music signal is processed by frame segmentation, and spectral features such as Mel-frequency cepstral coefficients (MFCCs) of the speech frames are extracted. By calculating the feature correlations between frames, a correlation matrix of the music signal can be obtained. Based on the correlation matrix, the speech signal can be segmented into multiple segments, and these segments can be clustered according to their correlations. After segmentation and clustering, for the music signal to be processed, a music structure similar to ABCBC and corresponding segment time points can be obtained. Finally, based on the repetition frequency of segments and acoustic features such as volume and brightness, the verses and choruses in the music can be identified.

[0086] The 3D metadata generation module can generate corresponding 3D spatial metadata for different audio track signals based on music beat information and music structure information. 3D spatial metadata corresponds to the information of the audio track signal in 3D space, specifically including 3D position coordinates and sound width, etc.

[0087] Users can customize templates for each sound source based on their understanding of the music content, or they can automatically generate three-dimensional spatial metadata for each sound source using some pre-set templates.

[0088] For example, pre-defined templates may include those that gradually move the distance between the vocal source and the listener from far to near in the verse, gradually increase the height of the vocal source in three-dimensional coordinates and the overall width of the sound in the transition between the verse and chorus, periodically change the instrumental source in three-dimensional space along a specific trajectory according to the beat, and increase the spatial rotation speed of the instrumental source in the chorus. However, the above examples are merely illustrative, and this disclosure is not limited thereto.

[0089] The 3D music signal generated by the spatial audio rendering module can include separate audio track signals and three-dimensional metadata corresponding to each track. The final 3D music effect can be obtained through spatial audio rendering technology. For example, for headphone playback devices, based on the Head Relevance Transfer Function (HRTF), each audio frame of the input track can be convolved with its corresponding HRTF in three-dimensional coordinates to obtain the corresponding binaural audio signal. For multi-speaker playback devices, based on Vector Amplitude Translation (VBAP) technology, the unit direction vector of sound can be represented as a linear combination of the unit direction vectors of the multiple speakers closest to the sound source, and the gain factor of each speaker can be calculated, thus rendering a 3D music effect for multiple speakers.

[0090] Figure 5This is a schematic diagram of the structure of an audio generation device in the hardware operating environment of this disclosure embodiment.

[0091] like Figure 5 As shown, the audio generation device 500 may include: a processing component 501, a communication bus 502, a network interface 503, an input / output interface 504, a memory 505, and a power supply component 506. The communication bus 502 is used to enable communication between these components. The input / output interface 504 may include a video display (such as a liquid crystal display), a microphone and speakers, and a user interaction interface (such as a keyboard, mouse, touch input device, etc.). Optionally, the input / output interface 504 may also include a standard wired interface or a wireless interface. The network interface 503 may optionally include a standard wired interface or a wireless interface (such as a Wi-Fi interface). The memory 505 may be a high-speed random access memory or a stable non-volatile memory. The memory 505 may also optionally be a storage device independent of the aforementioned processing component 501.

[0092] Those skilled in the art will understand that Figure 5 The structure shown does not constitute a limitation on the audio generation device 500, and may include more or fewer components than shown, or combine certain components, or have different component arrangements.

[0093] like Figure 5 As shown, the memory 505, which serves as a storage medium, may include an operating system (such as a MAC operating system), a data storage module, a network communication module, a user interface module, programs, and a database.

[0094] exist Figure 5 In the audio generation device 500 shown, the network interface 503 is mainly used for data communication with external electronic devices / terminals; the input / output interface 504 is mainly used for data interaction with users; the processing component 501 and the memory 505 in the audio generation device 500 can be set in the audio generation device 500. The audio generation device 500 calls the program stored in the memory 505 and various APIs provided by the operating system through the processing component 501 to execute the audio generation method provided in the embodiments of this disclosure.

[0095] Processing component 501 may include at least one processor, and memory 505 stores a set of computer-executable instructions that, when executed by the at least one processor, perform an audio generation method according to embodiments of the present disclosure. However, the above examples are merely illustrative, and the present disclosure is not limited thereto.

[0096] The processing component 501 can control the components included in the audio generation device 500 by executing a program.

[0097] As an example, the audio generating device 500 may be a PC, tablet, personal digital assistant, smartphone, or other device capable of executing the aforementioned set of instructions. Here, the audio generating device 500 is not necessarily a single electronic device, but may be a collection of any devices or circuits capable of executing the aforementioned instructions (or instruction sets) individually or in combination. The audio generating device 500 may also be part of an integrated control system or system manager, or may be configured to interface with a portable electronic device locally or remotely (e.g., via wireless transmission).

[0098] In the audio generation device 500, the processing component 501 may include a central processing unit (CPU), a graphics processing unit (GPU), a programmable logic device, a dedicated processor system, a microcontroller, or a microprocessor. By way of example and not limitation, the processing component 501 may also include an analog processor, a digital processor, a microprocessor, a multi-core processor, a processor array, a network processor, etc.

[0099] Processing component 501 can execute instructions or code stored in memory, wherein memory 505 can also store data. Instructions and data can also be sent and received over a network via network interface 503, wherein network interface 503 can employ any known transport protocol.

[0100] The memory 505 can be integrated with the processing component 501, for example, by placing RAM or flash memory within an integrated circuit microprocessor. Alternatively, the memory 505 can include a separate device, such as an external disk drive, a storage array, or other storage device that can be used by any database system. The memory and processing component 501 can be operatively coupled, or can communicate with each other, for example, via I / O ports, network connections, etc., enabling the processing component 501 to read data stored in the memory 505.

[0101] Figure 6 This is a block diagram of an audio generation apparatus according to an embodiment of the present disclosure.

[0102] Reference Figure 6 The audio generation apparatus 600 may include an acquisition module 601, an audio track separation module 602, a metadata generation module 603, and a rendering module 604. Each module in the audio generation apparatus 600 may be implemented by one or more modules, and the names of the corresponding modules may vary depending on the type of module. In various embodiments, some modules in the audio generation apparatus 600 may be omitted, or additional modules may be included. Furthermore, modules / elements according to various embodiments of this disclosure may be combined to form a single entity, and thus perform the functions of the respective modules / elements equivalently prior to the combination.

[0103] The acquisition module 601 can acquire the audio signal to be processed.

[0104] The audio track separation module 602 can obtain multiple audio track signals for multiple sound sources by separating the audio signals to be processed.

[0105] The metadata generation module 603 can determine the user's location information in three-dimensional space and generate three-dimensional space metadata corresponding to each of the separated multiple audio track signals. The three-dimensional space metadata may include the location information and sound width information of the corresponding sound source in three-dimensional space.

[0106] Optionally, the metadata generation module 603 can acquire feature information of the audio signal to be processed, wherein the feature information may include at least one of the tempo information of the audio signal to be processed and the structural information of each audio segment; and generate three-dimensional spatial metadata for each audio track signal based on the feature information.

[0107] Optionally, the metadata generation module 603 can obtain beat information by performing audio beat detection on the audio signal to be processed; and obtain structural information of each audio segment by performing audio structure analysis on the audio signal to be processed.

[0108] Optionally, the metadata generation module 603 can determine the position adjustment information of the sound source corresponding to the human voice signal relative to the user's orientation information in three-dimensional space based on the type of each audio segment in the structural information, and determine the three-dimensional spatial metadata of the human voice signal based on the position adjustment information.

[0109] For example, the metadata generation module 603 can, for human voice signals in multiple audio track signals, set the sound source corresponding to the human voice signal to move towards the user's location in three-dimensional space during an audio segment of a first preset type, and use the information related to the movement as the three-dimensional spatial metadata of the human voice signal.

[0110] For example, the metadata generation module 603 can, for human voice signals in multiple audio track signals, increase the height coordinates of the sound source corresponding to the human voice signal in three-dimensional space to a predetermined height and increase the sound width of the sound source to a predetermined sound width during an audio segment of a second preset type, and use the information related to the increase to the predetermined height and the predetermined sound width as the three-dimensional space metadata of the human voice signal.

[0111] Optionally, the metadata generation module 603 can determine the movement information of the sound source corresponding to the instrument signal in three-dimensional space based on the type of each audio segment in the beat information and structure information, and determine the three-dimensional space metadata of the instrument signal based on the movement information.

[0112] For example, the metadata generation module 603 can, for instrument signals in multiple audio tracks, set the sound source corresponding to the instrument signal to move periodically in three-dimensional space according to a predetermined trajectory based on the beat information, and use the information related to the movement as the three-dimensional spatial metadata of the instrument signal.

[0113] For example, the metadata generation module 603 can increase the rotation speed of the sound source corresponding to the instrument signal in three-dimensional space to a predetermined rotation speed during a second preset type of audio segment for the instrument signal in multiple audio track signals, and use the information related to the increase to the predetermined rotation speed as the three-dimensional space metadata of the instrument signal.

[0114] Optionally, the metadata generation module 603 can determine a preset template corresponding to each audio track signal from a plurality of preset templates, wherein the preset template includes at least one of the following: the movement trajectory information of the sound source in three-dimensional space, the movement speed information, and the sound width change information; and use the determined preset template to generate three-dimensional space metadata for each audio track signal.

[0115] Optionally, the metadata generation module 603 can generate three-dimensional spatial metadata for multiple audio track signals based on the feature information of the audio signal to be processed and according to a predetermined preset template.

[0116] Optionally, the metadata generation module 603 can obtain user-inputted setting information, wherein the setting information includes at least one of the movement trajectory, movement speed and sound width change value of each sound source corresponding to multiple audio track signals in three-dimensional space, and generates three-dimensional space metadata for each audio track signal based on the setting information.

[0117] For example, the metadata generation module 603 can receive user input, wherein the user input is used to set at least one of the following: the movement trajectory, movement speed, and sound width variation value of each sound source corresponding to multiple audio track signals in three-dimensional space; and generate three-dimensional space metadata for multiple audio track signals based on the user input.

[0118] The rendering module 604 can generate a three-dimensional audio signal corresponding to the audio signal to be processed based on the user's location information, the separated multiple audio track signals and the three-dimensional spatial metadata corresponding to each audio track signal.

[0119] The rendering module 604 can identify the type of playback device used to play the three-dimensional audio signal, obtain the rendering strategy corresponding to the type of playback device, and generate a three-dimensional audio signal corresponding to the audio signal to be processed based on the user's orientation information, the separated multiple audio track signals and the three-dimensional spatial metadata corresponding to each audio track signal.

[0120] For in-ear playback devices, the rendering module 604 can generate a three-dimensional audio signal corresponding to each of the multiple audio track signals based on the orientation information corresponding to each audio frame of the audio track signal and the user's orientation information; and adjust the sound width of the three-dimensional audio signal corresponding to the audio track signal based on the sound width information of each audio frame.

[0121] For external playback devices, the rendering module 604 can render each audio track signal based on the location information of the sound source corresponding to the audio track signal and the location information of multiple speakers to generate a three-dimensional audio signal corresponding to the audio track signal; and adjust the sound width of the three-dimensional audio signal corresponding to the audio track signal based on the sound width information of the sound source.

[0122] The above has been based on Figures 1 to 4 The method of converting traditional stereo signals into three-dimensional audio signals has been described in detail, and will not be described again here.

[0123] According to embodiments of this disclosure, an electronic device may be provided. Figure 7 This is a block diagram of an electronic device 700 according to an embodiment of the present disclosure. The electronic device 700 may include at least one memory 702 and at least one processor 701. The at least one memory 702 stores a set of computer-executable instructions. When the set of computer-executable instructions is executed by the at least one processor 701, an audio generation method according to an embodiment of the present disclosure is performed.

[0124] Processor 701 may include a central processing unit (CPU), a graphics processing unit (GPU), a programmable logic device, a dedicated processor system, a microcontroller, or a microprocessor. By way of example and not limitation, processor 701 may also include an analog processor, a digital processor, a microprocessor, a multi-core processor, a processor array, a network processor, etc.

[0125] The memory 702, which serves as a storage medium, may include an operating system, a data storage module, a network communication module, a user interface module, a program for executing the audio generation method of this disclosure, and a database.

[0126] The memory 702 can be integrated with the processor 701; for example, RAM or flash memory can be arranged within an integrated circuit microprocessor. Alternatively, the memory 702 can include a separate device, such as an external disk drive, a storage array, or other storage device that can be used by any database system. The memory 702 and the processor 701 can be operatively coupled, or can communicate with each other, for example, via I / O ports, network connections, etc., enabling the processor 701 to read files stored in the memory 702.

[0127] In addition, the electronic device 700 may also include a video display (such as a liquid crystal display) and a user interaction interface (such as a keyboard, mouse, touch input device, etc.). All components of the electronic device 700 can be interconnected via a bus and / or network.

[0128] As will be understood by those skilled in the art, Figure 7 The structure shown does not constitute a limitation on the structure and may include more or fewer components than shown, or combine certain components, or have different component arrangements.

[0129] According to embodiments of this disclosure, a computer-readable storage medium storing instructions may also be provided, wherein when the instructions are executed by at least one processor, they cause at least one processor to perform an audio generation method according to this disclosure. Examples of computer-readable storage media herein include: read-only memory (ROM), random access programmable read-only memory (PROM), electrically erasable programmable read-only memory (EEPROM), random access memory (RAM), dynamic random access memory (DRAM), static random access memory (SRAM), flash memory, non-volatile memory, CD-ROM, CD-R, CD+R, CD-RW, CD+RW, DVD-ROM, DVD-R, DVD+R, DVD-RW, DVD+RW, DVD-RAM, BD-ROM, BD-R, BD-R LTH, BD-RE, Blu-ray or optical disc storage, hard disk drive (HDD), solid-state drive (SSD), card storage (such as multimedia cards, secure digital (SD) cards, or ultra-fast digital (XD) cards), magnetic tape, floppy disk, magneto-optical data storage device, optical data storage device, hard disk, solid-state drive, and any other device configured to store a computer program and any associated data, data files, and data structures in a non-transitory manner and to provide the computer program and any associated data, data files, and data structures to a processor or computer so that the processor or computer can execute the computer program. The computer program in the aforementioned computer-readable storage medium can run in an environment deployed in computer devices such as clients, hosts, agent devices, servers, etc. Furthermore, in one example, the computer program and any associated data, data files, and data structures are distributed across a networked computer system, such that the computer program and any associated data, data files, and data structures are stored, accessed, and executed in a distributed manner through one or more processors or computers.

[0130] According to embodiments of this disclosure, a computer program product may also be provided, wherein the instructions in the computer program product can be executed by the processor of a computer device to complete the above-described audio generation method.

[0131] Other embodiments of this disclosure will readily occur to those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of this disclosure that follow the general principles of this disclosure and include common knowledge or customary techniques in the art not disclosed herein. The specification and examples are to be considered exemplary only, and the true scope and spirit of this disclosure are indicated by the following claims.

[0132] It should be understood that this disclosure is not limited to the precise structures described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of this disclosure is limited only by the appended claims.

Claims

1. An audio generation method, characterized by, The method comprises: obtaining a to-be-processed audio signal; obtaining a plurality of track signals corresponding to a plurality of sound sources by performing track separation on the to-be-processed audio signal; determining user orientation information in a three-dimensional space and generating three-dimensional space metadata corresponding to each of the plurality of track signals, wherein the three-dimensional space metadata comprises orientation information and sound width information of the corresponding sound source in the three-dimensional space; generating a three-dimensional audio signal corresponding to the to-be-processed audio signal based on the user orientation information, the plurality of separated track signals, and the three-dimensional space metadata corresponding to each of the plurality of track signals, wherein generating three-dimensional space metadata corresponding to each of the plurality of track signals comprises: obtaining feature information of the to-be-processed audio signal, wherein the feature information comprises at least one of beat information and structure information of the to-be-processed audio signal, and the structure information comprises type information of each audio segment of the to-be-processed audio signal; generating three-dimensional space metadata corresponding to each of the plurality of track signals based on the feature information, wherein generating three-dimensional space metadata corresponding to each of the plurality of track signals based on the feature information comprises: for an instrument signal in the plurality of track signals, determining movement information of a sound source corresponding to the instrument signal in the three-dimensional space according to the type of each audio segment in the beat information and the structure information, and determining three-dimensional space metadata of the instrument signal based on the movement information, wherein the movement information comprises a movement trajectory and a movement speed; for a vocal signal in the plurality of track signals, determining position adjustment information of a sound source corresponding to the vocal signal relative to the user orientation information in the three-dimensional space according to the type of each audio segment in the structure information, and determining three-dimensional space metadata of the vocal signal based on the position adjustment information.

2. The method of claim 1, wherein, Generating three-dimensional space metadata corresponding to each of the plurality of track signals further comprises: determining a preset template corresponding to each of the plurality of track signals from a plurality of preset templates, wherein the preset template comprises at least one of movement trajectory information, movement speed information, and sound width change information of a corresponding sound source in a three-dimensional space; generating three-dimensional space metadata for each of the plurality of track signals using the determined preset template.

3. The method of claim 1, wherein, Generating three-dimensional space metadata corresponding to each of the plurality of track signals further comprises: obtaining setting information input by a user, wherein the setting information comprises at least one of a movement trajectory, a movement speed, and a sound width change value of each sound source corresponding to the plurality of track signals in a three-dimensional space; generating three-dimensional space metadata for each of the plurality of track signals based on the setting information.

4. The method of claim 1, wherein, Generating a three-dimensional audio signal corresponding to the to-be-processed audio signal based on the user orientation information, the plurality of separated track signals, and the three-dimensional space metadata corresponding to each of the plurality of track signals comprises: identifying a type of a playback device for playing the three-dimensional audio signal; obtain a rendering strategy corresponding to a type of the playback device, and generate a three-dimensional audio signal corresponding to the to-be-processed audio signal based on the user orientation information, the separated plurality of audio track signals, and the three-dimensional space metadata corresponding to each of the audio track signals and the rendering strategy.

5. The method of claim 4, wherein, The generating of the three-dimensional audio signal corresponding to the to-be-processed audio signal based on the rendering strategy comprises: when the playback device is an in-ear playback device, for each of the audio track signals, generating a three-dimensional audio signal corresponding to the audio track signal based on orientation information corresponding to each audio frame of the audio track signal and the user orientation information; adjusting a sound width of the three-dimensional audio signal corresponding to the audio track signal based on the sound width information of the audio frame.

6. The method of claim 4, wherein, The generating of the three-dimensional audio signal corresponding to the to-be-processed audio signal based on the rendering strategy comprises: when the playback device is an in-ear playback device, for each of the audio track signals, generating a three-dimensional audio signal corresponding to the audio track signal based on orientation information corresponding to each audio frame of the audio track signal and the user orientation information; adjusting a sound width of the three-dimensional audio signal corresponding to the audio track signal based on the sound width information of the audio frame.

7. An audio generating apparatus, characterized by comprising: The apparatus comprises: an obtaining module configured to obtain a to-be-processed audio signal; an audio track separation module configured to obtain a plurality of audio track signals for a plurality of sound sources by performing audio track separation on the to-be-processed audio signal; a metadata generation module configured to determine user orientation information in a three-dimensional space and generate three-dimensional space metadata corresponding to each of the plurality of audio track signals, wherein the three-dimensional space metadata comprises orientation information and sound width information of a corresponding sound source in the three-dimensional space; a rendering module configured to generate a three-dimensional audio signal corresponding to the to-be-processed audio signal based on the user orientation information, the separated plurality of audio track signals, and the three-dimensional space metadata corresponding to each of the audio track signals, wherein the metadata generation module is configured to: obtain feature information of the to-be-processed audio signal, wherein the feature information comprises at least one of beat information and structure information of the to-be-processed audio signal, and the structure information comprises type information of each audio segment of the to-be-processed audio signal; generate the three-dimensional space metadata corresponding to each of the audio track signals based on the feature information, respectively, The metadata generation module is configured to: for an instrument signal in the plurality of audio track signals, determine movement information of a sound source corresponding to the instrument signal in the three-dimensional space according to the beat information and a type of each audio segment in the structure information, determine three-dimensional space metadata of the instrument signal based on the movement information, wherein the movement information includes a movement trajectory and a movement speed; for a vocal signal in the plurality of audio track signals, determine position adjustment information of a sound source corresponding to the vocal signal relative to the user orientation information in the three-dimensional space according to a type of each audio segment in the structure information, and determine three-dimensional space metadata of the vocal signal based on the position adjustment information.

8. The apparatus of claim 7, wherein, The metadata generation module is configured to: determine a preset template corresponding to each of the audio track signals from a plurality of preset templates, wherein the preset template includes at least one of movement trajectory information, movement speed information, and sound width change information of a corresponding sound source in the three-dimensional space; generate three-dimensional space metadata for each of the audio track signals using the determined preset template.

9. The apparatus of claim 7, wherein, The metadata generation module is configured to: obtain setting information input by a user, wherein the setting information includes at least one of a movement trajectory, a movement speed, and a sound width change value of each sound source corresponding to the plurality of audio track signals in the three-dimensional space; generate three-dimensional space metadata for each of the audio track signals based on the setting information.

10. The apparatus of claim 7, wherein, The rendering module is configured to: identify a type of a playback device used to play the three-dimensional audio signal; obtain a rendering strategy corresponding to the type of the playback device, and generate a three-dimensional audio signal corresponding to the audio signal to be processed by the rendering strategy based on the user orientation information, the separated plurality of audio track signals, and the three-dimensional space metadata corresponding to each of the audio track signals.

11. The apparatus of claim 10, wherein, The rendering module is configured to: when the playback device is an in-ear playback device, generate a three-dimensional audio signal corresponding to each of the audio track signals based on orientation information corresponding to each audio frame of the audio track signal and the user orientation information; adjust a sound width of the three-dimensional audio signal corresponding to the audio track signal based on sound width information of the audio frame.

12. The apparatus of claim 10, wherein, The rendering module is configured to: when the playback device is an external playback device, render the audio track signal based on orientation information of a sound source corresponding to the audio track signal and orientation information of a plurality of loudspeakers to generate a three-dimensional audio signal corresponding to the audio track signal; adjust a sound width of the three-dimensional audio signal corresponding to the audio track signal based on sound width information of the sound source.

13. An electronic device, comprising: comprise: at least one processor; at least one memory storing computer executable instructions, wherein the computer executable instructions, when executed by the at least one processor, cause the at least one processor to perform the audio generation method according to any one of claims 1 to 6.

14. A computer-readable storage medium storing instructions, wherein, When the instructions are executed by at least one processor, the at least one processor is caused to perform the audio generation method as claimed in any one of claims 1 to 6.

15. A computer program product, instructions in the computer program product being executed by at least one processor in an electronic device to perform the audio generation method as claimed in any one of claims 1 to 6.

Citation Information

Patent Citations

  • Three-Dimensional Audio Systems

    CN113784274A