Audio rendering method and related device

By constructing spheres for audio sampling points and simulating changes in sound amplitude, the problem of high computational complexity in HOA encoding of Jingcai Sound was solved, achieving real-time and stable audio rendering and avoiding audio-video synchronization issues.

CN121011202APending Publication Date: 2025-11-25HUNAN HAPPLY SUNSHINE INTERACTIVE ENTERTAINMENT MEDIA CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510969739.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-14
Publication Date
2025-11-25

AI Technical Summary

Technical Problem

In HOA encoding, the sound object occupies the same amount of resources, resulting in high computational complexity, which affects the real-time rendering speed, causing the player to be unable to play in real time, resulting in audio and video desynchronization or playback stuttering.

Method used

By acquiring the audio PCM data, Cartesian coordinates, and spherical coordinates of the audio sampling points of each sound object in the target sound stream, a sphere is constructed between adjacent sampling points to simulate the change in sound amplitude during spatial movement, thus obtaining the binaural rendering signal value at the midpoint.

Benefits of technology

It improves the real-time performance of audio rendering, reduces signal processing time, avoids audio-video desynchronization and playback stuttering, and ensures the real-time rendering performance of multiple sound objects.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121011202A_ABST
    Figure CN121011202A_ABST
Patent Text Reader

Abstract

The invention discloses an audio rendering method and a related device, and relates to the technical field of audio rendering, and the method comprises the steps: for every two adjacent sampling points in each audio sampling point of each sound object in a target sound stream, constructing spheres corresponding to the two sampling points according to the spherical coordinates of the two sampling points, according to the audio PCM data, Cartesian coordinates and spherical coordinates of the two sampling points, simulation calculation is carried out on the change of the sound amplitude of the previous sampling point in the two sampling points in the space moving process by combining the constructed sphere, the sound amplitude of each middle point in the space moving process is obtained, and the sound amplitude of each middle point in the space moving process is calculated according to the sound amplitude of each middle point. Obtaining a binaural rendering signal value of each intermediate point; and obtaining a binaural rendering signal value of the target sound stream. According to the method, the sound amplitude of the intermediate point can be quickly obtained only by constructing the two spheres and performing simple simulation calculation, and the real-time performance of audio rendering is improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of audio rendering, and in particular to an audio rendering method and related device. BACKGROUND

[0002] AudioVivid is the first audio coding standard based on AI (Artificial Intelligence) technology in the world. The standard creatively divides audio content into sound beds, sound objects and Ambisonic information, and encodes all audio content through Higher Order Ambisonics (HOA) and realizes AudioVivid loudspeaker playback and binaural rendering playback through different decoding methods.

[0003] Currently, in the HOA encoding of AudioVivid, sound objects and sound bed channels occupy equal resources, resulting in multiple channels of Ambisonic information after HOA encoding of sound objects, for example, in HOA three-order encoding, 12 sound objects will generate 192 channels of Ambisonic information after HOA encoding, and each channel of Ambisonic information is processed, which has extremely high calculation amount and calculation complexity, greatly affecting the speed of real-time rendering, causing the player to be unable to play in real time, causing audio and video out of synchronization or playing lag and other problems. SUMMARY

[0004] In view of the above problems, the present application provides an audio rendering method and related device to achieve the purpose of real-time audio rendering. The specific scheme is as follows:

[0005] The first aspect of the present application provides an audio rendering method, comprising:

[0006] obtaining audio PCM data, Cartesian coordinates and spherical coordinates of each audio sampling point of each sound object in a target sound stream, wherein the audio PCM data refers to audio pulse code modulation data;

[0007] for each adjacent two sampling points in each audio sampling point of each sound object, constructing a sphere corresponding to each of the two sampling points according to the spherical coordinates of the two sampling points, wherein the sphere corresponding to any sampling point represents the range of the sound amplitude of the sampling point; according to the audio PCM data, Cartesian coordinates and spherical coordinates of the two sampling points, and combining the spheres corresponding to the two sampling points, simulating and calculating the change of sound amplitude in the spatial movement process of the former one of the two sampling points, to obtain the sound amplitude of each intermediate point in the spatial movement process;

[0008] obtaining a binaural rendering signal value of the target sound stream according to the sound amplitude value of each intermediate point.

[0009] In a possible implementation, the simulation calculation of the sound amplitude value of each intermediate point during the spatial movement of the former one of the two sampling points in the corresponding sphere of the two sampling points according to the audio PCM data, the Cartesian coordinates and the spherical coordinates of the two sampling points respectively comprises:

[0010] determining the inclusion relationship of the corresponding spheres of the two sampling points and the positional relationship between the corresponding spheres of the two sampling points and each intermediate point according to the Cartesian coordinates and the spherical coordinates of the two sampling points respectively;

[0011] determining the sound amplitude value of each intermediate point according to the inclusion relationship, the positional relationship between the corresponding spheres of the two sampling points and each intermediate point, and the audio PCM data, the Cartesian coordinates and the spherical coordinates of the two sampling points respectively.

[0012] In a possible implementation, the determination of the inclusion relationship of the corresponding spheres of the two sampling points and the positional relationship between the corresponding spheres of the two sampling points and each intermediate point according to the Cartesian coordinates and the spherical coordinates of the two sampling points respectively comprises:

[0013] determining the sampling point distance between the two sampling points according to the Cartesian coordinates of the two sampling points respectively;

[0014] determining the inclusion relationship of the corresponding spheres of the two sampling points according to the radial distance in the spherical coordinates of the two sampling points and the sampling point distance;

[0015] obtaining the duration of each frame of the target sound stream, and determining the movement distance corresponding to each intermediate point according to the duration;

[0016] determining the positional relationship between the corresponding spheres of the two sampling points and each intermediate point according to the movement distance corresponding to each intermediate point, the radial distance in the spherical coordinates of the two sampling points and the sampling point distance.

[0017] In a possible implementation, the determination of the sound amplitude value of each intermediate point according to the inclusion relationship, the positional relationship between the corresponding spheres of the two sampling points and each intermediate point, and the audio PCM data, the Cartesian coordinates and the spherical coordinates of the two sampling points respectively comprises:

[0018] determine a sound amplitude function of each intermediate point according to the containing relationship and the position relationship between the two sampling points respectively corresponding to the sphere and each intermediate point;

[0019] determine an independent variable in the sound amplitude function of each intermediate point according to the audio PCM data, the Cartesian coordinates and the spherical coordinates of the two sampling points respectively;

[0020] substitute the independent variable in the sound amplitude function of each intermediate point into the sound amplitude function of each intermediate point for function calculation to obtain the sound amplitude of each intermediate point.

[0021] In a possible implementation, the obtaining of the binaural rendering signal value of each intermediate point according to the sound amplitude of each intermediate point comprises:

[0022] perform dynamic interpolation calculation based on the head-related transfer function value for each intermediate point in the space movement process according to the spherical coordinates of the two sampling points respectively to obtain the head-related transfer function value of each intermediate point, wherein the head-related transfer function value of each intermediate point represents the image information of each intermediate point;

[0023] obtain the binaural rendering signal value of each intermediate point according to the sound amplitude of each intermediate point and the head-related transfer function value of each intermediate point.

[0024] In a possible implementation, the performing of dynamic interpolation calculation based on the head-related transfer function value for each intermediate point in the space movement process according to the spherical coordinates of the two sampling points respectively to obtain the head-related transfer function value of each intermediate point comprises:

[0025] calculate the unit vector of the Cartesian coordinates of the two sampling points and the sound frequency impulse value of each of the two sampling points according to the pitch angle and the azimuth angle in the spherical coordinates of the two sampling points respectively;

[0026] determine the included angle between the two sampling points and the origin of the Cartesian coordinates according to the unit vector of the Cartesian coordinates of the two sampling points respectively;

[0027] obtain the duration of each target sound stream and the playing duration corresponding to each intermediate point, and determine the angle change value corresponding to each intermediate point according to the duration, the included angle and the playing duration corresponding to each intermediate point;

[0028] determine the spherical linear interpolation angle of each intermediate point according to the unit vector of the Cartesian coordinates of the two sampling points, the playing duration corresponding to each intermediate point and the angle change value;

[0029] determining a head-related transfer function value of each intermediate point according to the spherical linear interpolation angle of the intermediate point and the sound frequency impulse value of each of the two sampling points.

[0030] In a possible implementation, the determining the head-related transfer function value of each intermediate point according to the spherical linear interpolation angle of the intermediate point and the sound frequency impulse value of each of the two sampling points comprises:

[0031] for the each intermediate point:

[0032] if the spherical linear interpolation angle of the intermediate point is less than a preset first angle value, presetting a target unit time by 1, and determining a head-related transfer function value of a previous intermediate point of the intermediate point as the head-related transfer function value of the intermediate point, wherein if the intermediate point is the first intermediate point after the previous sampling point, the head-related transfer function value of the previous intermediate point of the intermediate point is the sound frequency impulse value of the previous sampling point;

[0033] if the spherical linear interpolation angle of the intermediate point is greater than or equal to the first angle value and less than or equal to a preset second angle value, or the target unit time is equal to a preset count threshold, determining an updated time length corresponding to the intermediate point according to the target unit time and a playing time length corresponding to the intermediate point, determining the head-related transfer function value of the intermediate point according to the updated time length corresponding to the intermediate point, a sound frequency impulse value of a latter sampling point of the two sampling points, and the head-related transfer function value of the previous intermediate point of the intermediate point, and setting the target unit time to 0;

[0034] if the spherical linear interpolation angle of the intermediate point is greater than the second angle value, determining a first point at which the spherical linear interpolation angle is greater than the second angle value, and determining the head-related transfer function value of the intermediate point according to an updated time length corresponding to the point, the sound frequency impulse value of the latter sampling point, and the head-related transfer function value of the previous intermediate point of the intermediate point.

[0035] In a possible implementation, the obtaining the binaural rendering signal value of each intermediate point according to the sound amplitude value of the intermediate point and the head-related transfer function value of the intermediate point comprises:

[0036] point-by-point convolution multiplication is performed on the sound amplitude value and the head-related transfer function value of the intermediate point to obtain the binaural rendering signal value of the intermediate point.

[0037] The second aspect of the present application provides an audio rendering device, comprising:

[0038] obtain audio PCM data, Cartesian coordinates and spherical coordinates of each audio sampling point of each sound object in the target sound stream, wherein the audio PCM data refers to audio pulse code modulation data;

[0039] construct a first sphere according to the spherical coordinates of a first sampling point of each adjacent two sampling points of the audio sampling points of the sound object, and construct a second sphere according to the spherical coordinates of a second sampling point of the two sampling points, wherein the first sampling point is a previous sampling point of the second sampling point, and the sphere corresponding to any sampling point represents the action range of the sound amplitude of the sampling point;

[0040] simulate and calculate the change of the sound amplitude in the spatial movement process from the first sampling point to the second sampling point according to the audio PCM data, the Cartesian coordinates and the spherical coordinates of the two sampling points, the first sphere and the second sphere, to obtain the sound amplitude of each intermediate point in the spatial movement process;

[0041] obtain the binaural rendering signal value of each intermediate point according to the sound amplitude of the intermediate point, to obtain the binaural rendering signal value of the target sound stream.

[0042] The third aspect of the present application provides a computer program product, comprising computer readable instructions, when the computer readable instructions run on an electronic device, the electronic device implements the audio rendering method of the first aspect or any implementation manner of the first aspect.

[0043] The fourth aspect of the present application provides an electronic device, comprising at least one processor and a memory connected with the processor, wherein:

[0044] The memory is used to store a computer program;

[0045] The processor is used to execute the computer program, so that the electronic device can implement the audio rendering method of the first aspect or any implementation manner of the first aspect.

[0046] The fifth aspect of the present application provides a computer storage medium, the storage medium carries one or more computer programs, when the one or more computer programs are executed by an electronic device, the electronic device can implement the audio rendering method of the first aspect or any implementation manner of the first aspect.

[0047] By the technical scheme, the audio rendering method provided by the application obtains the audio PCM data, the Cartesian coordinates and the spherical coordinates of each audio sampling point of each sound object in the target sound stream. Since the sound amplitude of each audio sampling point has a certain range of action and only acts on the points after the audio sampling point, the application constructs the spherical bodies corresponding to each adjacent two sampling points of each audio sampling point of each sound object according to the spherical coordinates of the two sampling points. Here, the spherical body corresponding to any sampling point represents the range of action of the sound amplitude of the sampling point. Then, the application simulates and calculates the change of the sound amplitude of the former sampling point in the spatial movement process of the two sampling points according to the audio PCM data, the Cartesian coordinates and the spherical coordinates of the two sampling points, combines the spherical bodies corresponding to the two sampling points, and obtains the sound amplitude of each intermediate point in the spatial movement process. Finally, the application obtains the binaural rendering signal value of each intermediate point according to the sound amplitude of each intermediate point, and obtains the binaural rendering signal value of the target sound stream. As can be seen, the application only needs to construct two spherical bodies and perform simple simulation calculation, so that the sound amplitude of each intermediate point between the adjacent two audio sampling points can be quickly obtained, and the real-time performance of the audio rendering is improved. BRIEF DESCRIPTION OF DRAWINGS

[0048] The above and other features, advantages and aspects of the present disclosure will become more apparent by describing in detail some embodiments thereof with reference to the attached drawings in which:

[0049] Figure 1 A flowchart of an audio rendering method provided by the application;

[0050] Figure 2 A schematic diagram in which a first spherical body and a second spherical body are completely contained;

[0051] Figure 3 A schematic diagram in which a first spherical body and a second spherical body are completely contained;

[0052] Figure 4 A schematic diagram in which a first spherical body and a second spherical body are partially contained;

[0053] Figure 5 A schematic diagram in which a first spherical body and a second spherical body are not contained at all;

[0054] Figure 6 A structural schematic diagram of an audio rendering device provided by the application;

[0055] Figure 7 A structural schematic diagram of an electronic device provided by the application. DETAILED DESCRIPTION

[0056] The embodiments of this application are described below with reference to the accompanying drawings. The terminology used in the implementation section of this application is for explaining specific embodiments only and is not intended to limit the scope of this application.

[0057] The embodiments of this application will now be described with reference to the accompanying drawings. Those skilled in the art will recognize that, with technological advancements and the emergence of new scenarios, the technical solutions provided in the embodiments of this application are equally applicable to similar technical problems.

[0058] The terms "first," "second," etc., used in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such terms are interchangeable where appropriate; this is merely a way of distinguishing objects with the same attributes in the embodiments of this application. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion, so that a process, method, system, product, or apparatus that comprises a series of elements is not necessarily limited to those elements, but may include other elements not explicitly listed or inherent to those processes, methods, products, or apparatuses.

[0059] This application provides an audio rendering method that can be applied to a terminal or server with audio data processing capabilities, and can also be applied to a system composed of the aforementioned terminal and server.

[0060] Of course, the audio rendering method provided in this application can also be applied to others, and this application does not make any specific limitations.

[0061] To enable those skilled in the art to better understand this application, the audio rendering method of the embodiments of this application will be described in detail below with reference to the accompanying drawings.

[0062] Reference Figure 1 , Figure 1 This is a flowchart illustrating an audio rendering method provided in an embodiment of this application, as shown below. Figure 1 As shown, the audio rendering method may include:

[0063] Step S101: Obtain the audio PCM data, Cartesian coordinates, and spherical coordinates of each audio sampling point of each sound object in the target audio stream.

[0064] In this embodiment, the target audio stream refers to the target audio stream that has been decoded by the decoder, such as the Jingcai audio stream. Optionally, the decoder refers to the AV3A decoder.

[0065] For example, the embodiment can parse the original sound stream from the player, then decode the original sound stream using the av3a decoder to obtain the decoded target sound stream, and then sequentially obtain the audio PCM data, the Cartesian coordinates and the spherical coordinates of each frame of the target sound stream of each sound object in the target sound stream. Here, the audio PCM data refers to audio Pulse Code Modulation (PCM) data.

[0066] It can be understood that one frame of audio data can include a plurality of audio sampling points, such as 1024 audio sampling points in one frame of target sound stream, and thus the audio PCM data, the Cartesian coordinates and the spherical coordinates of each frame of the target sound stream include the audio PCM data, the Cartesian coordinates and the spherical coordinates of each audio sampling point included in the frame of target sound stream.

[0067] It should be noted that the Cartesian coordinates and the spherical coordinates of one audio sampling point can be converted to each other.

[0068] In step S102, for each adjacent two sampling points of each audio sampling point of each sound object, a first sphere corresponding to the first sampling point and a second sphere corresponding to the second sampling point are constructed according to the spherical coordinates of the two sampling points.

[0069] Specifically, the embodiment can construct a first sphere according to the spherical coordinates of a first sampling point in the two sampling points, and construct a second sphere according to the spherical coordinates of a second sampling point in the two sampling points. The first sampling point is the previous sampling point of the second sampling point, and the sphere corresponding to any sampling point represents the range of the sound amplitude of the sampling point.

[0070] Considering that the sound amplitude of each audio sampling point has a certain range of action and only acts on the points after the audio sampling point, that is, taking the adjacent first sampling point and the second sampling point as an example, the sound amplitude of the first sampling point gradually weakens as the first sampling point moves to the second sampling point, and when the first sampling point moves to the action range of the second sampling point, the sound amplitude of the second sampling point is added with the distance from the second sampling point becoming smaller and smaller, so that when the first sampling point moves to the second sampling point, only the sound amplitude of the second sampling point is included (at this time the sound amplitude of the first sampling point is attenuated to 0).

[0071] Based on this, the embodiment can determine the sound amplitude action range of each adjacent two sampling points of each audio sampling point of each sound object, that is, construct a first sphere representing the sound amplitude action range of the previous sampling point (defined as the first sampling point in the embodiment) according to the spherical coordinates of the previous sampling point, and construct a second sphere representing the sound amplitude action range of the next sampling point (defined as the second sampling point in the embodiment) according to the spherical coordinates of the next sampling point.

[0072] In an alternative embodiment, the process of constructing the first sphere according to the spherical coordinates of the first sampling points can include: determining the sphere center according to the spherical coordinates (or Cartesian coordinates) of the first sampling points, and determining the sphere radius according to the radial distance in the spherical coordinates of the first sampling points, thereby constructing the first sphere; similarly, the process of constructing the second sphere according to the spherical coordinates of the second sampling points can include: determining the sphere center according to the spherical coordinates (or Cartesian coordinates) of the second sampling points, and determining the sphere radius according to the radial distance in the spherical coordinates of the second sampling points, thereby constructing the second sphere.

[0073] Taking any sphere as an example, the sound amplitude of a point (hereinafter referred to as a current point) on the sphere surface and inside the sphere can be obtained by the following formulas (1)-(5).

[0074] Formula (1);

[0075] Formula (2);

[0076] Formula (3);

[0077] Formula (4);

[0078] Formula (5);

[0079] wherein, represents the distance between the current point and the sphere center point, represents the coordinates of the current point, represents the coordinates of the sphere center point (such as the coordinates of the first sampling point or the coordinates of the second sampling point), represents the sound amplitude intensity of the current point, which is inversely related to the square of the distance, represents the inverse relationship, represents the loudness difference caused by the change of the distance between the current point and the sphere center point, represents the sound amplitude intensity of the sphere center point (obtained by calculating the sound amplitude squared through the audio PCM data and the bit depth of the sphere center point, where the bit depth refers to the number of binary numbers used to represent each audio sampling point, which is a known quantity), represents the amplitude intensity attenuation value caused by the change of the distance between the current point and the sphere center point, which is a decay coefficient, represents the sound amplitude of the sphere center point, represents the sound amplitude of the current point.

[0080] Step S103, according to the audio PCM data, the Cartesian coordinates and the spherical coordinates of the two sampling points respectively, the spherical body corresponding to the first sampling point and the spherical body corresponding to the second sampling point respectively simulate the change of the sound amplitude in the spatial movement process of the first sampling point, to obtain the sound amplitude of each intermediate point in the spatial movement process.

[0081] That is, the embodiment can simulate the change of the sound amplitude in the spatial movement process of the first sampling point to the second sampling point according to the audio PCM data, the Cartesian coordinates and the spherical coordinates of the two sampling points respectively, in combination with the first spherical body and the second spherical body, to obtain the sound amplitude of each intermediate point in the spatial movement process.

[0082] In the embodiment, the spatial movement process of the first sampling point to the second sampling point is taken as an example with a playing unit time of 1 ms, and the spatial movement process of the first sampling point to the second sampling point is represented by a plurality of intermediate points, that is, the first sampling point moves 1 ms to the first intermediate point, moves 2 ms to the second intermediate point, and so on, until the first sampling point moves to the second sampling point.

[0083] Of course, the above 1 ms is only an example, and the playing unit time can also be set to other values, which is not limited in the application.

[0084] Then, the embodiment can simulate the change of the sound amplitude in the spatial movement process of the first sampling point to the second sampling point according to whether each intermediate point is in the first spherical body or the second spherical body, to obtain the sound amplitude of each intermediate point in the spatial movement process.

[0085] Step S104, obtaining the binaural rendering signal value of each intermediate point according to the sound amplitude of each intermediate point; to obtain the binaural rendering signal value of the target sound stream.

[0086] In the embodiment, the binaural sound signal of each intermediate point can be rendered according to the sound amplitude of each intermediate point, to obtain the binaural rendering signal value of each intermediate point.

[0087] The above processing is performed for all adjacent two sampling points of the target sound stream, that is, the binaural rendering signal value of the target sound stream can be obtained.

[0088] The audio PCM data, the Cartesian coordinates and the spherical coordinates of each audio sampling point of each sound object in the target sound stream are acquired. Since the sound amplitude of each audio sampling point has a certain range of action and only acts on the points after the audio sampling point, for each adjacent two audio sampling points of each audio sampling point of each sound object, the spherical bodies corresponding to the two audio sampling points are constructed according to the spherical coordinates of the two audio sampling points, here, the spherical body corresponding to any audio sampling point represents the range of action of the sound amplitude of the audio sampling point, and then the change of the sound amplitude of the former audio sampling point in the spatial movement process is simulated and calculated according to the audio PCM data, the Cartesian coordinates and the spherical coordinates of the two audio sampling points, in combination with the spherical bodies corresponding to the two audio sampling points, to obtain the sound amplitude of each intermediate point in the spatial movement process, and finally the binaural rendering signal value of each intermediate point is obtained according to the sound amplitude of each intermediate point; so as to obtain the binaural rendering signal value of the target sound stream. As can be seen, the two spherical bodies are only needed to be constructed and the simple simulation calculation is needed to be performed, so that the sound amplitudes of the intermediate points between the adjacent two audio sampling points can be quickly obtained, and the real-time performance of the audio rendering is improved.

[0089] In some embodiments of the present application, the process of the foregoing step S103 “the change of the sound amplitude of the former audio sampling point in the spatial movement process is simulated and calculated according to the audio PCM data, the Cartesian coordinates and the spherical coordinates of the two audio sampling points, in combination with the spherical bodies corresponding to the two audio sampling points, to obtain the sound amplitude of each intermediate point in the spatial movement process” is introduced in detail.

[0090] As introduced before, the change of the sound amplitude in the spatial movement process from the first audio sampling point to the second audio sampling point can be simulated and calculated according to whether each intermediate point is in the first spherical body or the second spherical body, and meanwhile, the first spherical body and the second spherical body also have a certain inclusion relationship, which also affects the sound amplitude change of each intermediate point, therefore, the inclusion relationship of the first spherical body and the second spherical body and the positional relationship of each intermediate point with the first spherical body and the second spherical body can be determined according to the Cartesian coordinates and the spherical coordinates of the two audio sampling points, and the sound amplitude of each intermediate point can be determined according to the inclusion relationship and the positional relationship of each intermediate point with the first spherical body and the second spherical body, and the audio PCM data, the Cartesian coordinates and the spherical coordinates of the two audio sampling points.

[0091] In one possible implementation, the process of "determining the inclusion relationship of the spheres corresponding to the two sampling points and the positional relationship between the spheres corresponding to the two sampling points and each intermediate point based on their respective Cartesian and spherical coordinates" (i.e., "determining the inclusion relationship of the first sphere and the second sphere and the positional relationship between each intermediate point and the first and second spheres based on their respective Cartesian and spherical coordinates") can include: determining the sampling point distance between the two sampling points based on their respective Cartesian coordinates; determining the inclusion relationship of the spheres corresponding to the two sampling points based on the radial distance and sampling point distance in the spherical coordinates of the two sampling points; obtaining the duration of the target sound stream in each frame; determining the movement distance corresponding to each intermediate point based on the duration; and determining the positional relationship between the spheres corresponding to the two sampling points and each intermediate point based on the movement distance corresponding to each intermediate point, the radial distance and sampling point distance in the spherical coordinates of the two sampling points, and the sampling point distance.

[0092] It is understandable that the inclusion relationship between the first sphere and the second sphere includes: complete inclusion, partial inclusion, and no inclusion at all, such as... Figures 2-5 This is a schematic diagram illustrating three types of inclusion relationships between the first sphere and the second sphere. Figure 2 This is a schematic diagram showing how a first sphere and a second sphere completely contain each other. Figure 3 This is a schematic diagram showing how a first sphere and a second sphere can be completely contained within each other. Figure 4 This is a schematic diagram showing the components of the first and second spheres. Figure 5 This is a schematic diagram where the first and second spheres are not included at all.

[0093] like Figures 2-5 , This represents the distance between the center point of the first sphere and the center point of the second sphere, which is also the distance between the first sampling point and the second sampling point. Represents the radius of the first sphere. This represents the radius of the second sphere.

[0094] Optionally, in this embodiment, the distance between two sampling points can be calculated by performing Euclidean distance calculation based on their respective Cartesian coordinates. For ease of explanation later, this distance will be defined as the sampling point distance.

[0095] Therefore, in this embodiment, the radial distance in the spherical coordinates of the two sampling points can be used as a basis. and and the distance of the sampling point Determine the inclusion relationship between the first sphere and the second sphere. Wherein, if the following conditions are met... Then, the first sphere and the second sphere are determined to be completely contained if the following conditions are met: , then it is determined that the first sphere and the second sphere are in a partial containing relationship, if , then it is determined that the first sphere and the second sphere are in a complete non-containing relationship.

[0096] Optionally, the process of determining the moving distance corresponding to each intermediate point according to the duration can include: dividing the sampling point distance by the duration to obtain a quotient value as the unit moving distance of each intermediate point, and multiplying the unit moving distance of each intermediate point by the playing duration corresponding to each intermediate point to obtain the moving distance corresponding to each intermediate point, see formula (6) as follows.

[0097] Formula (6);

[0098] wherein, represents the moving distance corresponding to the intermediate point, that is, the distance of the first sampling point moving to the second sampling point, taking the Cartesian coordinates of the first sampling point as the origin. For example, It can also be understood as the distance between the intermediate point and . represents the duration, represents the playing duration corresponding to the intermediate point.

[0099] Optionally, the number of sampling points of a frame of target sound stream is 1024, and the sampling rate is 48000, so the duration of a frame of target sound stream is : 1024 / 48000≈21ms.

[0100] Taking 1ms as the playing unit time, the playing duration corresponding to the first intermediate point (the first intermediate point refers to the position reached after the first sampling point moves 1ms) is 1ms, the playing duration corresponding to the second intermediate point (the second intermediate point refers to the position reached after the first sampling point moves 2ms) is 2ms, the playing duration corresponding to the third intermediate point (the third intermediate point refers to the position reached after the first sampling point moves 3ms) is 3ms, and so on.

[0101] The following takes any intermediate point as an example to introduce the process of determining the positional relationship of each intermediate point with the first sphere and the second sphere according to the moving distance corresponding to each intermediate point, the radial distance in the spherical coordinates of the two sampling points, and the sampling point distance.

[0102] When the first sphere and the second sphere are in a complete containing relationship: if and , it is determined that the intermediate point does not fall within the second sphere; if and​​​ , or satisfying , then it is determined that the intermediate point falls within the second sphere.

[0103] When the first sphere and the second sphere are in partial containing relationship: if satisfying , then it is determined that the intermediate point only falls within the first sphere; if satisfying , then it is determined that the intermediate point falls within both the first sphere and the second sphere; if satisfying , then it is determined that the intermediate point only falls within the second sphere.

[0104] When the first sphere and the second sphere are in complete non-containing relationship: if satisfying , then it is determined that the intermediate point only falls within the first sphere; if satisfying , then it is determined that the intermediate point does not fall within both the first sphere and the second sphere; if satisfying , then it is determined that the intermediate point only falls within the second sphere.

[0105] Optionally, the process of determining the sound amplitude of each intermediate point according to the containing relationship and the position relationship between the two sampling points respectively corresponding to the spheres and each intermediate point, and the audio PCM data, Cartesian coordinates and spherical coordinates of the two sampling points can include: determining a sound amplitude function of each intermediate point according to the containing relationship and the position relationship between the two sampling points respectively corresponding to the spheres and each intermediate point, determining the independent variable in the sound amplitude function of each intermediate point according to the audio PCM data, Cartesian coordinates and spherical coordinates of the two sampling points, and substituting the independent variable in the sound amplitude function of each intermediate point into the sound amplitude function of each intermediate point to perform function calculation, so as to obtain the sound amplitude of each intermediate point.

[0106] Optionally, for each intermediate point, the process of determining the sound amplitude function of the intermediate point according to the containing relationship and the position relationship between the intermediate point and the first sphere and the second sphere can include the following several cases.

[0107] The first case: when the first sphere and the second sphere are in complete containing relationship and the intermediate point does not fall within the second sphere, since the intermediate point does not fall within the second sphere, the sound amplitude of the intermediate point is not affected by the sound amplitude of the second sampling point, and thus a function of only attenuating the sound amplitude of the first sampling point with the increase of the moving distance corresponding to the intermediate point can be generated as the sound amplitude function of the intermediate point.

[0108] That is, when satisfying , and , the sound amplitude function of the intermediate point is as follows.

[0109] Formula (7);

[0110] in, This represents the sound amplitude at the midpoint, i.e., the movement of the first sampling point. The sound amplitude after the time interval; Indicates will As calculated by formula (1) Substituting into formula (2) - formula (5) yields (At this point, the first sampling point) (The center of the ball).

[0111] The second method: When the first and second spheres are completely contained within each other and the midpoint falls within the second sphere, the sound amplitude at this midpoint is affected not only by the sound amplitude at the first sampling point but also by the sound amplitude at the second sampling point. Therefore, it is possible to generate a sound amplitude that is influenced by both the movement distance corresponding to the midpoint and the amplitude calculated with the second sampling point as the center of the sphere. (hereinafter referred to as) As the first sampling point gradually moves towards the second sampling point, The function of the effect of gradually increasing until it becomes 1) is used as the sound amplitude function at that intermediate point.

[0112] That is, when both conditions are met , and or simultaneously satisfy and At that time, the sound amplitude function at the intermediate point is as follows (8).

[0113] Formula (8);

[0114] in, Indicates will As calculated by formula (1) Substituting into formula (2) - formula (4) to calculate (At this point, the second sampling point) (for the center of the ball) Indicates will As calculated by formula (1) Substituting into formula (2) - formula (5) to calculate (At this point, the second sampling point) (The center of the ball).

[0115] The third scenario: when the first sphere and the second sphere are partially contained within each other, and the intermediate point falls only within the first sphere (i.e., the intermediate point falls within the first sphere). Figure 4When the sound amplitude at the intermediate point is within the first sphere (but not within the overlapping area of ​​the first and second spheres), since it does not fall within the second sphere, the sound amplitude at that intermediate point is not affected by the sound amplitude at the second sampling point. Therefore, a function that attenuates only the sound amplitude at the first sampling point as the moving distance corresponding to that intermediate point increases can be generated as the sound amplitude function at that intermediate point.

[0116] That is, when both conditions are met and At that time, the sound amplitude function at the intermediate point is as shown in the above formula (7).

[0117] The fourth type: When the first sphere and the second sphere are partially contained within each other, and the intermediate point falls within both the first sphere and the second sphere (i.e., the intermediate point falls within the first sphere and the second sphere). Figure 4 When the sound amplitude at the midpoint (within the overlapping area of ​​the first and second spheres shown) is affected by both the sound amplitude at the first sampling point and the sound amplitude at the second sampling point, the amplitude can be adjusted accordingly. Alternatively, bilinear gain control can be used to allow these two amplitudes to influence each other, i.e., the amplitudes of the spheres centered at the first sampling point can be calculated separately. and the second sampling point is spherical The function shown in formula (9) is obtained as the sound amplitude function at the intermediate point.

[0118] That is, when both conditions are met and At that time, the sound amplitude function at the intermediate point is as follows (9).

[0119] Formula (9);

[0120] in, Indicates will As calculated by formula (1) Substituting into formula (2) - formula (4) to calculate (At this point, the first sampling point) (The center of the ball).

[0121] The fifth scenario occurs when the first sphere and the second sphere are partially contained within each other, and the intermediate point falls only within the second sphere (i.e., the intermediate point falls within the second sphere). Figure 4 When the second sphere is located within the area where the first and second spheres overlap (but not within the overlapping region of the first and second spheres), as previously explained, this intermediate point is essentially a position where the first sampling point has moved. Therefore, the sound amplitude at this intermediate point will still be affected by the sound amplitude of the first sampling point, and will also be affected by... The influence of this can generate a function as shown in the above formula (8), which serves as the sound amplitude function at the intermediate point.

[0122] That is, when both conditions are met and When the above conditions are satisfied, the sound amplitude function of the intermediate point is as shown in the above formula (7).

[0123] Sixth, when the first sphere and the second sphere are completely non-inclusive and the intermediate point only falls within the first sphere, it is only affected by the sound amplitude of the first sampling point, so a function as shown in formula (7) can be generated as the sound amplitude function of the intermediate point.

[0124] That is, when the above conditions are satisfied, and the sound amplitude function of the intermediate point is as shown in the above formula (7).

[0125] Seventh, when the first sphere and the second sphere are completely non-inclusive and the intermediate point is not in the first sphere and the second sphere at the same time, a spatial attenuation mechanism can be introduced to obtain a function as shown in the following formula (10) as the sound amplitude function of the intermediate point.

[0126] That is, when the above conditions are satisfied, and the sound amplitude function of the intermediate point is as shown in the above formula (10).

[0127] Formula (10);

[0128] wherein, represents the point on the surface of the first sphere with the first sampling point as the center of the sphere as the current point, and the , calculated according to the above formula (1)-formula (5), represents the point on the surface of the second sphere with the second sampling point as the center of the sphere as the current point, and the calculated according to the above formula (1)-formula (5), represents a preset attenuation value coefficient, which can be selected as 0.8.

[0129] Eighth, when the first sphere and the second sphere are completely non-inclusive and the intermediate point only falls within the second sphere, a function as shown in formula (8) can be generated as the sound amplitude function of the intermediate point.

[0130] That is, when the above conditions are satisfied, and the sound amplitude function of the intermediate point is as shown in the above formula (8).

[0131] In summary, the embodiment of the present application can simulate the sound amplitude change (i.e. sound loudness difference) in the process of moving the first sampling point to the second sampling point based on the two constructed spheres, can accurately determine the sound amplitude of the first sampling point at the moving position corresponding to each playing unit time, and since the multi-channel Ambisonics information signal processing process is not required, the time for binaural rendering is saved, and it is ensured that the multi-sound object vivid sound content can be played in real time in the player without the problems of audio and video out of synchronization or playing lag.

[0132] It can be understood that the phase change also needs to ensure that the first sampling point can move uniformly to the second sampling point, and therefore, the following formula (11) can be used to calculate the phase change degree of the first sampling point after moving to the second sampling point at the moment.

[0133] Formula (11);

[0134] wherein, represents the position of the first sampling point (i.e. the position of the center of the first sphere), represents the position of the second sampling point (i.e. the position of the center of the second sphere), represents the phase change degree of the first sampling point after moving to the second sampling point at the moment.

[0135] Taking HOA third-order coding of 12 sound objects as an example, the signal processing on the channel reduces 180 (180 = 12 x 16 - 12) times of coding processing, which can maximize the performance indicators of the sound objects in real-time rendering and ensure that the phase movement will not cause large image jumps.

[0136] In some other embodiments of the present application, the process of the foregoing step S104 "obtaining the binaural rendering signal value of each intermediate point according to the sound amplitude of each intermediate point" is introduced.

[0137] It can be understood that in the current implementation, the corresponding elevation angle, horizontal angle and spatial impulse signal value (i.e. HRTF data, here, the full name of HRTF is Head-Related Transfer Function, which represents head-related transfer function) are usually measured by the laboratory. However, for the HRTF data measured by the laboratory, such as the HRTF data in the commonly used KEMAR dummy head microphone, the interval angle between the measured angles is usually more than 10°, and the elevation angle is even more than 15°.

[0138] ​​In this embodiment, the phase angle changes between the intermediate points between the first sampling point and the second sampling point are often small. Therefore, the HRTF dataset currently tested in the laboratory has insufficient HRTF data and cannot meet the requirements of this application for HRTF data corresponding to image position angles with small range changes.

[0139] To achieve more accurate spatial mapping, spatial interpolation of the HRTF is required to achieve a spatial accuracy of 5° or even less. Specifically, this embodiment can perform dynamic interpolation calculations based on the head-related transfer function (HRTF) values ​​for each intermediate point during spatial movement, using the spherical coordinates of the two sampling points. This yields the HRTF value for each intermediate point, where the HRTF value represents the image position information of that point. Then, based on the sound amplitude and HRTF value of each intermediate point, the binaural rendering signal value for each intermediate point is obtained.

[0140] The following describes the process of "using the spherical coordinates of the two sampling points to perform dynamic interpolation calculations based on the head-related transfer function values ​​for each intermediate point during spatial movement, and obtaining the head-related transfer function value for each intermediate point".

[0141] First, in this embodiment, the unit vector of the Cartesian coordinates of the two sampling points can be calculated based on the pitch and azimuth angles in the spherical coordinates of the two sampling points, as well as the audio pulse values ​​(i.e., HRTF data) of the two sampling points.

[0142] In this embodiment, the elevation and azimuth angles of the two sampling points can be converted into the following forms according to the existing HRTF measurement method: ,in, Indicates the pitch angle of the sampling point. Indicates the azimuth angle of the sampling point. and This represents the sound frequency pulse value at the sampling point, more specifically, Indicates in and Below, the left ear acoustic pulse value corresponding to HRTF. Indicates in and Below, the right ear acoustic pulse value corresponding to HRTF.

[0143] For ease of explanation below, the first sampling point Recorded as , the first sampling point Recorded as The second sampling point Recorded as The second sampling point denoted as .

[0144] Taking the elevation angle and azimuth angle of the first sampling point as (0°, 0°) and the elevation angle and azimuth angle of the second sampling point as (0°, 15°) as an example, due to measurement problems, the HRTF data in the range of 0° to 15° of the azimuth angle cannot be queried from the HRTF data set. Therefore, the following dynamic interpolation can be performed.

[0145] The embodiment can also convert the elevation angle and azimuth angle of each of the two sampling points into a unit vector of Cartesian coordinates according to an existing calculation method of the unit vector of Cartesian coordinates. In order to facilitate the description hereinafter, the unit vector of Cartesian coordinates of the first sampling point is denoted as , and the unit vector of Cartesian coordinates of the second sampling point is denoted as .

[0146] Then, the included angle between the two sampling points and the origin of the Cartesian coordinates can be determined according to the unit vector of Cartesian coordinates of each of the two sampling points, that is, the dot product calculation of the unit vector of Cartesian coordinates of each of the two sampling points is performed, that is, , and the included angle β between the two sampling points and the origin of the Cartesian coordinates is .

[0147] Further, the duration of each target sound stream and the playback duration corresponding to each intermediate point can be obtained, and the angle change value corresponding to each intermediate point can be determined according to the duration, the included angle, and the playback duration corresponding to each intermediate point.

[0148] Optionally, the angle change value corresponding to each intermediate point can be calculated by using the following formula (12) .

[0149] Formula (12).

[0150] Therefore, the spherical linear interpolation angle of each intermediate point can be determined according to the unit vector of Cartesian coordinates of each of the two sampling points, the playback duration corresponding to each intermediate point, and the angle change value.

[0151] Optionally, the spherical linear interpolation angle of each intermediate point can be calculated by using the following formula (13).

[0152] Formula (13);

[0153] wherein, represents the spherical linear interpolation angle of the intermediate point with the playback duration t.

[0154] Finally, the embodiment can determine the head-related transfer function value of each intermediate point according to the spherical linear interpolation angle of each intermediate point and the sound frequency pulse value of each of the two sampling points.

[0155] Optionally, the process of "determining the head-related transfer function value of each intermediate point according to the spherical linear interpolation angle of each intermediate point and the sound frequency pulse value of each of the two sampling points" can include:

[0156] For each intermediate point:

[0157] If the spherical linear interpolation angle of the intermediate point is less than a preset first angle value, the preset target unit time is increased by 1, and the head-related transfer function value of the previous intermediate point of the intermediate point is determined as the head-related transfer function value of the intermediate point, wherein if the intermediate point is the first intermediate point after the previous sampling point, the head-related transfer function value of the previous intermediate point of the intermediate point is the sound frequency pulse value of the previous sampling point;

[0158] If the spherical linear interpolation angle of the intermediate point is greater than or equal to the first angle value and less than or equal to a preset second angle value, or the target unit time is equal to a preset count threshold, the updated time length corresponding to the intermediate point is determined according to the target unit time and the playing time length corresponding to the intermediate point, the head-related transfer function value of the intermediate point is determined according to the updated time length corresponding to the intermediate point, the sound frequency pulse value of the latter one of the two sampling points, and the head-related transfer function value of the previous intermediate point of the intermediate point, and the target unit time is set to 0;

[0159] If the spherical linear interpolation angle of the intermediate point is greater than the second angle value, the first point that makes the spherical linear interpolation angle greater than the second angle value is determined, and the head-related transfer function value of the intermediate point is determined according to the updated time length corresponding to the point, the sound frequency pulse value of the latter one of the two sampling points, and the head-related transfer function value of the previous intermediate point of the intermediate point.

[0160] The above process is explained in detail as follows.

[0161] Optionally, in the embodiment, for each intermediate point, the following dynamic interpolation strategy can be used: according to If the angle change does not exceed the first angle value, the preset target unit time is increased by 1, the number n of times of increasing the target unit time is recorded (n represents the target unit time), and the head-related transfer function value of the intermediate point is obtained by convolving the head-related transfer function value of the previous intermediate point of the intermediate point (if the intermediate point is the first intermediate point after the first sampling point, the head-related transfer function value of the previous intermediate point of the intermediate point is the sound frequency pulse value of the first sampling point).

[0162] In a first possible case, if the spherical linear interpolation angle of the intermediate point is less than the preset first angle value, the preset target unit time is increased by 1, and the head-related transfer function value of the previous intermediate point of the intermediate point is determined as the head-related transfer function value of the intermediate point, wherein if the intermediate point is the first intermediate point after the first sampling point, the head-related transfer function value of the previous intermediate point of the intermediate point is the sound frequency pulse value of the first sampling point.

[0163] Optionally, the first angle value is 1°. Then, if <1°, n = n + 1, , wherein n represents the target unit time, and the initial value is 0; represents the left ear component of the head-related transfer function value of the intermediate point, represents the right ear component of the head-related transfer function value of the intermediate point; if the intermediate point is the first intermediate point after the first sampling point (for example, the first sampling point moves to the second sampling point after sequentially passing through the intermediate points c, d, e and f, and the intermediate point is c), then is the left ear sound frequency pulse value of the first sampling point, is the right ear sound frequency pulse value of the first sampling point, if the intermediate point is not the first intermediate point after the first sampling point (for example, the first sampling point moves to the second sampling point after sequentially passing through the intermediate points c, d, e and f, and the intermediate point is d or e or f), then is the left ear component of the head-related transfer function value of the previous intermediate point of the intermediate point, is a left ear component of the head-related transfer function value of the previous intermediate point of the intermediate point.

[0164] The second possible case: if the spherical linear interpolation angle of the intermediate point is greater than or equal to the first angle value and less than or equal to a preset second angle value, or the target unit time is equal to a preset count threshold, then the updated time length corresponding to the intermediate point is determined according to the target unit time and the playing time corresponding to the intermediate point, the head-related transfer function value of the intermediate point is determined according to the updated time length corresponding to the intermediate point, the sound frequency pulse value of the second sampling point, and the head-related transfer function value of the previous intermediate point of the intermediate point, and the target unit time is set to 0.

[0165] Optionally, the second angle value is 5°, and the count threshold is 5. Then, if 1°≤ ≤5° or (notably, n is gradually accumulated from 0 to 5 and then returns to 0, and n will not be greater than 5), then the updated time length corresponding to the intermediate point is , the left ear component of the head-related transfer function value of the intermediate point is: , the right ear component of the head-related transfer function value of the intermediate point is: , and n=0. Wherein, if the intermediate point is the first intermediate point after the first sampling point, then is the left ear sound frequency pulse value of the first sampling point, is the right ear sound frequency pulse value of the first sampling point, if the intermediate point is not the first intermediate point after the first sampling point, then is a left ear component of the head-related transfer function value of the previous intermediate point of the intermediate point, is a right ear component of the head-related transfer function value of the previous intermediate point of the intermediate point.

[0166] The third possible case: if the spherical linear interpolation angle of the intermediate point is greater than the second angle value, then the first point whose spherical linear interpolation angle is greater than the second angle value is determined, and the head-related transfer function value of the intermediate point is determined according to the updated time length corresponding to the point, the sound frequency pulse value of the second sampling point, and the head-related transfer function value of the previous intermediate point of the intermediate point.

[0167] Optionally, if >5°, then the left ear component of the head-related transfer function value of the intermediate point is: , the right ear component of the head-related transfer function value of the intermediate point is: , and denotes the updated time length corresponding to the first point at which the spherical linear interpolation angle is greater than the second angle value, if the intermediate point is the first intermediate point after the first sampling point, then denotes the left ear sound frequency pulse value of the first sampling point, denotes the right ear sound frequency pulse value of the first sampling point, if the intermediate point is not the first intermediate point after the first sampling point, then denotes the left ear component of the head-related transfer function value of the previous intermediate point of the intermediate point, denotes the right ear component of the head-related transfer function value of the previous intermediate point of the intermediate point.

[0168] Through the above process, the interpolated head-related transfer function values of each of the three possible cases can be obtained, that is, .

[0169] Optionally, the process of "obtaining the binaural rendering signal value of each intermediate point according to the sound amplitude of each intermediate point and the head-related transfer function value of each intermediate point" can include: performing point-to-point convolution multiplication on the sound amplitude of each intermediate point and the head-related transfer function value to obtain the binaural rendering signal value of each intermediate point.

[0170] In summary, in the embodiment, in order to avoid the auditory discontinuity phenomenon caused by the sudden change of sound image position, the spherical linear interpolation technology is used to realize the smooth transition of the head-related transfer function (HRTF). First, the azimuth angle and the elevation angle corresponding to the two HRTF data are converted into unit direction vectors in three-dimensional space. Then, the included angle between the vectors and the interpolation weight are calculated to perform spherical linear interpolation on the shortest path on the unit sphere, so as to ensure the natural smoothness of the sound image moving track. At the same time, the sound frequency pulse values of the left and right ear HRTFs are linearly mixed in the time domain, and the smooth transition of the frequency spectrum is realized through cross-fade. This method of combining spatial interpolation and time domain mixing can effectively eliminate the auditory discomfort caused by the jumping of the sound image.

[0171] The point-to-point convolution multiplication of the sound amplitude of each intermediate point and the head-related transfer function value makes the binaural rendering signal value of each intermediate point be accurately and timely obtained, so as to ensure that the multi-sound object vivid sound content is played in real time in the player and has a better binaural rendering effect.

[0172] The above introduces an audio rendering method provided by the embodiment of the application, and the following will introduce a device for executing the above audio rendering method.

[0173] Please refer to Figure 6 , Figure 6 for a structural schematic diagram of an audio rendering device provided by the embodiment of the application. As shown inFigure 6 The audio rendering device can include:

[0174] The base data acquisition module 601 is configured to acquire audio PCM data, Cartesian coordinates, and spherical coordinates of each audio sampling point of each sound object in the target sound stream, where the audio PCM data refers to audio pulse code modulation data.

[0175] The sphere construction module 602 is configured to construct a sphere corresponding to each of the two sampling points according to the spherical coordinates of the two sampling points, where the sphere corresponding to any sampling point represents a range of sound amplitude of the sampling point.

[0176] The sound amplitude simulation module 603 is configured to simulate the change in sound amplitude of a first sampling point in a spatial movement process of the two sampling points according to the audio PCM data, Cartesian coordinates, and spherical coordinates of the two sampling points, and the spheres corresponding to the two sampling points, to obtain the sound amplitude of each intermediate point in the spatial movement process.

[0177] The binaural rendering module 604 is configured to obtain a binaural rendering signal value of each intermediate point according to the sound amplitude of each intermediate point, to obtain a binaural rendering signal value of the target sound stream.

[0178] In a possible implementation, when the sound amplitude simulation module simulates the change in sound amplitude of a first sampling point in a spatial movement process of the two sampling points according to the audio PCM data, Cartesian coordinates, and spherical coordinates of the two sampling points, and the spheres corresponding to the two sampling points, to obtain the sound amplitude of each intermediate point in the spatial movement process, the sound amplitude simulation module can be specifically configured to:

[0179] determine the inclusion relationship of the spheres corresponding to the two sampling points, and the positional relationship between the spheres corresponding to the two sampling points and each intermediate point according to the Cartesian coordinates and the spherical coordinates of the two sampling points.

[0180] determine the sound amplitude of each intermediate point according to the inclusion relationship, the positional relationship between the spheres corresponding to the two sampling points and each intermediate point, and the audio PCM data, Cartesian coordinates, and spherical coordinates of the two sampling points.

[0181] In a possible implementation, when the sound amplitude simulation module determines the inclusion relationship of the spheres corresponding to the two sampling points, and the positional relationship between the spheres corresponding to the two sampling points and each intermediate point according to the Cartesian coordinates and the spherical coordinates of the two sampling points, the sound amplitude simulation module can be specifically configured to:

[0182] determine the sampling point distance between the two sampling points according to the Cartesian coordinates of the two sampling points.

[0183] determine the containing relationship of the two spherical bodies corresponding to the two sampling points respectively according to the radial distance in the spherical coordinates of the two sampling points respectively and the sampling point distance;

[0184] obtain the duration of each target sound stream, and determine the moving distance corresponding to each intermediate point according to the duration;

[0185] determine the positional relationship of the two spherical bodies corresponding to the two sampling points respectively and each intermediate point respectively according to the moving distance corresponding to each intermediate point, the radial distance in the spherical coordinates of the two sampling points respectively and the sampling point distance.

[0186] In a possible implementation, when the sound amplitude simulation module determines the sound amplitude of each intermediate point according to the containing relationship and the positional relationship of the two spherical bodies corresponding to the two sampling points respectively and each intermediate point respectively, and the audio PCM data, the Cartesian coordinates and the spherical coordinates of the two sampling points respectively, the sound amplitude simulation module can be specifically used for:

[0187] determine the sound amplitude function of each intermediate point according to the containing relationship and the positional relationship of the two spherical bodies corresponding to the two sampling points respectively and each intermediate point respectively;

[0188] determine the independent variable in the sound amplitude function of each intermediate point according to the audio PCM data, the Cartesian coordinates and the spherical coordinates of the two sampling points respectively;

[0189] substitute the independent variable in the sound amplitude function of each intermediate point into the sound amplitude function of each intermediate point for function calculation to obtain the sound amplitude of each intermediate point.

[0190] In a possible implementation, when the binaural rendering module obtains the binaural rendering signal value of each intermediate point according to the sound amplitude of each intermediate point, the binaural rendering module can be specifically used for:

[0191] perform dynamic interpolation calculation based on the head-related transfer function value for each intermediate point in the spatial movement process according to the spherical coordinates of the two sampling points respectively to obtain the head-related transfer function value of each intermediate point, wherein the head-related transfer function value of each intermediate point represents the image information of each intermediate point;

[0192] obtain the binaural rendering signal value of each intermediate point according to the sound amplitude of each intermediate point and the head-related transfer function value of each intermediate point.

[0193] In a possible implementation, when the binaural rendering module performs dynamic interpolation calculation based on the head-related transfer function value for each intermediate point in the spatial movement process according to the spherical coordinates of the two sampling points respectively to obtain the head-related transfer function value of each intermediate point, the binaural rendering module can be specifically used for:

[0194] According to the pitch angle and the azimuth angle in the spherical coordinates of the two sampling points respectively, a unit vector of the Cartesian coordinates of the two sampling points respectively is calculated, and a sound frequency pulse value of the two sampling points respectively is calculated;

[0195] According to the unit vector of the Cartesian coordinates of the two sampling points respectively, an included angle between the two sampling points and the origin of the Cartesian coordinates is determined;

[0196] A duration of each target sound stream and a playing duration corresponding to each intermediate point are obtained, and according to the duration, the included angle and the playing duration corresponding to each intermediate point, an angle change value corresponding to each intermediate point is determined;

[0197] According to the unit vector of the Cartesian coordinates of the two sampling points respectively, the playing duration corresponding to each intermediate point and the angle change value, a spherical linear interpolation angle of each intermediate point is determined;

[0198] According to the spherical linear interpolation angle of each intermediate point and the sound frequency pulse value of the two sampling points respectively, a head-related transfer function value of each intermediate point is determined.

[0199] In a possible implementation, when the binaural rendering module determines the head-related transfer function value of each intermediate point according to the spherical linear interpolation angle of each intermediate point and the sound frequency pulse value of the two sampling points respectively, the binaural rendering module can be specifically configured to:

[0200] For each intermediate point:

[0201] If the spherical linear interpolation angle of the intermediate point is less than a preset first angle value, a preset target unit time is increased by 1, and a head-related transfer function value of a previous intermediate point of the intermediate point is determined as the head-related transfer function value of the intermediate point, wherein if the intermediate point is the first intermediate point after a previous sampling point, the head-related transfer function value of the previous intermediate point of the intermediate point is a sound frequency pulse value of the previous sampling point;

[0202] If the spherical linear interpolation angle of the intermediate point is greater than or equal to a first angle value and less than or equal to a preset second angle value, or the target unit time is equal to a preset count threshold, an updated duration corresponding to the intermediate point is determined according to the target unit time and the playing duration corresponding to the intermediate point, the head-related transfer function value of the intermediate point is determined according to the updated duration corresponding to the intermediate point, a sound frequency pulse value of a later sampling point of the two sampling points and the head-related transfer function value of the previous intermediate point of the intermediate point, and the target unit time is set to 0;

[0203] If the spherical linear interpolation angle of the intermediate point is greater than the second angle value, a first point is determined, which makes the spherical linear interpolation angle greater than the second angle value, and a head-related transfer function value of the intermediate point is determined according to the updated time length corresponding to the point, the sound frequency pulse value of the next sampling point, and the head-related transfer function value of the previous intermediate point of the intermediate point.

[0204] In a possible implementation, when the binaural rendering module obtains the binaural rendering signal value of each intermediate point according to the sound amplitude value of each intermediate point and the head-related transfer function value of each intermediate point, the binaural rendering module can be specifically used for:

[0205] convolving and multiplying the sound amplitude value and the head-related transfer function value of each intermediate point point by point to obtain the binaural rendering signal value of each intermediate point.

[0206] The modules in the audio rendering device can be all or partially implemented by software, hardware, and a combination thereof. The modules can be embedded in or independent of a processor in a computer device in a hardware form, or can be stored in a memory in a computer device in a software form, so as to be called and executed by a processor to perform operations corresponding to the modules.

[0207] In an embodiment of the present application, an electronic device is also provided. The electronic device can include at least one processor and a memory connected to the processor, wherein:

[0208] The memory is configured to store a computer program.

[0209] The processor is configured to execute the computer program, so that the electronic device can implement any of the audio rendering methods provided in the embodiments of the present application.

[0210] Reference Figure 7 As shown in FIG. 1, a structure diagram of an electronic device suitable for implementing the electronic device in the embodiments of the present application is shown. The electronic device in the embodiments of the present application can include, but is not limited to, a fixed terminal such as a mobile phone, a notebook computer, a PDA (Personal Digital Assistant), a PAD (Tablet Personal Computer), a desktop computer, and the like. Figure 7 The electronic device shown is only an example, and should not impose any limitation on the functions and use range of the embodiments of the present application.

[0211] As Figure 7As shown, the electronic device can include a processing device (e.g., a central processing unit, a graphics processing unit, etc.) 701 that can perform various appropriate actions and processes according to programs stored in a read-only memory (ROM) 702 or loaded from a storage device 708 into a random access memory (RAM) 703. In a state in which the electronic device is powered on, various programs and data required for operation of the electronic device are also stored in the RAM 703. The processing device 701, the ROM 702, and the RAM 703 are connected to each other through a bus 704. An input / output (I / O) interface 705 is also connected to the bus 704.

[0212] Generally, the following devices can be connected to the I / O interface 705: input devices 706 including, for example, a touch screen, a touch pad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, and the like; output devices 707 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, and the like; storage devices 708 including, for example, a memory card, a hard disk, and the like; and communication devices 709. The communication devices 709 can allow the electronic device to communicate wirelessly or wiredly with other devices to exchange data. Although Figure 7 An electronic device having various devices is shown, but it is understood that all of the shown devices are not required to be implemented or possessed. More or fewer devices can be alternatively implemented or possessed.

[0213] The embodiments of the present application also provide a computer program product including computer readable instructions, which, when executed on an electronic device, cause the electronic device to implement any of the audio rendering methods provided by the embodiments of the present application.

[0214] The embodiments of the present application also provide a computer readable storage medium carrying one or more computer programs, which, when executed by an electronic device, can cause the electronic device to implement any of the audio rendering methods provided by the embodiments of the present application.

[0215] In addition, it should be noted that the apparatus embodiments described above are merely illustrative, wherein the units described as separate components can or can not be physically separate, and the components displayed as units can or can not be physical units, i.e., they can be located in one place, or distributed on a plurality of network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the embodiments. In addition, the connection relationship between the modules in the apparatus embodiments provided by the present application indicates that there is a communication connection between them, which can be implemented as one or more communication buses or signal lines.

[0216] Those skilled in the art can clearly understand that the application can be implemented by means of software plus necessary universal hardware, and of course can also be implemented by means of dedicated hardware including special integrated circuit, special CPU, special memory, special component, etc. Generally, any function completed by computer program can be easily implemented by corresponding hardware, and the specific hardware structure for implementing the same function can also be various, such as analog circuit, digital circuit or special circuit, etc. However, for the application, software program implementation is a better embodiment. Based on such understanding, the technical solution of the application or the part of the application which makes contribution to the prior art can be embodied in the form of software product, which is stored in readable storage medium, such as computer floppy disk, U disk, mobile hard disk, ROM, RAM, magnetic disk or optical disk, etc., and includes a plurality of instructions for making a computer device (which can be personal computer, training device or network device, etc.) execute the method described in various embodiments of the application.

[0217] In the above embodiments, the implementation can be achieved by software, hardware, firmware or any combination thereof, entirely or partially. When implemented by software, the implementation can be achieved in the form of a computer program product, entirely or partially.

[0218] The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the flow or function described in the embodiments of the application is generated entirely or partially. The computer can be a general-purpose computer, a special-purpose computer, a computer network or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another, for example, the computer instructions can be transmitted from one website, computer, training device or data center to another website, computer, training device or data center through wired (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (such as infrared, wireless, microwave, etc.) mode. The computer-readable storage medium can be any available medium that can be stored by the computer or a data storage device such as training device, data center, etc. integrated with one or more available media sets. The available medium can be magnetic medium (such as floppy disk, hard disk, magnetic tape), optical medium (such as DVD) or semiconductor medium (such as solid state disk (SSD)) etc.

Claims

1. An audio rendering method, characterized in that, include: Acquire the audio PCM data, Cartesian coordinates, and spherical coordinates of each audio sampling point of each sound object in the target sound stream, wherein the audio PCM data refers to audio pulse code modulation data; For each audio sampling point of each sound object, for every two adjacent sampling points, a sphere is constructed according to the spherical coordinates of the two sampling points respectively, wherein the sphere corresponding to any sampling point represents the range of the sound amplitude of that sampling point; Based on the audio PCM data, Cartesian coordinates, and spherical coordinates of the two sampling points, and combined with the spheres corresponding to the two sampling points, the change in sound amplitude during the spatial movement of the first sampling point is simulated and calculated to obtain the sound amplitude at each intermediate point during the spatial movement process. Based on the sound amplitude at each intermediate point, the binaural rendering signal value of each intermediate point is obtained; thus, the binaural rendering signal value of the target sound stream is obtained.

2. The audio rendering method according to claim 1, characterized in that, The method involves simulating and calculating the change in sound amplitude during the spatial movement of the first of the two sampling points based on their respective audio PCM data, Cartesian coordinates, and spherical coordinates, combined with the spheres corresponding to each sampling point. This yields the sound amplitude at each intermediate point during the spatial movement process, including: Based on the Cartesian coordinates and spherical coordinates of the two sampling points, determine the inclusion relationship of the spheres corresponding to the two sampling points, and the positional relationship between the spheres corresponding to the two sampling points and each intermediate point. Based on the inclusion relationship and the positional relationship between the spheres corresponding to the two sampling points and each intermediate point, as well as the audio PCM data, Cartesian coordinates, and spherical coordinates of the two sampling points, the sound amplitude of each intermediate point is determined.

3. The audio rendering method according to claim 2, characterized in that, The step of determining the inclusion relationship of the spheres corresponding to the two sampling points, and the positional relationship between the spheres corresponding to the two sampling points and each intermediate point, based on their respective Cartesian and spherical coordinates, includes: The sampling point distance between the two sampling points is determined based on their respective Cartesian coordinates. The inclusion relationship of the spheres corresponding to the two sampling points is determined based on the radial distance in the spherical coordinates of the two sampling points and the distance between the sampling points. The duration of each frame of the target audio stream is obtained, and the movement distance corresponding to each intermediate point is determined based on the duration. Based on the movement distance corresponding to each intermediate point, the radial distance in the spherical coordinates of the two sampling points, and the distance between the sampling points, the positional relationship between the spheres corresponding to the two sampling points and each intermediate point is determined.

4. The audio rendering method according to claim 3, characterized in that, The step of determining the sound amplitude of each intermediate point based on the inclusion relationship, the positional relationship between the spheres corresponding to the two sampling points and each intermediate point, and the audio PCM data, Cartesian coordinates, and spherical coordinates of each of the two sampling points includes: Based on the inclusion relationship and the positional relationship between the spheres corresponding to the two sampling points and each intermediate point, the sound amplitude function of each intermediate point is determined; Based on the audio PCM data, Cartesian coordinates, and spherical coordinates of the two sampling points, determine the independent variable in the sound amplitude function of each intermediate point; The independent variable in the sound amplitude function of each intermediate point is substituted into the sound amplitude function of each intermediate point to perform function calculation, so as to obtain the sound amplitude of each intermediate point.

5. The audio rendering method according to claim 1, characterized in that, The step of obtaining the binaural rendering signal value of each intermediate point based on the sound amplitude of each intermediate point includes: Based on the spherical coordinates of the two sampling points, dynamic interpolation calculation based on the head-related transfer function value is performed on each intermediate point in the spatial movement process to obtain the head-related transfer function value of each intermediate point, wherein the head-related transfer function value of each intermediate point represents the image position information of each intermediate point; The binaural rendering signal value of each intermediate point is obtained based on the sound amplitude of each intermediate point and the head-related transfer function value of each intermediate point.

6. The audio rendering method according to claim 5, characterized in that, The step of performing dynamic interpolation calculations based on the head-related transfer function (HRF) values ​​for each intermediate point during the spatial movement, according to the spherical coordinates of the two sampling points respectively, to obtain the HRF values ​​for each intermediate point, includes: Based on the pitch and azimuth angles in the spherical coordinates of the two sampling points, calculate the unit vectors of the Cartesian coordinates of the two sampling points, as well as the audio pulse values ​​of the two sampling points. Based on the unit vector of the Cartesian coordinates of the two sampling points, determine the angle between the two sampling points and the origin of the Cartesian coordinates; The duration of each frame of the target audio stream and the playback duration corresponding to each intermediate point are obtained. Based on the duration, the included angle, and the playback duration corresponding to each intermediate point, the angle change value corresponding to each intermediate point is determined. Based on the unit vector of the Cartesian coordinates of the two sampling points, the playback duration and angle change value corresponding to each intermediate point, the spherical linear interpolation angle of each intermediate point is determined. The head-related transfer function value of each intermediate point is determined based on the spherical linear interpolation angle of each intermediate point and the audio pulse values ​​of the two sampling points.

7. The audio rendering method according to claim 6, characterized in that, The step of determining the head-related transfer function value of each intermediate point based on the spherical linear interpolation angle of each intermediate point and the audio-visual impulse values ​​of the two sampling points includes: For each of the intermediate points: If the spherical linear interpolation angle of the intermediate point is less than the preset first angle value, then the preset target unit time is incremented by 1, and the head-related transfer function value of the previous intermediate point is determined as the head-related transfer function value of the intermediate point. If the intermediate point is the first intermediate point after the previous sampling point, then the head-related transfer function value of the previous intermediate point is the audio frequency pulse value of the previous sampling point. If the spherical linear interpolation angle of the intermediate point is greater than or equal to the first angle value and less than or equal to the preset second angle value, or if the target unit time is equal to the preset counting threshold, then the updated duration corresponding to the intermediate point is determined according to the target unit time and the playback duration corresponding to the intermediate point. The head-related transfer function value of the intermediate point is determined according to the updated duration corresponding to the intermediate point, the audio frequency pulse value of the second sampling point among the two sampling points, and the head-related transfer function value of the previous intermediate point. The target unit time is then set to 0. If the spherical linear interpolation angle of the intermediate point is greater than the second angle value, then the first point that makes the spherical linear interpolation angle greater than the second angle value is determined. Based on the updated duration corresponding to the point, the audio frequency pulse value of the next sampling point, and the head-related transfer function value of the intermediate point preceding the intermediate point, the head-related transfer function value of the intermediate point is determined.

8. The audio rendering method according to claim 5, characterized in that, The step of obtaining the binaural rendering signal value of each intermediate point based on the sound amplitude of each intermediate point and the head-related transfer function value of each intermediate point includes: The sound amplitude and head-related transfer function value at each intermediate point are multiplied by point-to-point convolution to obtain the binaural rendering signal value at each intermediate point.

9. An audio rendering device, characterized in that, include: The basic data acquisition module is used to acquire the audio PCM data, Cartesian coordinates, and spherical coordinates of each audio sampling point of each sound object in the target sound stream, wherein the audio PCM data refers to audio pulse code modulation data; The sphere construction module is used to construct spheres corresponding to the two sampling points respectively based on their respective spherical coordinates. The sphere corresponding to any sampling point represents the range of the sound amplitude of that sampling point. The sound amplitude simulation module is used to simulate and calculate the change in sound amplitude during the spatial movement of the first of the two sampling points based on the audio PCM data, Cartesian coordinates, and spherical coordinates of each of the two sampling points, combined with the spheres corresponding to the two sampling points, so as to obtain the sound amplitude at each intermediate point during the spatial movement process. The binaural rendering module is used to obtain the binaural rendering signal value of each intermediate point based on the sound amplitude of each intermediate point, so as to obtain the binaural rendering signal value of the target sound stream.

10. An electronic device, characterized in that, It includes at least one processor and a memory connected to the processor, wherein: The memory is used to store computer programs; The processor is used to execute the computer program to enable the electronic device to implement the audio rendering method as described in any one of claims 1 to 8.