A loudspeaker rendering method, apparatus, device and storage medium
Patent Information
- Application Number
- CN202611088059.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-22
- Publication Date
- 2026-09-18
- Estimated Expiration
- 2046-07-22
AI Technical Summary
[0003]但经典VBAP在实际使用中存在明显局限:第一,算法仅利用声像方向的单位向量,所有声像均被约束于扬声器所在球面,无法编码声源与听音者之间的距离信息,导致声像距离感缺失,难以还原真实空间听觉感受
[0015]In this application, a target cube space model is first constructed with the listening position as the center of the bottom surface of the cube space model, and physical speakers are mapped onto the edges of the target cube space model to obtain the mapping position information of the physical speakers. Then, the target cube space model is divided into grids to obtain several grid points, and corresponding virtual speakers are generated at each grid point. The virtual speakers include edge virtual speakers corresponding to grid points on the edges of the target cube space model, face virtual speakers corresponding to grid points on the faces, and spatial virtual speakers corresponding to grid points inside the model. Next, based on the type of each virtual speaker and the mapping position information, the gain corresponding to each virtual speaker is determined, and a corresponding lookup table is constructed based on the gain corresponding to each virtual speaker. The gain is used to represent the volume of the physical speaker corresponding to the position of the virtual speaker. Finally, based on the position of the virtual sound source, the target grid point corresponding to the virtual sound source in the target cube space model is located, and the gain of the target grid point is determined from the lookup table. A target gain corresponding to the virtual sound source is generated based on the gain of the target grid point, so as to control the physical speaker to perform audio rendering based on the target gain. As can be seen from the above, this application constructs a target cubic spatial model with the listening position as the center of the bottom face of the cubic spatial model, and maps physical speakers onto the edges of the target cubic spatial model to obtain the mapping position information. Then, the target cubic spatial model is meshed, and three types of virtual speakers—edge, face, and space—are generated at each mesh point. Next, the gain of the corresponding virtual speaker is calculated based on the virtual speaker type and mapping position information, and a gain lookup table is constructed. Finally, the target mesh point to which the virtual sound source belongs is located based on the virtual sound source position, the gain of the target mesh point is obtained from the lookup table, and the target gain corresponding to the virtual sound source is calculated and generated, thereby driving the physical speaker to complete audio rendering. In this way, this application transforms the complex VBAP solution into pre-computation and lookup table interpolation through cubic mesh modeling and hierarchical cascaded computation. While retaining the high efficiency characteristics of classic VBAP, it achieves more stable spatial sound image rendering, effectively improves the listening effect at non-listening positions, supports distance perception, has low computational load, strong adaptability, and can meet the needs of real-time spatial audio playback.
Smart Images

Figure CN122602061B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of audio rendering technology, and in particular to a speaker rendering method, apparatus, device, and storage medium. Background Technology
[0002] VBAP (Vector Base Amplitude Panning) is one of the most widely used speaker rendering algorithms in the field of spatial audio. Its core is to represent the sound image direction vector as a weighted combination of several speaker direction vectors. The gain of each speaker is obtained by solving a system of linear equations. In 3D scenes, it follows the principle of "minimum activation" and usually only uses 2 to 3 speakers to form a triangular rendering area to complete the sound image reproduction.
[0003] However, the classic VBAP algorithm has significant limitations in practical applications: First, it only utilizes unit vectors of sound image direction, constraining all sound images to the sphere where the speaker is located. This fails to encode distance information between the sound source and the listener, resulting in a lack of sound image distance perception and difficulty in reproducing a realistic spatial auditory experience. Second, its rendering is based on a unit spherical model, while content creation and actual playback scenarios often use cube or shoebox models, and speaker layouts are also cubic. This mismatch between the spherical and cubic models introduces directional distortion, reducing the accuracy of sound image localization. Third, limited by a small number of speaker activation mechanisms, the sound image jumps noticeably and lacks a sense of immersion. When the listener deviates from the center sweet spot, i.e., the optimal center listening position, the sound image stability drops significantly. Existing improvements attempt to increase the number of activated speakers and use overlapping triangular regions for optimization, but they fail to fundamentally solve the problems of distance modeling, model adaptation, and stability in non-sweet spots. Furthermore, some solutions significantly increase computational complexity, lacking versatility and practicality.
[0004] In conclusion, how to improve the rendering stability of non-auditory positions while maintaining low computational complexity is a technical problem that urgently needs to be solved. Summary of the Invention
[0005] In view of this, the purpose of this invention is to provide a speaker rendering method, apparatus, device, and storage medium that can improve the rendering stability of non-listening locations while maintaining low computational complexity. The specific solution is as follows: Firstly, this application provides a speaker rendering method, including: A target cube space model is constructed with the listening position as the center of the bottom surface of the cube space model, and the physical speaker is mapped onto the edge of the target cube space model to obtain the mapping position information of the physical speaker; The target cube spatial model is divided into a grid to obtain several grid points, and a corresponding virtual speaker is generated on each grid point; the virtual speakers include edge virtual speakers corresponding to grid points located on the edges of the target cube spatial model, face virtual speakers corresponding to grid points on the faces, and spatial virtual speakers corresponding to grid points inside the model. Based on the type of each virtual speaker and the mapping location information, the gain corresponding to each virtual speaker is determined, and a corresponding lookup table is constructed based on the gain corresponding to each virtual speaker; the gain is used to represent the volume of the physical speaker corresponding to the position of the virtual speaker; Based on the location of the virtual sound source, the target grid point corresponding to the virtual sound source in the target cube space model is located, and the gain of the target grid point is determined from the lookup table. Based on the gain of the target grid point, the target gain corresponding to the virtual sound source is generated so as to control the physical speaker to perform audio rendering based on the target gain.
[0006] Optionally, the step of constructing a target cube space model with the listening position as the center of the bottom surface of the cube space model, and mapping the physical speaker onto the edge of the target cube space model to obtain the mapping position information of the physical speaker, includes: In a preset spatial coordinate system, the listening position is taken as the center of the bottom surface of the cube spatial model, the origin of the coordinate system is taken as any vertex of the bottom surface of the cube spatial model, and the target cube spatial model is constructed based on the unit edge length. A ray is emitted from the listening position toward the physical speaker, and the intersection point of the ray and the surface of the target cube spatial model is obtained; From each edge of the target cube space model, determine the target edge that is closest to the intersection point, and project the intersection point onto the target edge of the target cube space model to obtain the mapping position information of the physical speaker; The preset spatial coordinate system is a spatial coordinate system constructed with the vertical direction of the listening position as the Z-axis, the left and right direction as the X-axis, and the front and back direction as the Y-axis.
[0007] Optionally, the step of dividing the target cube spatial model into a mesh to obtain several mesh points includes: The target cube spatial model is divided into a grid to obtain several cubes of the same size, and each grid point is determined based on the vertices of each cube.
[0008] Optionally, determining the gain corresponding to each virtual speaker based on the type of each virtual speaker and the mapping location information includes: If the virtual speaker is of the type of prism virtual speaker, then the gain corresponding to the prism virtual speaker is determined based on the mapping position information of the physical speaker using a two-dimensional VBAP method or a three-dimensional VBAP method.
[0009] Optionally, determining the gain corresponding to each virtual speaker based on the type of each virtual speaker and the mapping location information includes: If the type of the virtual speaker is the surface virtual speaker, then the first weight corresponding to the surface virtual speaker at the first grid point is calculated using the two-dimensional VBAP method, and the first gain corresponding to the edge virtual speaker at the first grid point is obtained; The second gain of the surface virtual loudspeaker at the first grid point is determined by the product of the first coefficient and the first weight, and the products of the first gain and the second gain at each of the same first grid points are added together to obtain the gain corresponding to the surface virtual loudspeaker. The first grid point includes a first left intersection point and a first right intersection point between a line generated along the X-axis centered on the virtual loudspeaker and the edge of the target cube space model, and a first front intersection point and a first back intersection point between a line generated along the Y-axis and the edge of the target cube space model. The X-axis coordinate of the first left intersection point is smaller than the X-axis coordinate of the first right intersection point, and the Y-axis coordinates of the first front intersection point and the first back intersection point are larger than the Y-axis coordinate of the back intersection point.
[0010] Optionally, determining the gain corresponding to each virtual speaker based on the type of each virtual speaker and the mapping location information includes: If the type of virtual speaker is the spatial virtual speaker, then the second weight corresponding to the spatial virtual speaker at the second grid point is calculated using the two-dimensional VBAP method, and the third gain corresponding to the surface virtual speaker at the second grid point is obtained; The fourth gain of the spatial virtual loudspeaker at the second grid point is determined by the product of the second coefficient and the second weight, and the products of the third gain and the fourth gain at each of the same second grid points are added together to obtain the gain corresponding to the spatial virtual loudspeaker. The second grid point includes the second left and second right intersection points between the line generated along the X-axis centered on the spatial virtual speaker and the face of the target cube spatial model; the second front and second back intersection points between the line generated along the Y-axis and the face of the target cube spatial model; and the upper and lower intersection points between the line generated along the Z-axis and the face of the target cube spatial model. The X-axis coordinate of the second left intersection point is smaller than the X-axis coordinate of the second right intersection point, the Y-axis coordinate of the second front intersection point is larger than the Y-axis coordinate of the second back intersection point, and the Z-axis coordinate of the upper intersection point is larger than the Z-axis coordinate of the lower intersection point.
[0011] Optionally, the step of locating the target grid point corresponding to the virtual sound source in the target cube space model based on the location of the virtual sound source, determining the gain of the target grid point from the lookup table, and generating the target gain corresponding to the virtual sound source based on the gain of the target grid point includes: Based on the location of the virtual sound source, the target cube corresponding to the virtual sound source in the target cube space model is determined, and each target grid point is determined based on each vertex of the target cube; the target cube is a cube obtained by meshing the target cube space model. The weights corresponding to each target grid point are calculated using the trilinear interpolation formula, and the gains of each target grid point are determined from the lookup table. Based on the weights corresponding to each target grid point, the gains of each target grid point are weighted and summed to obtain the initial gain corresponding to the virtual sound source, and the initial gain is normalized to obtain the target gain.
[0012] Secondly, this application provides a speaker rendering apparatus, comprising: The cube space model construction module is used to construct a target cube space model with the listening position as the center of the bottom surface of the cube space model, and to map the physical speaker onto the edge of the target cube space model to obtain the mapping position information of the physical speaker; The virtual speaker generation module is used to divide the target cube spatial model into a grid to obtain a number of grid points, and generate a corresponding virtual speaker on each of the grid points; the virtual speakers include edge virtual speakers corresponding to the grid points located on the edges of the target cube spatial model, face virtual speakers corresponding to the grid points on the faces, and spatial virtual speakers corresponding to the grid points inside the model. The lookup table construction module is used to determine the gain corresponding to each virtual speaker based on the type of each virtual speaker and the mapping position information, and to construct a corresponding lookup table based on the gain corresponding to each virtual speaker; the gain is used to represent the volume of the physical speaker corresponding to the position of the virtual speaker; The audio rendering module is used to locate the target grid point corresponding to the virtual sound source in the target cube space model based on the location of the virtual sound source, determine the gain of the target grid point from the lookup table, generate the target gain corresponding to the virtual sound source based on the gain of the target grid point, so as to control the physical speaker to perform audio rendering based on the target gain.
[0013] Thirdly, this application provides an electronic device, comprising: Memory, used to store computer programs; A processor is used to execute the computer program to implement the aforementioned speaker rendering method.
[0014] Fourthly, this application provides a computer-readable storage medium for storing a computer program; wherein, when the computer program is executed by a processor, it implements the aforementioned speaker rendering method.
[0015] In this application, a target cube space model is first constructed with the listening position as the center of the bottom surface of the cube space model, and physical speakers are mapped onto the edges of the target cube space model to obtain the mapping position information of the physical speakers. Then, the target cube space model is divided into grids to obtain several grid points, and corresponding virtual speakers are generated at each grid point. The virtual speakers include edge virtual speakers corresponding to grid points on the edges of the target cube space model, face virtual speakers corresponding to grid points on the faces, and spatial virtual speakers corresponding to grid points inside the model. Next, based on the type of each virtual speaker and the mapping position information, the gain corresponding to each virtual speaker is determined, and a corresponding lookup table is constructed based on the gain corresponding to each virtual speaker. The gain is used to represent the volume of the physical speaker corresponding to the position of the virtual speaker. Finally, based on the position of the virtual sound source, the target grid point corresponding to the virtual sound source in the target cube space model is located, and the gain of the target grid point is determined from the lookup table. A target gain corresponding to the virtual sound source is generated based on the gain of the target grid point, so as to control the physical speaker to perform audio rendering based on the target gain. As can be seen from the above, this application constructs a target cubic spatial model with the listening position as the center of the bottom face of the cubic spatial model, and maps physical speakers onto the edges of the target cubic spatial model to obtain the mapping position information. Then, the target cubic spatial model is meshed, and three types of virtual speakers—edge, face, and space—are generated at each mesh point. Next, the gain of the corresponding virtual speaker is calculated based on the virtual speaker type and mapping position information, and a gain lookup table is constructed. Finally, the target mesh point to which the virtual sound source belongs is located based on the virtual sound source position, the gain of the target mesh point is obtained from the lookup table, and the target gain corresponding to the virtual sound source is calculated and generated, thereby driving the physical speaker to complete audio rendering. In this way, this application transforms the complex VBAP solution into pre-computation and lookup table interpolation through cubic mesh modeling and hierarchical cascaded computation. While retaining the high efficiency characteristics of classic VBAP, it achieves more stable spatial sound image rendering, effectively improves the listening effect at non-listening positions, supports distance perception, has low computational load, strong adaptability, and can meet the needs of real-time spatial audio playback. Attached Figure Description
[0016] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on the provided drawings without creative effort.
[0017] Figure 1 A flowchart of a speaker rendering method provided in this application; Figure 2 A schematic diagram of a virtual loudspeaker in a specific cubic space model provided in this application; Figure 3 A specific speaker rendering flowchart is provided for this application; Figure 4 This application provides a schematic diagram of the structure of a speaker rendering device; Figure 5 This application provides a structural diagram of an electronic device. Detailed Implementation
[0018] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0019] VBAP is one of the most widely used speaker rendering algorithms in the field of spatial audio. Its core is to represent the sound image direction vector as a weighted combination of several speaker direction vectors. The gain of each speaker is obtained by solving a system of linear equations. In 3D scenes, it follows the "minimum activation" principle, typically using only 2 to 3 speakers to form a triangular rendering area to complete sound image reproduction. However, classic VBAP has significant limitations in practical use: First, the algorithm only utilizes unit vectors of sound image direction, constraining all sound images to the sphere where the speakers are located. It cannot encode the distance information between the sound source and the listener, resulting in a lack of sound image distance perception and difficulty in reproducing a realistic spatial auditory experience. Second, its rendering is based on a unit spherical model, while content creation and actual playback scenes often use cube or shoebox models, and the speaker layout is also cubic. This mismatch between the spherical and cubic models introduces directional distortion, reducing the accuracy of sound image localization. Third, limited by the limited number of speaker activation mechanisms, the sound image has a noticeable jumpiness and insufficient immersion. When the listener deviates from the center sweet spot, the stability of the sound image drops significantly. While existing improvement schemes attempt to optimize by increasing the number of active speakers and using overlapping triangular regions, they fail to fundamentally solve the problems of distance modeling, model adaptation, and stability in non-sweet spots. Furthermore, some schemes significantly increase computational complexity, lacking versatility and practicality. Therefore, this application provides a speaker rendering scheme that can improve rendering stability in non-listening positions while maintaining low computational complexity.
[0020] See Figure 1 As shown, an embodiment of the present invention discloses a speaker rendering method, which may include: Step S11: Construct a target cube space model with the listening position as the center of the bottom surface of the cube space model, and map the physical speaker onto the edge of the target cube space model to obtain the mapping position information of the physical speaker.
[0021] In this embodiment, a target cube space model is constructed with the preset optimal center listening position, i.e., the sweet spot, as the center of the bottom surface of the cube space model. The physical speaker is then mapped onto the edge of the target cube space model to obtain the mapping position information of the physical speaker. The specific process may include: First, in a preset spatial coordinate system, the listening position is taken as the center of the bottom surface of the cube space model, and the origin is taken as any vertex of the bottom surface of the cube space model. The target cube space model is constructed based on a unit edge length. Then, a ray is emitted from the listening position to the physical speaker to obtain the intersection point of the ray with the surface of the target cube space model. Next, from each edge of the target cube space model, the target edge that is closest to the intersection point is determined, and the intersection point is projected onto the target edge of the target cube space model to obtain the mapping position information of the physical speaker. The preset spatial coordinate system is a spatial coordinate system constructed with the vertical direction of the listening position as the Z-axis, the left-right direction as the X-axis, and the front-back direction as the Y-axis.
[0022] Specifically, this embodiment uses a cube coordinate system, with a unit cube as the rendering space, an edge length of 1, and coordinates of the 8 vertices. to Construct a spatial model of the target cube. Setting: Origin The target cube spatial model is located at its lower left rear corner; the x-axis points to the right (x=0 for left, x=1 for right); the y-axis points forward (y=0 for rear, y=1 for front); the z-axis points upward (z=0 for ground, z=1 for ceiling), and the listening position (sweet spot) is the center of the bottom surface of the target cube spatial model. The listener is located on the ground, and the space above covers the entire upper hemisphere. In the actual spatial audio scene, the sound image will not appear below the listener. Therefore, the listening position on the ground will not cause a mapping loss.
[0023] Any point on the surface of the target cube spatial model The direction vector relative to the listener is defined as: ; in, Indicates the coordinates of the listening position. Represents the surface points of the target cube spatial model The unit direction vector relative to the listener.
[0024] For each physical speaker ( , (The total number of physical loudspeakers), a ray is emitted from the listening position along the direction of the physical loudspeakers, and the intersection with the surface of the target cube spatial model yields the mapping point. Then, project the mapping point onto the edge closest to the mapping point (keeping the coordinates along the edge direction, and setting the coordinates perpendicular to the edge to 0 or 1), so that the physical speaker falls on the edge. After mapping, the mapping position information of the physical speaker is obtained, including its position on the edge. and direction vector ;in, Indicates physical loudspeaker The mapping point on the edge.
[0025] It should be noted that when At that time, the physical speaker is pointing below the ground, and the elevation angle is clamped to... (i.e., projected onto the ground plane) The mapping point will fall directly on one of the edges of the bottom surface of the target cube spatial model; where, This indicates the direction from the dessert position to the physical speaker. The direction vector.
[0026] Step S12: Divide the target cube space model into a grid to obtain several grid points, and generate corresponding virtual speakers on each grid point; the virtual speakers include edge virtual speakers corresponding to grid points on the edges of the target cube space model, face virtual speakers corresponding to grid points on the faces, and space virtual speakers corresponding to grid points inside the model.
[0027] In this embodiment, the target cube spatial model is meshed to obtain several cubes of the same size, and each mesh point is determined based on the vertices of each cube. Specifically, the resolution... User-defined, representing the mesh spacing. It is typically taken as the minimum distance between the physical speakers on the surface of the target cubic spatial model. or .Require Let be a positive integer, denoted as That is, each edge is divided into equal parts. Segments, formed on each edge There are 1 grid point, and the coordinates of the grid point are... ;in .
[0028] A corresponding virtual speaker is generated at each grid point. Based on the location of the grid point, the virtual speakers are divided into grid points located on the edges of the target cube space model, i.e., edge virtual speakers corresponding to edge grid points. The grid points located on the surface, i.e., the virtual loudspeakers corresponding to the grid points on the surface. The grid points located inside, that is, the spatial virtual loudspeakers corresponding to the spatial grid points. See also Figure 2 As shown, Figure 2 The blue dots in the image represent the virtual speakers corresponding to the edge grid points. Figure 2 The green dots in the image represent virtual speakers corresponding to the grid points. Figure 2 The purple dots in the image represent the virtual speakers corresponding to the spatial grid points. Figure 2 The red dots in the image indicate the dessert section.
[0029] Step S13: Based on the type of each virtual speaker and the mapping position information, determine the gain corresponding to each virtual speaker, and construct a corresponding lookup table based on the gain corresponding to each virtual speaker; the gain is used to represent the volume of the physical speaker corresponding to the position of the virtual speaker.
[0030] To determine the gain of the virtual speakers, in a first specific embodiment, the gain corresponding to each virtual speaker is determined based on the type and mapping position information of each virtual speaker. Specifically, this may include: if the type of the virtual speaker is the prism virtual speaker, then the gain corresponding to the prism virtual speaker is determined based on the mapping position information of the physical speaker using a two-dimensional VBAP method or a three-dimensional VBAP method.
[0031] Specifically, for the prism virtual speaker The gain vector is obtained by calculating the gain from the prism virtual speaker to each physical speaker using VBAP (2D or 3D) with adjacent physical speakers. Virtual speaker The grid point where it is located happens to have a physical speaker. ,but For the first A unit vector with each element equal to 1; otherwise, calculate the gain from the prism virtual speaker to each physical speaker using VBAP (2D or 3D) based on the adjacent physical speakers, and fill in the space. .
[0032] In a second specific implementation, the gain corresponding to each virtual speaker is determined based on the type and mapping position information of each virtual speaker. Specifically, this may include: if the type of the virtual speaker is the surface virtual speaker, then the first weight corresponding to the surface virtual speaker at the first grid point is calculated using the two-dimensional VBAP method, and the first gain corresponding to the edge virtual speaker at the first grid point is obtained; the second gain of the surface virtual speaker at the first grid point is determined according to the product of the first coefficient and the first weight, and the products of the first gain and the second gain corresponding to each of the same first grid points are added together to obtain the gain corresponding to the surface virtual speaker; wherein, the first grid point includes the first left intersection point and the first right intersection point between the line generated along the X-axis centered on the surface virtual speaker and the edge of the target cube space model, and the first front intersection point and the first back intersection point between the line generated along the Y-axis and the edge of the target cube space model, and the X-axis coordinate of the first left intersection point is less than the X-axis coordinate of the first right intersection point, and the Y-axis coordinates of the first front intersection point and the first back intersection point are greater than the Y-axis coordinate of the back intersection point.
[0033] Specifically, taking the bottom surface of the target cube space model ( Taking one face as an example, the other five faces follow the same pattern. The bottom face already has equally spaced grid points on its four edges, along the vertical lines... ( ) and horizontal lines ( The grid formed by connecting lines in two directions has grid points. This allows us to obtain the virtual loudspeakers corresponding to the grid points located on the bottom surface of the target cube spatial model. .
[0034] Virtual speaker located on the bottom surface For example, its energy is distributed 50% in each of the two directions. Along the horizontal direction (along the horizontal line)... The first left intersection point is obtained by intersecting the edge of the target cube spatial model. (Left endpoint) and the first right intersection point (Right endpoint), along the vertical direction (along the vertical line) The first front intersection point is obtained by intersecting the edges of the target cube spatial model. (Front-end point) and the first back intersection point (Rear endpoint), obtain the first grid point.
[0035] Then, the first weight corresponding to the virtual loudspeaker at the first grid point is calculated using 2D VBAP: The weight of the first left intersection point. The weight of the first right intersection point. The weight of the first anterior intersection point. The weight of the first intersection point.
[0036] Next, obtain the first gain corresponding to the edge virtual speaker at the first grid point: The first gain of the prism virtual loudspeaker at the first left intersection point. The first gain of the prism virtual loudspeaker at the first right intersection point. The first gain of the prism virtual loudspeaker at the first front intersection point. The first gain of the prism virtual loudspeaker at the first intersection point.
[0037] Then, the first coefficient is obtained based on the energy equalization weight. Then, the second gain of the virtual loudspeaker at the first grid point is determined based on the product of the first coefficient and the first weight, as shown below: ; ; ; ; in, The second gain of the virtual loudspeaker at the first left intersection point. The second gain of the virtual loudspeaker at the first right intersection point. The second gain of the virtual loudspeaker at the first front intersection point. The second gain of the virtual loudspeaker at the first post-intersection point.
[0038] Finally, the gains of the virtual loudspeaker to the physical loudspeaker are cascaded. The products of the first gain and the second gain at each identical first grid point are added together to obtain the virtual loudspeaker. To physical speakers Gain As shown below: ; As can be seen above, by performing the above operations on the six faces of the target cube spatial model, the gain vector from the virtual speaker to the physical speaker on each face can be obtained.
[0039] In the third specific implementation, based on the type and mapping location information of each virtual speaker, the gain corresponding to each virtual speaker is determined. Specifically, this may include: if the type of the virtual speaker is the spatial virtual speaker, then using the two-dimensional VBAP method, the second weight corresponding to the spatial virtual speaker at the second grid point is calculated, and the third gain corresponding to the surface virtual speaker at the second grid point is obtained; the fourth gain of the spatial virtual speaker at the second grid point is determined according to the product of the second coefficient and the second weight, and the products between the third gain and the fourth gain corresponding to each identical second grid point are added together to obtain the spatial virtual... The gain corresponding to the loudspeaker; wherein, the second grid point includes the second left intersection point and the second right intersection point between the line generated along the X-axis centered on the spatial virtual loudspeaker and the face of the target cube spatial model, the second front intersection point and the second back intersection point between the line generated along the Y-axis and the face of the target cube spatial model, and the upper intersection point and the lower intersection point between the line generated along the Z-axis and the face of the target cube spatial model, and the X-axis coordinate of the second left intersection point is less than the X-axis coordinate of the second right intersection point, the Y-axis coordinate of the second front intersection point is greater than the Y-axis coordinate of the second back intersection point, and the Z-axis coordinate of the upper intersection point is greater than the Z-axis coordinate of the lower intersection point.
[0040] Specifically, for each space virtual speaker Its energy is distributed 1 / 3 in each of the three directions. The second left intersection point is obtained by intersecting the face of the target cube spatial model along the left-right direction (along the x-axis). (Left endpoint) and the second right intersection point (Right endpoint); the second front intersection point is obtained by intersecting the face of the target cube spatial model along the front-back direction (along the y-axis). (Front-end point) and second rear intersection point (Rear endpoint); the upper intersection point is obtained by intersecting the face of the target cube spatial model along the vertical direction (along the z-axis). (Upper endpoint) and lower intersection point (Lower endpoint), thus obtaining the second grid point.
[0041] Next, the second weight of the spatial virtual loudspeaker at the second grid point is calculated using 2D VBAP: The weight of the second left intersection point. The weight of the second right intersection point. The weight of the second anterior intersection point. The weight of the second subsequent intersection point. The weight of the intersection point. The weight of the lower intersection point.
[0042] Next, obtain the third gain corresponding to the virtual speaker at the second grid point. That is, the virtual speaker at the second grid point. To physical speakers The gain.
[0043] Then, the second coefficient is obtained based on the energy equalization weight. Then, the fourth gain of the spatial virtual loudspeaker at the second grid point is determined based on the product of the second coefficient and the second weight, as shown below: ; ; ; in, The fourth gain for the spatial virtual loudspeaker at the second left intersection point. The fourth gain of the spatial virtual loudspeaker at the second right intersection point. The fourth gain for the spatial virtual loudspeaker at the second front intersection point. The fourth gain of the spatial virtual loudspeaker at the second rear intersection point. The fourth gain of the spatial virtual loudspeaker at the upper intersection point. The fourth gain of the spatial virtual loudspeaker at the lower intersection.
[0044] Finally, gain cascading is performed from the spatial virtual speaker to the physical speaker, and the third gain is applied to the corresponding second grid points. and fourth gain The products of these products are added together to obtain the gain corresponding to the spatial virtual loudspeaker, as shown below: .
[0045] In this embodiment, the gains of the virtual speakers at all grid points are iterated and summarized into a three-dimensional gain lookup table. : ; in, For grid points arrive The gain vector of each physical loudspeaker.
[0046] Due to the cascaded structure, the gain vector of each grid point is typically non-zero only for a small number of physical loudspeakers (e.g., grid points on edges correspond to only 23 physical loudspeakers, while grid points on faces correspond to 48 physical loudspeakers). Therefore, lookup tables... Sparse storage can be used, storing only non-zero gain and its corresponding physical speaker index to reduce memory usage.
[0047] Step S14: Locate the target grid point corresponding to the virtual sound source in the target cube space model based on the location of the virtual sound source, determine the gain of the target grid point from the lookup table, generate the target gain corresponding to the virtual sound source based on the gain of the target grid point, so as to control the physical speaker to perform audio rendering based on the target gain.
[0048] During real-time rendering, the target grid points corresponding to the virtual sound source in the target cube space model are located based on the virtual sound source's position. The gain of the target grid points is determined from a lookup table, and the target gain corresponding to the virtual sound source is generated based on the gain of the target grid points. The specific process may include: first, determining the target cube corresponding to the virtual sound source in the target cube space model based on the virtual sound source's position, and determining each target grid point based on each vertex of the target cube; the target cube is a cube obtained by meshing the target cube space model; then, calculating the weights corresponding to each target grid point using a trilinear interpolation formula, and determining the gain of each target grid point from the lookup table; finally, based on the weights corresponding to each target grid point, the gains of each target grid point are weighted and summed to obtain the initial gain corresponding to the virtual sound source, and the initial gain is normalized to obtain the target gain.
[0049] Specifically, the location information of the sound image, i.e., the virtual sound source, is first mapped to coordinates in the target cube spatial model. If the input is spherical coordinates (azimuth, elevation, distance), it can be converted to the corresponding rectangular coordinates of the target cube spatial model; if the input is rectangular coordinates, it can be used directly.
[0050] The position of the virtual sound source in the target cube spatial model encodes distance information: the closer the sound image is to the surface of the target cube spatial model (farr distance), the more speakers in a few directions are strongly activated after subsequent interpolation, resulting in a strong sense of localization; the closer the sound image is to the center of the target cube spatial model (closer distance), the more speakers are uniformly activated, resulting in a strong sense of immersion.
[0051] Traditional VBAP only utilizes direction information, and the acoustic image always lies on the spherical surface, making distance modeling impossible. In this embodiment, direction and distance are uniformly encoded using the coordinates of the target cube spatial model, naturally supporting distance modeling. If no distance information is available, the acoustic image is located on the surface of the target cube spatial model by default.
[0052] Map the location information of the virtual sound source to coordinates in the target cube space model. Next, locate the index of the target cube corresponding to the virtual sound source in the target cube space model: ; in, The index of the target cube where the sound image is located; Clamp the index of the target cube to .
[0053] The eight vertices of the target cube are the target grid points. ;in, .
[0054] Next, calculate the interpolation parameters: ; in, This represents the relative position parameter of the virtual sound source within the target cube.
[0055] Next, the virtual sound source is represented as a weighted sum of the eight vertices of the target cube using trilinear interpolation. The weight of each target grid point is as follows: ; in, Indicates the weights of the trilinear interpolation. This represents the offset of a target grid point relative to the smallest corner point of the target cube, satisfying the following condition: .
[0056] Next, from the lookup table Find each target grid point in the middle Gain And the weights of the 8 target grid points are compared with the gain table. The initial gain of the original physical loudspeaker is obtained by multiplying and summing the gain vectors of corresponding grid points. : ; Finally, the initial gain is normalized to obtain the target gain of the physical loudspeaker, and then output. : ; in, Indicates physical loudspeaker The target gain, This is the gain vector of the physical speaker, used to control the audio rendering of the physical speaker based on the target gain.
[0057] It should be noted that the computational complexity of the target gain during real-time rendering is independent of the grid spacing. Only steps such as cube positioning and interpolation calculation are required, which eliminates the need for speaker region division and matrix solving in traditional VBAP, making it suitable for embedded real-time implementation.
[0058] In one specific implementation, see Figure 3 As shown, the speaker rendering process is divided into a preprocessing stage and a real-time rendering stage, which can specifically include: The preprocessing stage includes steps S1 to S4 (executed only when the speaker configuration changes): S1 maps the physical speakers onto the edges of the target cube spatial model and performs mesh generation on the target cube spatial model and generates edge virtual speakers; S2 meshes the surface of the target cube spatial model and generates face virtual speakers; S3 meshes the internal space of the target cube spatial model and generates spatial virtual speakers; S4 summarizes the gain data of each virtual speaker and constructs a gain lookup table.
[0059] The real-time rendering stage includes steps S5 and S6: S5 maps the virtual sound source coordinates to the target cube spatial model, locates the target cube, calculates the interpolation parameters using a trilinear interpolation algorithm, calculates the weights of each vertex of the target cube, looks up the gain of each vertex from the lookup table, and calculates the gain of the physical speaker by weighting the gain of each vertex according to the weights; S6 performs energy normalization processing on the gain; finally, it outputs a gain vector that can directly drive the physical speaker, completing the real-time rendering of spatial audio.
[0060] As can be seen from the above, in this embodiment, a target cube space model is first constructed with the listening position as the center of the bottom surface of the cube space model, and the physical speaker is mapped onto the edge of the target cube space model to obtain the mapping position information of the physical speaker; then, the target cube space model is divided into a grid to obtain several grid points, and a corresponding virtual speaker is generated on each grid point; the virtual speaker includes edge virtual speakers corresponding to the grid points on the edge of the target cube space model, face virtual speakers corresponding to the grid points on the face, and space virtual speakers corresponding to the grid points inside; next, based on the type of each virtual speaker and the mapping position information, the gain corresponding to each virtual speaker is determined, and a corresponding lookup table is constructed based on the gain corresponding to each virtual speaker; the gain is used to represent the volume of the physical speaker corresponding to the position of the virtual speaker; finally, based on the position of the virtual sound source, the target grid point corresponding to the virtual sound source in the target cube space model is located, and the gain of the target grid point is determined from the lookup table, and the target gain corresponding to the virtual sound source is generated based on the gain of the target grid point, so as to control the physical speaker to perform audio rendering based on the target gain. As can be seen from the above, in this embodiment, the target cube spatial model is constructed with the listening position as the center of the bottom surface of the cube spatial model, and the physical speaker is mapped onto the edge of the target cube spatial model to obtain the mapping position information. Then, the target cube spatial model is meshed, and three types of virtual speakers—edge, face, and space—are generated at each mesh point. Then, the gain of the corresponding virtual speaker is calculated according to the virtual speaker type and mapping position information, and a gain lookup table is constructed. Finally, the target mesh point to which the virtual sound source belongs is located according to the virtual sound source position, the gain of the target mesh point is obtained from the lookup table, and the target gain corresponding to the virtual sound source is calculated and generated to drive the physical speaker to complete the audio rendering. In this way, this embodiment transforms the complex VBAP solution into pre-computation and lookup table interpolation through cube mesh modeling and hierarchical cascaded calculation. While retaining the high efficiency characteristics of classic VBAP, it achieves more stable spatial sound image rendering, effectively improves the listening effect at non-listening positions and supports distance perception. It has low computational load, strong adaptability, and can meet the needs of real-time spatial audio playback.
[0061] Accordingly, see Figure 4 As shown in the illustration, this application also provides a speaker rendering apparatus, which may include: The cube space model construction module 11 is used to construct a target cube space model with the listening position as the center of the bottom surface of the cube space model, and to map the physical speaker onto the edge of the target cube space model to obtain the mapping position information of the physical speaker; The virtual speaker generation module 12 is used to divide the target cube space model into a grid to obtain a number of grid points, and generate a corresponding virtual speaker on each of the grid points; the virtual speakers include edge virtual speakers corresponding to the grid points located on the edges of the target cube space model, face virtual speakers corresponding to the grid points on the faces, and space virtual speakers corresponding to the grid points inside the target cube space model. The lookup table construction module 13 is used to determine the gain corresponding to each of the virtual speakers based on the type of each virtual speaker and the mapping position information, and to construct a corresponding lookup table based on the gain corresponding to each of the virtual speakers; the gain is used to represent the volume of the physical speaker corresponding to the position of the virtual speaker; The audio rendering module 14 is used to locate the target grid point corresponding to the virtual sound source in the target cube space model based on the location of the virtual sound source, determine the gain of the target grid point from the lookup table, generate the target gain corresponding to the virtual sound source based on the gain of the target grid point, so as to control the physical speaker to perform audio rendering based on the target gain.
[0062] In some specific embodiments, the cube space model construction module 11 may include: A cube space model construction unit is used to construct the target cube space model in a preset spatial coordinate system, with the listening position as the center of the bottom surface of the cube space model and the origin of the coordinate system as any vertex of the bottom surface of the cube space model, based on a unit edge length. An intersection point determination unit is used to emit a ray from the listening position toward the physical speaker and obtain the intersection point of the ray with the surface of the target cube space model; The mapping position information determination unit is used to determine the target edge that is closest to the intersection point from each edge of the target cube space model, and project the intersection point onto the target edge of the target cube space model to obtain the mapping position information of the physical speaker; wherein, the preset spatial coordinate system is a spatial coordinate system constructed with the vertical direction of the listening position as the Z-axis direction, the left and right direction as the X-axis direction, and the front and back direction as the Y-axis direction.
[0063] In some specific embodiments, the virtual speaker generation module 12 may include: The grid point determination unit is used to divide the target cube space model into grids to obtain several cubes of the same size, and to determine each grid point based on the vertices of each cube.
[0064] In some specific embodiments, the lookup table construction module 13 may include: The gain determination unit is used to determine the gain corresponding to the prism virtual speaker based on the mapping position information of the physical speaker, using a two-dimensional VBAP method or a three-dimensional VBAP method if the type of the virtual speaker is the prism virtual speaker.
[0065] In some specific embodiments, the lookup table construction module 13 may include: The first gain acquisition unit is used to calculate the first weight of the surface virtual speaker at the first grid point using the two-dimensional VBAP method if the type of the virtual speaker is the surface virtual speaker, and to obtain the first gain of the edge virtual speaker at the first grid point. The second gain determination unit is used to determine the second gain of the surface virtual loudspeaker at the first grid point based on the product of the first coefficient and the first weight, and to add the products between the first gain and the second gain corresponding to each of the same first grid points to obtain the gain corresponding to the surface virtual loudspeaker. The first grid point includes a first left intersection point and a first right intersection point between a line generated along the X-axis centered on the virtual loudspeaker and the edge of the target cube space model, and a first front intersection point and a first back intersection point between a line generated along the Y-axis and the edge of the target cube space model. The X-axis coordinate of the first left intersection point is smaller than the X-axis coordinate of the first right intersection point, and the Y-axis coordinates of the first front intersection point and the first back intersection point are larger than the Y-axis coordinate of the back intersection point.
[0066] In some specific embodiments, the lookup table construction module 13 may include: The third gain acquisition unit is used to calculate the second weight of the spatial virtual speaker at the second grid point using the two-dimensional VBAP method if the type of the virtual speaker is the spatial virtual speaker, and to obtain the third gain of the surface virtual speaker at the second grid point. The fourth gain determination unit is used to determine the fourth gain of the spatial virtual loudspeaker at the second grid point based on the product of the second coefficient and the second weight, and to add the products between the third gain and the fourth gain corresponding to each of the same second grid points to obtain the gain corresponding to the spatial virtual loudspeaker. The second grid point includes the second left and second right intersection points between the line generated along the X-axis centered on the spatial virtual speaker and the face of the target cube spatial model; the second front and second back intersection points between the line generated along the Y-axis and the face of the target cube spatial model; and the upper and lower intersection points between the line generated along the Z-axis and the face of the target cube spatial model. The X-axis coordinate of the second left intersection point is smaller than the X-axis coordinate of the second right intersection point, the Y-axis coordinate of the second front intersection point is larger than the Y-axis coordinate of the second back intersection point, and the Z-axis coordinate of the upper intersection point is larger than the Z-axis coordinate of the lower intersection point.
[0067] In some specific embodiments, the audio rendering module 14 may include: The target grid point determination unit is used to determine the target cube corresponding to the virtual sound source in the target cube space model based on the location of the virtual sound source, and to determine each target grid point based on each vertex of the target cube; the target cube is a cube obtained by meshing the target cube space model; The weight calculation unit is used to calculate the weight corresponding to each of the target grid points using the trilinear interpolation formula, and to determine the gain of each of the target grid points from the lookup table. The target gain determination unit is used to perform a weighted summation of the gains of each target grid point based on the weights corresponding to each target grid point to obtain the initial gain corresponding to the virtual sound source, and to normalize the initial gain to obtain the target gain.
[0068] Furthermore, embodiments of this application also disclose an electronic device, Figure 5 This is a structural diagram of an electronic device 20 according to an exemplary embodiment. The content of the diagram should not be construed as limiting the scope of this application. The electronic device 20 may specifically include: at least one processor 21, at least one memory 22, a power supply 23, a communication interface 24, an input / output interface 25, and a communication bus 26. The memory 22 stores a computer program, which is loaded and executed by the processor 21 to implement the relevant steps in the speaker rendering method disclosed in any of the foregoing embodiments. Furthermore, the electronic device 20 in this embodiment may specifically be an electronic computer.
[0069] In this embodiment, the power supply 23 is used to provide operating voltage for each hardware device on the electronic device 20; the communication interface 24 can create a data transmission channel between the electronic device 20 and external devices, and the communication protocol it follows can be any communication protocol applicable to the technical solution of this application, and is not specifically limited here; the input / output interface 25 is used to acquire external input data or output data to the outside world, and its specific interface type can be selected according to specific application needs, and is not specifically limited here.
[0070] In addition, the memory 22, as a carrier for resource storage, can be a read-only memory, random access memory, disk or optical disk, etc. The resources stored thereon can include operating system 221, computer program 222, etc., and the storage method can be temporary storage or permanent storage.
[0071] The operating system 221 is used to manage and control the various hardware devices on the electronic device 20 and the computer program 222, which may be Windows Server, Netware, Unix, Linux, etc. In addition to including a computer program capable of performing the speaker rendering method executed by the electronic device 20 as disclosed in any of the foregoing embodiments, the computer program 222 may further include computer programs capable of performing other specific tasks.
[0072] Furthermore, this application also discloses a computer-readable storage medium for storing a computer program; wherein, when the computer program is executed by a processor, it implements the aforementioned speaker rendering method. Specific steps of this method can be found in the corresponding content disclosed in the foregoing embodiments, and will not be repeated here.
[0073] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on its differences from other embodiments. Similar or identical parts between embodiments can be referred to interchangeably. For the apparatus disclosed in the embodiments, since it corresponds to the method disclosed in the embodiments, the description is relatively simple; relevant parts can be referred to in the method section.
[0074] Those skilled in the art will further recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of the various examples have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0075] The steps of the methods or algorithms described in conjunction with the embodiments disclosed herein can be implemented directly by hardware, a software module executed by a processor, or a combination of both. The software module can be located in random access memory (RAM), main memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, hard disk, removable disk, CD-ROM, or any other form of storage medium known in the art.
[0076] Finally, it should be noted that in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.
[0077] The technical solutions provided in this application have been described in detail above. Specific examples have been used to illustrate the principles and implementation methods of this application. The descriptions of the above embodiments are only for the purpose of helping to understand the methods and core ideas of this application. At the same time, for those skilled in the art, there will be changes in the specific implementation methods and application scope based on the ideas of this application. Therefore, the content of this specification should not be construed as a limitation of this application.
Claims
1. A speaker rendering method, characterized in that, include: A target cube space model is constructed with the listening position as the center of the bottom surface of the cube space model, and the physical speaker is mapped onto the edge of the target cube space model to obtain the mapping position information of the physical speaker; The target cube spatial model is divided into a grid to obtain several grid points, and a corresponding virtual speaker is generated on each grid point; the virtual speakers include edge virtual speakers corresponding to grid points located on the edges of the target cube spatial model, face virtual speakers corresponding to grid points on the faces, and spatial virtual speakers corresponding to grid points inside the model. Based on the type of each virtual speaker and the mapping location information, the gain corresponding to each virtual speaker is determined, and a corresponding lookup table is constructed based on the gain corresponding to each virtual speaker. The gain is used to represent the volume of the physical speaker corresponding to the position of the virtual speaker; Based on the location of the virtual sound source, the target grid point corresponding to the virtual sound source in the target cube space model is located, and the gain of the target grid point is determined from the lookup table. Based on the gain of the target grid point, the target gain corresponding to the virtual sound source is generated so as to control the physical speaker to perform audio rendering based on the target gain. The step of determining the gain corresponding to each virtual speaker based on the type of each virtual speaker and the mapping position information includes: If the virtual speaker is of the type of prism virtual speaker, then the gain corresponding to the prism virtual speaker is determined based on the mapping position information of the physical speaker using a two-dimensional VBAP method or a three-dimensional VBAP method. The step of determining the gain corresponding to each virtual speaker based on the type of each virtual speaker and the mapping position information includes: If the type of the virtual speaker is the surface virtual speaker, then the first weight corresponding to the surface virtual speaker at the first grid point is calculated using the two-dimensional VBAP method, and the first gain corresponding to the edge virtual speaker at the first grid point is obtained; The second gain of the surface virtual loudspeaker at the first grid point is determined by the product of the first coefficient and the first weight, and the products of the first gain and the second gain at each of the same first grid points are added together to obtain the gain corresponding to the surface virtual loudspeaker. Wherein, the first coefficient is The first grid point includes a first left intersection point and a first right intersection point between the line generated along the X-axis centered on the virtual loudspeaker and the edge of the target cube space model, and a first front intersection point and a first back intersection point between the line generated along the Y-axis and the edge of the target cube space model. The X-axis coordinate of the first left intersection point is smaller than the X-axis coordinate of the first right intersection point, and the Y-axis coordinate of the first front intersection point is larger than the Y-axis coordinate of the first back intersection point. The step of determining the gain corresponding to each virtual speaker based on the type of each virtual speaker and the mapping position information includes: If the type of virtual speaker is the spatial virtual speaker, then the second weight corresponding to the spatial virtual speaker at the second grid point is calculated using the two-dimensional VBAP method, and the third gain corresponding to the surface virtual speaker at the second grid point is obtained; The fourth gain of the spatial virtual loudspeaker at the second grid point is determined by the product of the second coefficient and the second weight, and the products of the third gain and the fourth gain at each of the same second grid points are added together to obtain the gain corresponding to the spatial virtual loudspeaker. Wherein, the second coefficient is The second grid point includes the second left and second right intersection points between the line generated along the X-axis centered on the spatial virtual speaker and the face of the target cube spatial model; the second front and second back intersection points between the line generated along the Y-axis and the face of the target cube spatial model; and the upper and lower intersection points between the line generated along the Z-axis and the face of the target cube spatial model. The X-axis coordinate of the second left intersection point is smaller than the X-axis coordinate of the second right intersection point, the Y-axis coordinate of the second front intersection point is larger than the Y-axis coordinate of the second back intersection point, and the Z-axis coordinate of the upper intersection point is larger than the Z-axis coordinate of the lower intersection point.
2. The speaker rendering method according to claim 1, characterized in that, The process of constructing a target cube space model with the listening position as the center of the bottom surface of the cube space model, and mapping the physical speaker onto the edges of the target cube space model to obtain the mapping position information of the physical speaker includes: In a preset spatial coordinate system, the listening position is taken as the center of the bottom surface of the cube spatial model, the origin of the coordinate system is taken as any vertex of the bottom surface of the cube spatial model, and the target cube spatial model is constructed based on the unit edge length. A ray is emitted from the listening position toward the physical speaker, and the intersection point of the ray and the surface of the target cube spatial model is obtained; From each edge of the target cube space model, determine the target edge that is closest to the intersection point, and project the intersection point onto the target edge of the target cube space model to obtain the mapping position information of the physical speaker; The preset spatial coordinate system is a spatial coordinate system constructed with the vertical direction of the listening position as the Z-axis, the left and right direction as the X-axis, and the front and back direction as the Y-axis.
3. The speaker rendering method according to claim 2, characterized in that, The process of dividing the target cube spatial model into a grid to obtain several grid points includes: The target cube spatial model is divided into a grid to obtain several cubes of the same size, and each grid point is determined based on the vertices of each cube.
4. The speaker rendering method according to claim 3, characterized in that, The process of locating the target grid point corresponding to the virtual sound source in the target cube space model based on the virtual sound source's location, determining the gain of the target grid point from the lookup table, and generating the target gain corresponding to the virtual sound source based on the gain of the target grid point includes: Based on the location of the virtual sound source, the target cube corresponding to the virtual sound source in the target cube space model is determined, and each target grid point is determined based on each vertex of the target cube; the target cube is a cube obtained by meshing the target cube space model. The weights corresponding to each target grid point are calculated using the trilinear interpolation formula, and the gains of each target grid point are determined from the lookup table. Based on the weights corresponding to each target grid point, the gains of each target grid point are weighted and summed to obtain the initial gain corresponding to the virtual sound source, and the initial gain is normalized to obtain the target gain.
5. A speaker rendering device, characterized in that, include: The cube space model construction module is used to construct a target cube space model with the listening position as the center of the bottom surface of the cube space model, and to map the physical speaker onto the edge of the target cube space model to obtain the mapping position information of the physical speaker; The virtual speaker generation module is used to divide the target cube spatial model into a grid to obtain a number of grid points, and generate a corresponding virtual speaker on each of the grid points; the virtual speakers include edge virtual speakers corresponding to the grid points located on the edges of the target cube spatial model, face virtual speakers corresponding to the grid points on the faces, and spatial virtual speakers corresponding to the grid points inside the model. The lookup table construction module is used to determine the gain corresponding to each virtual speaker based on the type of each virtual speaker and the mapping position information, and to construct a corresponding lookup table based on the gain corresponding to each virtual speaker; The gain is used to represent the volume of the physical speaker corresponding to the position of the virtual speaker; The audio rendering module is used to locate the target grid point corresponding to the virtual sound source in the target cube space model based on the location of the virtual sound source, determine the gain of the target grid point from the lookup table, generate the target gain corresponding to the virtual sound source based on the gain of the target grid point, so as to control the physical speaker to perform audio rendering based on the target gain. The lookup table construction module includes: The gain determination unit is used to determine the gain corresponding to the prism virtual speaker based on the mapping position information of the physical speaker by using a two-dimensional VBAP method or a three-dimensional VBAP method if the type of the virtual speaker is the prism virtual speaker. The first gain acquisition unit is used to calculate the first weight of the surface virtual speaker at the first grid point using the two-dimensional VBAP method if the type of the virtual speaker is the surface virtual speaker, and to obtain the first gain of the edge virtual speaker at the first grid point. The second gain determination unit is configured to determine the second gain of the surface virtual loudspeaker at the first grid point based on the product of a first coefficient and a first weight, and to add the products of the first gain and the second gain at each identical first grid point to obtain the gain corresponding to the surface virtual loudspeaker; wherein, the first coefficient is... The first grid point includes a first left intersection point and a first right intersection point between the line generated along the X-axis centered on the virtual loudspeaker and the edge of the target cube space model, and a first front intersection point and a first back intersection point between the line generated along the Y-axis and the edge of the target cube space model. The X-axis coordinate of the first left intersection point is smaller than the X-axis coordinate of the first right intersection point, and the Y-axis coordinate of the first front intersection point is larger than the Y-axis coordinate of the first back intersection point. The third gain acquisition unit is used to calculate the second weight of the spatial virtual speaker at the second grid point using the two-dimensional VBAP method if the type of the virtual speaker is the spatial virtual speaker, and to obtain the third gain of the surface virtual speaker at the second grid point. The fourth gain determination unit is used to determine the fourth gain of the spatial virtual loudspeaker at the second grid point based on the product of the second coefficient and the second weight, and to add the products of the third gain and the fourth gain corresponding to each identical second grid point to obtain the gain corresponding to the spatial virtual loudspeaker; wherein, the second coefficient is... The second grid point includes the second left and second right intersection points between the line generated along the X-axis centered on the spatial virtual speaker and the face of the target cube spatial model; the second front and second back intersection points between the line generated along the Y-axis and the face of the target cube spatial model; and the upper and lower intersection points between the line generated along the Z-axis and the face of the target cube spatial model. The X-axis coordinate of the second left intersection point is smaller than the X-axis coordinate of the second right intersection point, the Y-axis coordinate of the second front intersection point is larger than the Y-axis coordinate of the second back intersection point, and the Z-axis coordinate of the upper intersection point is larger than the Z-axis coordinate of the lower intersection point.
6. An electronic device, characterized in that, The electronic device includes a processor and a memory; wherein the memory is used to store a computer program, which is loaded and executed by the processor to implement the speaker rendering method as described in any one of claims 1 to 4.
7. A computer-readable storage medium, characterized in that, Used to store a computer program, which, when executed by a processor, implements the speaker rendering method as described in any one of claims 1 to 4.
Citation Information
Patent Citations
Rendering of audio objects with apparent size to arbitrary loudspeaker layouts
CN105075292A
Rendering audio objects having apparent size
CN110603821A