Surround sound generation method and device, readable storage medium and computer program product
By detecting changes in user posture and regenerating surround sound signals, the problem of spatial and orientational perception in surround sound technology during posture changes is solved, dynamic sound field adjustment is achieved, and surround sound effects are improved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-09-26
- Publication Date
- 2026-03-27
AI Technical Summary
Existing surround sound technology cannot effectively guarantee spatial awareness and orientation when the listener's position changes, resulting in poor surround sound effects.
By detecting changes in the user's pose, the current pose is obtained, and a surround sound signal is regenerated based on the current pose. The signal is then output using a preset sound source array to ensure that the user is always in the sweet spot of the sound field, thus achieving dynamic adjustment of the sound source.
Maintain the spatial and directional sense of the surround sound signal as the user's posture changes, ensuring that the sweet spot of the sound field changes with the user's posture, thus enhancing the surround sound effect.
Smart Images

Figure CN121751072A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of spatial audio technology, and in particular to a surround sound generation method, apparatus, readable storage medium, and computer program product. Background Technology
[0002] Spatial audio, also known as 3D audio, is an audio technology designed to simulate or reproduce how sound propagates in real three-dimensional space. By simulating the location, distance, motion, and spatial environmental characteristics of sound, it provides listeners with an immersive auditory experience, just as if sound were naturally propagating in physical space.
[0003] Surround sound technology plays a crucial role in spatial audio, creating an immersive auditory experience by simulating sound propagation in the real world. Typically, surround sound technology utilizes multiple sound sources arranged in a specific layout around the listener to create a multi-dimensional surround sound with a sense of space and direction, giving the listener a feeling of being present. However, because surround sound signals are always output using a pre-defined rendering method, the resulting surround sound field is also fixed. When the listener moves or rotates, causing changes in posture, the final surround sound may lack a sense of space and direction, resulting in poor surround sound effects.
[0004] Therefore, how to ensure the spatial and directional sense of surround sound when the listener's posture changes, so as to ensure surround sound effect, is a technical problem that urgently needs to be solved in this field. Summary of the Invention
[0005] The main purpose of this application is to provide a surround sound generation method, device, readable storage medium, and computer program product, which aims to solve the technical problem of how to ensure the spatial sense and orientation of surround sound when the listener's posture changes, so as to ensure surround sound effect.
[0006] To achieve the above objectives, this application provides a surround sound generation method, the surround sound generation method comprising:
[0007] When a change in the user's pose is detected, the user's current pose is obtained, wherein the current pose includes the current position and the current posture angle;
[0008] The current pose is used to regenerate a surround sound signal, and the regenerated surround sound signal is output through a preset sound source array so that the current pose is located in the sweet spot of the sound field formed by the regenerated surround sound signal.
[0009] In one embodiment, the step of regenerating the surround sound signal based on the current pose includes:
[0010] Obtain the user's reference pose, wherein the reference pose includes a reference position and a reference attitude angle;
[0011] Calculate the pose change between the reference pose and the current pose;
[0012] The virtual position of the sound source in the preset sound source array is determined based on the pose change amount, wherein the virtual position is the position of the sound source in the preset coordinate system after the pose change amount is adjusted.
[0013] The virtual location is input into a preset rendering system to obtain a surround sound signal.
[0014] In one embodiment, the pose change includes a position change and an attitude angle change, and the step of determining the virtual position of the sound source in the preset sound source array based on the pose change includes:
[0015] Determine the first position in the preset coordinate system after the position change of the sound source in the preset sound source array is moved by the amount of the position change, and / or determine the second position in the preset coordinate system after the attitude angle change of the sound source in the preset sound source array is rotated by the amount of the attitude angle change;
[0016] The virtual positions of the sound sources in the preset sound source array are determined based on the first position and / or the second position.
[0017] In one embodiment, the attitude angle change includes pitch angle change, roll angle change, and yaw angle change. The step of determining the second position of the sound sources in the preset sound source array in the preset coordinate system after rotating by the attitude angle change includes:
[0018] After determining the pitch angle change, roll angle change, and yaw angle change of the sound sources in the preset sound source array around the horizontal axis of the preset coordinate system, the second position in the preset coordinate system is determined.
[0019] In one embodiment, the step of determining the second position in the preset coordinate system after the change in the attitude angle of the sound sources in the preset sound source array is rotated includes:
[0020] Obtain the initial angle information of the sound sources in the preset sound source array, wherein the initial angle information includes the azimuth angle and the elevation angle;
[0021] The initial angle information is converted into a quaternion to obtain a first position vector, and the current pose is converted into a quaternion to obtain a second position vector;
[0022] The target position vector is obtained by multiplying the first position vector and the second position vector.
[0023] Extract the imaginary part of the target position vector, and determine the extracted imaginary part as the second position of the sound source in the preset coordinate system after the change in the attitude angle.
[0024] In one embodiment, the step of converting the initial angle information into a quaternion to obtain a first position vector includes:
[0025] The initial angle information is converted to a preset coordinate system to obtain the initial position information;
[0026] The initial position information is converted into a quaternion to obtain a first position vector, wherein the real part of the first position vector is zero, and the imaginary part of the first position vector is the initial position information.
[0027] In one embodiment, the steps of determining the first position in the preset coordinate system after the sound source in the preset sound source array moves by the change in position, and / or determining the second position in the preset coordinate system after the sound source in the preset sound source array rotates by the change in attitude angle, include:
[0028] If the change in position is not zero and the change in attitude angle is zero, then determine the first position of the sound source after the change in position.
[0029] If the position change is zero and the attitude angle change is not zero, then determine the second position of the sound source after rotating by the attitude angle change.
[0030] If the change in position is not zero and the change in attitude angle is not zero, then determine the first position of the sound source after the change in position and the second position of the sound source after the change in attitude angle.
[0031] In addition, to achieve the above objectives, this application also provides a surround sound generation device, which includes: a memory, a processor, and a computer program stored in the memory and executable on the processor, the computer program being configured to implement the steps of the surround sound generation method as described above.
[0032] In addition, to achieve the above objectives, this application also provides a readable storage medium, which is a computer-readable storage medium, on which a program implementing the surround sound generation method is stored, and the program implementing the surround sound generation method is executed by a processor to implement the steps of the surround sound generation method as described above.
[0033] This application also provides a computer program product, including a computer program that, when executed by a processor, implements the steps of the surround sound generation method described above.
[0034] One or more technical solutions proposed in this application have at least the following technical effects:
[0035] When a change in user pose is detected, the user's current pose is acquired, including the current position and current posture angle. A surround sound signal is regenerated based on the current pose, and the regenerated surround sound signal is output through a preset sound source array, so that the current pose is located within the sweet spot of the sound field formed by the regenerated surround sound signal. Thus, in this embodiment, when a change in user pose is detected, the user's current pose is acquired, and a surround sound signal is regenerated based on the user's current pose, ensuring that the user always remains within the sweet spot of the sound field formed by the surround sound signal. It can be understood that the sweet spot refers to the area within a certain spatial region where the sound quality, loudness, stereo effect, and sound image localization all reach their optimal listening experience. In other words, when the user moves and / or rotates, resulting in a change in posture, the surround sound signal is always regenerated based on the user's current posture, ensuring that the user is located in the sweet spot of the sound field. This gives the user (or listener) the feeling that the sound source is "moving" with their movement and / or "rotating" with their rotation (even though the sound source does not actually move and / or rotate). The sound source no longer outputs the surround sound signal in a fixed rendering manner, thus ensuring that even when the user moves and / or rotates, resulting in a change in posture, the sweet spot of the sound field changes with the user's posture. This results in a surround sound signal with a better sense of space and orientation, and a better surround sound effect. Attached Figure Description
[0036] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.
[0037] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0038] Figure 1 This is a schematic flowchart of the first embodiment of the surround sound generation method of this application;
[0039] Figure 2 This is a schematic flowchart of the second embodiment of the surround sound generation method of this application;
[0040] Figure 3 This is a schematic flowchart of the third embodiment of the surround sound generation method of this application;
[0041] Figure 4 This is a schematic diagram of the surround sound generation device of this application;
[0042] Figure 5 This is a schematic diagram of the hardware operating environment of the surround sound generation device in the embodiments of this application.
[0043] The purpose, features, and advantages of this application will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation
[0044] To make the above-mentioned objects, features, and advantages of the present invention more apparent and understandable, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0045] Spatial audio, also known as 3D audio, allows users to be fully immersed in a virtual three-dimensional audio space. Dynamic spatial audio, on the other hand, refers to the ability to dynamically track the sound field by capturing the user's head movements in real time, resulting in a more vivid and immersive spatial audio effect.
[0046] Surround sound technology plays a crucial role in spatial audio, creating an immersive auditory experience by simulating sound propagation in the real world. Surround sound technology utilizes multiple sound sources arranged in a specific layout around the listener to create a multi-dimensional sound space. The positions of the multiple sound sources in surround sound are fixed relative to the listener. When the listener moves or rotates, the position of each channel relative to the user changes accordingly. Properly handling these changes can make the spatial sense and orientation of surround sound more realistic and vivid.
[0047] Virtual surround sound technology based on head-related transfer function (HRTF) can process stereo sound through algorithms and play it back using ordinary headphones to achieve the effect of the sound source moving with the head. However, the computational cost of HRTF function is very high, especially when real-time processing and personalized calibration are required. To achieve a realistic experience, customized steps and personalized data are needed, which may require frequent testing and adjustment of an individual's HRTF, making it difficult and complex.
[0048] Based on this, the main solution of this application is: when a change in the user's pose is detected, the user's current pose is obtained, wherein the current pose includes the current position and the current attitude angle; a surround sound signal is regenerated based on the current pose, and the regenerated surround sound signal is output through a preset sound source array so that the current pose is located in the sweet spot of the sound field formed by the regenerated surround sound signal.
[0049] This application, upon detecting a change in user pose, acquires the user's current pose and regenerates a surround sound signal based on that pose. This ensures the user remains within the sweet spot of the surround sound field. The sweet spot, as understood, refers to the area within a spatial region where sound quality, loudness, stereo effect, and sound image localization are at their optimal listening experience. In other words, when the user moves and / or rotates, causing a change in pose, the surround sound signal is always regenerated based on the user's current pose, ensuring the user is within the sweet spot. This gives the user (or listener) the perception that the sound source "moves" with their movement and / or "rotates" with their rotation (even though the sound source itself does not move or rotate). The sound source no longer outputs the surround sound signal in a fixed rendering manner, thus ensuring that even when the user moves and / or rotates, causing a change in pose, the sweet spot changes with the user's pose. This results in a surround sound signal with good spatial awareness and directionality, and a superior surround sound effect.
[0050] It should be noted that the execution subject of the various embodiments of the surround sound generation method of this application can be a computing service device with data processing, network communication, and program execution functions, such as a tablet computer, personal computer, or mobile phone, or a surround sound generation device capable of realizing the above functions, such as headphones, AR (Augmented Reality) glasses, VR (Virtual Reality) glasses, AR helmets, VR helmets, etc. The various embodiments of the surround sound generation method of this application do not impose specific limitations in this regard.
[0051] Based on this, this application proposes a surround sound generation method according to the first embodiment, please refer to... Figure 1 The surround sound generation method includes steps S10 to S20:
[0052] Step S10: When a change in the user's pose is detected, the user's current pose is obtained, wherein the current pose includes the current position and the current pose angle;
[0053] The user's pose can be periodically acquired, and the current pose is the user's pose acquired in this pose acquisition cycle.
[0054] It should be noted that the pose acquisition cycle described in this embodiment starts from execution step S10 and ends after step S20 is completed, and then returns to execution step S10 to enter the next pose acquisition cycle.
[0055] To acquire the user's pose, known methods can be used to collect the user's pose, thus enabling the acquisition of the collected pose. In one specific embodiment, the user can wear a head-mounted smart device, such as smart audio glasses, an AR head-mounted all-in-one device, or a VR head-mounted all-in-one device. The head-mounted smart device is equipped with a pose sensor for collecting the user's pose. Specifically, the pose sensor may include a positioning sensor and an inertial measurement unit (IMU). The positioning sensor can collect the user's position, and the IMU can collect the user's attitude angle. The positioning sensor can specifically be a GPS (Global Positioning System) sensor.
[0056] Pose includes the user's position and attitude angles. The user's position can be the user's position in the world coordinate system, specifically the position of the user's head in the world coordinate system. This world coordinate system can be a right-handed coordinate system, and the user's position in the world coordinate system can be represented as (x, y, z). The attitude angles can be Euler angles, specifically the Euler angles of the user's head, which can be denoted as (P, R, Y), where P represents the pitch angle, Y represents the yaw angle, and R represents the roll angle (also known as the roll angle, yaw angle, etc.).
[0057] It should be noted that the user's pose from the previous pose acquisition cycle (hereinafter referred to as the historical pose) can be obtained. If the historical pose is inconsistent with the current pose, it is determined that the user's pose has changed. If the historical pose is consistent with the current pose, it is determined that the user's pose has not changed, and the surround sound generation process for this cycle can be terminated. Since the historical pose is consistent with the current pose, the user's pose has not changed, so the surround sound signal does not need to be regenerated, thus avoiding the repeated generation of the surround sound signal.
[0058] Step S20: Regenerate the surround sound signal based on the current pose, and output the regenerated surround sound signal through a preset sound source array so that the current pose is located in the sweet spot of the sound field formed by the regenerated surround sound signal.
[0059] It should be noted that, in order to make the current pose located in the sweet spot of the sound field formed by the regenerated surround sound signal, specifically, when the user is in the current pose, the user is located in the sweet spot of the sound field formed by the regenerated surround sound signal output by the preset sound source array. Specifically, when the user's head is in the current pose, the user's head is located in the sweet spot of the sound field formed by the regenerated surround sound signal output by the preset sound source array.
[0060] The surround sound signal is regenerated based on the current pose. Specifically, the surround sound signal can be re-rendered and generated using a certain rendering system, such as the Ambisonic rendering system.
[0061] The preset sound source array can be a pre-arranged sound source array, which can consist of multiple sound sources. For example, for a 5.1 surround sound system, the preset sound source array can specifically include a left speaker, a front right speaker, a center speaker, a rear left speaker, and a rear right speaker.
[0062] The surround sound signal can be output through a preset sound source array, meaning the surround sound signal can be output by each sound source in the preset sound source array. Specifically, if the regenerated surround sound signal is a digital signal, it can be converted into an analog signal using a digital-to-analog converter (DAC), because speakers are analog devices and require analog signals to function. After conversion, these signals are usually very weak and can be amplified by a power amplifier to drive the speakers to output sufficient sound. The amplified analog signal is then distributed to the corresponding sound source according to the channel. For example, in a 5.1 surround sound system, each speaker receives the signal for its corresponding channel. Finally, each sound source (speaker) receives the amplified analog signal and outputs sound. These sounds mix in space to create a surround sound effect.
[0063] In this embodiment, upon detecting a change in the user's pose, the user's current pose is obtained, including the current position and current posture angle. A surround sound signal is regenerated based on the current pose, and output through a preset sound source array, so that the current pose is located within the sweet spot of the sound field formed by the regenerated surround sound signal. Thus, in this embodiment, when a change in the user's pose is detected, the user's current pose is obtained, and a surround sound signal is regenerated based on the user's current pose, ensuring that the user always remains within the sweet spot of the sound field formed by the surround sound signal. It can be understood that the sweet spot refers to the area within a certain spatial region where the sound quality, loudness, stereo effect, and sound image localization all reach their optimal listening experience. In other words, when the user moves and / or rotates, resulting in a change in posture, the surround sound signal is always regenerated based on the user's current posture, ensuring that the user is located in the sweet spot of the sound field. This gives the user (or listener) the feeling that the sound source is "moving" with their movement and / or "rotating" with their rotation (even though the sound source does not actually move and / or rotate). The sound source no longer outputs the surround sound signal in a fixed rendering manner, thus ensuring that even when the user moves and / or rotates, resulting in a change in posture, the sweet spot of the sound field changes with the user's posture. This results in a surround sound signal with a better sense of space and orientation, and a better surround sound effect.
[0064] Based on the first embodiment of this application, in the second embodiment of this application, the content that is the same as or similar to that in Embodiment 1 above can be referred to the above description, and will not be repeated hereafter. On this basis, refer to Figure 2 As shown, the step of regenerating the surround sound signal based on the current pose includes:
[0065] Step A10: Obtain the user's reference pose, wherein the reference pose includes a reference position and a reference attitude angle;
[0066] Understandably, when initially setting up the sound source array, the method for rendering and generating audio signals is usually set, and the rendered audio signals (hereinafter referred to as the initial audio signals) are given to the sound source array for output. Based on this, any pose that can be located in the sweet spot of the sound field formed when the preset sound source array outputs the initial audio signals is denoted as the initial pose.
[0067] In one implementation, the reference pose can be the initial pose.
[0068] As another implementation, the reference pose can also be a historical pose (if the current pose acquisition cycle is the first pose acquisition cycle, then the reference pose can be the initial pose). In this case, the position in the reference pose is recorded as the historical position, and the attitude angle in the reference pose is recorded as the historical attitude angle.
[0069] Step A20: Calculate the pose change between the reference pose and the current pose;
[0070] The pose change can specifically be the pose difference between the initial pose and the current pose. For example, if the initial pose is (x0, y0, z0, P0, R0, Y0) and the current pose is (x1, y1, z1, P1, R1, Y1), then the pose change can be (x1-x0, y1-y0, z1-z0, P1-P0, R1-R0, Y1-Y0). Among them, (x1-x0, y1-y0, z1-z0) can be denoted as △(x,y,z), and (P1-P0, R1-R0, Y1-Y0) can be denoted as △θ.
[0071] Step A30: Determine the virtual position of the sound source in the preset sound source array based on the pose change amount, wherein the virtual position is the position of the sound source in the preset coordinate system after adjusting the pose change amount;
[0072] Specifically, the virtual position can be the position of the assumed sound source in a preset coordinate system after adjusting its pose by this amount of change from its initial pose. It can be understood that after the sound source array is laid out, the poses of each sound source in the array are also fixed, and this pose can be considered the initial pose of the sound source. Considering that the actual pose of the sound source is usually fixed and cannot be changed, this embodiment calculates the position of the assumed sound source in the preset coordinate system after adjusting its pose, thus realizing a virtual change in the sound source pose.
[0073] It can record the virtual pose of the sound source in each pose acquisition cycle, that is, the virtual position and virtual attitude angle after setting the change in pose of the sound source.
[0074] If the reference pose is the user's initial pose, the pose change is adjusted based on the sound source's initial pose. If the reference pose is the user's historical pose, the pose change is adjusted based on the sound source's virtual pose.
[0075] The preset coordinate system can be a pre-set global coordinate system, such as the world coordinate system.
[0076] Step A40: Input the virtual position into the preset rendering system to obtain the surround sound signal.
[0077] The preset rendering system is capable of rendering surround sound. In a preferred embodiment, the preset rendering system is an Ambisonic rendering system, a technology for three-dimensional spatial sound field reconstruction that can capture and reproduce sounds from different directions, providing users with an immersive auditory experience. It can be used to create more realistic and engaging audio environments, enhancing the user's sense of immersion.
[0078] Virtual position represents the dynamic position of the sound source relative to the user, taking into account the movement and rotation of the user's head. During signal rendering, it can be used to calculate the time difference and intensity difference of sound waves reaching the listener's two ears, and to adjust the direction of the sound source according to the user's head rotation, ensuring that the sense of sound direction is consistent with the user's visual and head orientation, thus enhancing the realism of surround sound. In this way, no matter how the listener moves or turns their head, they can hear sound that accurately reflects the location and distance of the sound source, thereby obtaining a natural and immersive auditory experience.
[0079] In one possible implementation, the pose change includes a position change and an attitude angle change, and the step of determining the virtual position of the sound source in the preset sound source array based on the pose change includes:
[0080] Step B10: Determine the first position in the preset coordinate system after the position change of the sound source in the preset sound source array is moved by the amount of the position change, and / or determine the second position in the preset coordinate system after the attitude angle change of the sound source in the preset sound source array is rotated by the amount of the attitude angle change.
[0081] For example, taking the historical pose as the reference pose, the change in position can be the difference between the user's current position and the historical position. Let the user's current position be head_position_now(x,y,z), the user's historical position be head_position_former(x,y,z), and the change in position be Δ(x,y,z). Then, the change in position can be expressed by the formula Δ(x,y,z) = head_position_now(x,y,z) - head_position_former(x,y,z). Let the historical sound source position be denoted as `source_position_former(x,y,z)`, and the first position be denoted as `source_position_now(x,y,z)`. The first position can be expressed by the formula: `source_position_now(x,y,z) = source_position_former(x,y,z) + Δ(x,y,z) = source_position_former(x,y,z) + head_position_now(x,y,z) - head_position_former(x,y,z)`. After calculating the first position, the historical sound source position is updated to this first position so that the historical sound source position can be successfully obtained in the subsequent pose acquisition.
[0082] It should be noted that the historical sound source position is the virtual position of the sound source in the previous pose acquisition cycle, that is, the virtual position of the sound source in the previous pose acquisition cycle as the sound source "moves" with the user. For example, assuming the initial position of the sound source is (0, 0, 0), after the first pose acquisition cycle, the sound source position is updated to (1, 1, 1) as the user "moves". Then, in the second pose acquisition cycle, the historical sound source position is the virtual position (1, 1, 1). In the second pose acquisition cycle, the sound source "moves" from (1, 1, 1) to (2, 2, 2) as the user. Then, in the third pose acquisition cycle, the historical sound source position is the virtual position (2, 2, 2). It can be understood that the actual position of the sound source is still (0, 0, 0). The actual position has not changed. What is updated is the virtual sound field position rendered by the sound source, or more specifically, the optimal sound effect position in the sound field rendered by the sound source, or the position where the audio listening effect level reaches or exceeds the preset sound effect level.
[0083] If the current pose acquisition cycle is the first pose acquisition cycle, then the historical position of the sound source can be determined as the initial position of the sound source.
[0084] Step B20: Determine the virtual position of the sound source in the preset sound source array based on the first position and / or the second position.
[0085] It should be noted that if only the first position is calculated, then the first position is determined to be a virtual position; if only the second position is calculated, then the second position is determined to be a virtual position; if both the first and second positions are calculated, then the sum of the first and second positions is determined to be a virtual position.
[0086] In one possible implementation, the attitude angle change includes pitch angle change, roll angle change, and yaw angle change. The step of determining the second position of the sound sources in the preset sound source array in the preset coordinate system after rotating by the attitude angle change includes:
[0087] Step C10: Determine the second position in the preset coordinate system after the sound source in the preset sound source array rotates around the horizontal axis of the preset coordinate system by the pitch angle change, around the vertical axis of the preset coordinate system by the roll angle change, and around the depth axis of the preset coordinate system by the yaw angle change.
[0088] It should be noted that the horizontal axis can specifically be the x-axis of a preset coordinate system, the vertical axis can specifically be the axis of a preset coordinate system, and the depth axis can specifically be the z-axis of a preset coordinate system. The pitch angle change is P1-P0, the roll angle change is R1-R0, and the yaw angle change is Y1-Y0.
[0089] For example, if the reference pose is a historical pose, the virtual pose of the sound source is rotated around the horizontal axis, vertical axis and depth axis, and the attitude angle of the sound source after rotating on the three axes can be recorded and recorded as the virtual attitude angle of the sound source.
[0090] In this embodiment, the rotation of the sound source is decomposed into rotation in three coordinate axes, which allows the sound source to achieve more precise "rotation" as the user rotates, thereby further improving the surround sound effect.
[0091] In this embodiment, the user's position change and attitude angle change are obtained to determine the first position after the sound source's position change and / or the second position after the sound source's attitude angle change. It can be understood that the position change is caused by the user's position change and the attitude angle change is caused by the user's rotation. The surround sound signal is rendered based on the virtual position after the sound source moves and / or rotates. This method of rendering the surround sound signal based on the virtual position achieves the effect that the sound source "moves" with the user's movement and / or "rotates" with the user's rotation (the actual sound source does not move and / or rotate, but the signal is re-rendered at the position after the movement and / or rotation, giving the listener the feeling that the sound source is "moving" and / or "rotating" with them). It is no longer rendering surround sound with a fixed distribution position of the sound source, thus ensuring that even when the user's position changes due to movement and / or rotation, the rendered surround sound signal also moves and rotates with the user's movement, thereby making the rendered surround sound signal have a better sense of space and direction, and a better surround sound effect.
[0092] Based on the first and / or second embodiments of this application, in the third embodiment of this application, the content that is the same as or similar to that in embodiments one and two above can be referred to the above description, and will not be repeated hereafter. On this basis, refer to Figure 3 As shown, the step of determining the second position in the preset coordinate system after the change in the attitude angle of the sound source in the preset sound source array includes:
[0093] Step D10: Obtain the initial angle information of the sound sources in the preset sound source array, wherein the initial angle information includes the azimuth angle and the elevation angle;
[0094] Specifically, this initial angle information can be the angle information of the sound sources set when arranging the sound source array, which can be represented as (azimuth, elevation), where azimuth represents the azimuth angle and elevation represents the elevation angle. It can be understood that the initial angle information of the sound sources is set in advance; for example, the initial angle information of the first sound source in a 5.1 surround sound system is (30, 0), where 30 is the azimuth angle and 0 is the elevation angle.
[0095] Step D20: Convert the initial angle information into a quaternion to obtain a first position vector, and convert the current pose into a quaternion to obtain a second position vector;
[0096] In three-dimensional space, using Euler angles to represent rotation may encounter gimbal lock, while quaternions can avoid this problem. Based on this, in this embodiment, the initial angle information is converted into quaternions to obtain the first position vector, and the current pose is converted into quaternions to obtain the second position vector.
[0097] Furthermore, when the position of the object (given by x, y, z coordinates) and attitude angle are known, they can be converted into quaternions based on known methods. This embodiment will not describe in detail the specific process of converting the current position and attitude angle into quaternions.
[0098] Step D30: Calculate the product between the first position vector and the second position vector to obtain the target position vector;
[0099] It should be noted that the product between the first position vector and the second position vector is specifically a quaternion multiplication. Let the first position vector be Pq, the second position vector be q, and the target position vector be P'q. Then the target position vector can be expressed by the formula P'q = q * Pq * q - 1, where q - 1 represents the inverse of q, which is equal to the conjugate of q.
[0100] Step D40: Extract the imaginary part of the target position vector, and determine the extracted imaginary part as the second position of the sound source in the preset coordinate system after the change in the attitude angle.
[0101] Let P' be the imaginary part of the target position vector. Then P' is the second position. It can be understood that a quaternion has three imaginary parts, that is, P' can be represented as (x', y', z'), where (x', y', z') is the second position, representing the coordinates of the sound source on each axis in the world coordinate system.
[0102] In one possible implementation, the step of converting the initial angle information into a quaternion to obtain a first position vector includes:
[0103] Step E10: Convert the initial angle information into a preset coordinate system to obtain the initial position information;
[0104] Step E20: Convert the initial position information into a quaternion to obtain a first position vector, wherein the real part of the first position vector is zero, and the imaginary part of the first position vector is the initial position information.
[0105] The initial angle information is transformed into the preset coordinate system to obtain the initial position information. Specifically, the transformation process can be: x0 = -cos(elevation)*sin(azimuth), y0 = sin(elevation), z0 = -cos(elevation)*cos(azimuth), that is, the initial position information can be represented as (x0, y0, z0).
[0106] The first position vector can be represented as Pq = 0 + x0i + y0i + z0k.
[0107] In this embodiment, the initial angle information of the sound source is converted into a quaternion first position vector. Quaternions require only four values to represent a rotation, and quaternion operations are generally more efficient than matrix operations, especially when performing compound rotations, where quaternion multiplication requires less computation than matrix multiplication. Furthermore, Euler angle representation of rotation may encounter gimbal lock issues, which occur when two rotation axes are aligned, resulting in the loss of a degree of freedom and preventing normal rotation. Quaternions do not have this problem and are unaffected by rotation axis alignment. Therefore, converting the angle information into quaternions and subsequently calculating the sound source's "rotated" position in the world coordinate system using quaternion products improves the accuracy of the calculated second position, further enhancing the surround sound effect.
[0108] In one possible implementation, the steps of determining the first position in the preset coordinate system after the sound source in the preset sound source array moves by the amount of position change, and / or determining the second position in the preset coordinate system after the sound source in the preset sound source array rotates by the amount of attitude angle change, include:
[0109] Step F10: If the position change is not zero and the attitude angle change is zero, then determine the first position of the sound source after the position change.
[0110] It should be noted that if the reference pose is the user's historical pose, then steps F10 to F30 are executed.
[0111] If the position change is not zero and the attitude angle change is zero, it means that the user only moves and does not rotate. At this time, the first position after the change in the position of the sound source is determined, that is, the sound source "moves" accordingly, and the second position does not need to be calculated, that is, the sound source does not "rotate", thus avoiding invalid or repeated calculation of the second position.
[0112] Step F20: If the position change is zero and the attitude angle change is not zero, then determine the second position of the sound source after the attitude angle change.
[0113] If the change in position is zero and the change in attitude angle is zero, it means that the user only rotates and does not move. At this time, the second position after the change in attitude angle of the sound source is determined, that is, the sound source "rotates" accordingly, and the first position does not need to be calculated, that is, the sound source does not "move", thus avoiding invalid or repeated calculation of the first position.
[0114] Step F30: If the position change is not zero and the attitude angle change is not zero, then determine the first position of the sound source after the position change and the second position of the sound source after the attitude angle change.
[0115] If the change in position is not zero and the change in attitude angle is not zero, it means that the user is both moving and rotating. At this time, the first position after the change in the position of the sound source and the second position after the change in the attitude angle of the sound source are determined. That is, the sound source "moves" and "rotates" accordingly, so as to further improve the surround sound effect of the surround sound signal obtained after subsequent rendering.
[0116] If both the position change and the attitude angle change are zero, it means that the user has neither moved nor rotated. In this case, there is no need to recalculate the first and second positions, and the default rendering system can simply maintain the current rendering method to render the signal.
[0117] Furthermore, embodiments of this application also propose a surround sound generation device, referring to... Figure 4 As shown, the surround sound generating device includes:
[0118] The acquisition module 10 is used to acquire the current pose of the user when a change in the user's pose is detected, wherein the current pose includes the current position and the current posture angle;
[0119] The generation module 20 is used to regenerate a surround sound signal based on the current pose and output the regenerated surround sound signal through a preset sound source array so that the current pose is located in the sweet spot of the sound field formed by the regenerated surround sound signal.
[0120] In one embodiment, the generation module 20 is further configured to:
[0121] Obtain the user's reference pose, wherein the reference pose includes a reference position and a reference attitude angle;
[0122] Calculate the pose change between the reference pose and the current pose;
[0123] The virtual position of the sound source in the preset sound source array is determined based on the pose change amount, wherein the virtual position is the position of the sound source in the preset coordinate system after the pose change amount is adjusted.
[0124] The virtual location is input into a preset rendering system to obtain a surround sound signal.
[0125] In one embodiment, the pose change includes a position change and an attitude angle change, and the generation module 20 is further configured to:
[0126] Determine the first position in the preset coordinate system after the position change of the sound source in the preset sound source array is moved by the amount of the position change, and / or determine the second position in the preset coordinate system after the attitude angle change of the sound source in the preset sound source array is rotated by the amount of the attitude angle change;
[0127] The virtual positions of the sound sources in the preset sound source array are determined based on the first position and / or the second position.
[0128] In one embodiment, the generation module 20 is further configured to:
[0129] After determining the pitch angle change, roll angle change, and yaw angle change of the sound sources in the preset sound source array around the horizontal axis of the preset coordinate system, the second position in the preset coordinate system is determined.
[0130] In one embodiment, the generation module 20 is further configured to:
[0131] Obtain the initial angle information of the sound sources in the preset sound source array, wherein the initial angle information includes the azimuth angle and the elevation angle;
[0132] The initial angle information is converted into a quaternion to obtain a first position vector, and the current pose is converted into a quaternion to obtain a second position vector;
[0133] The target position vector is obtained by multiplying the first position vector and the second position vector.
[0134] Extract the imaginary part of the target position vector, and determine the extracted imaginary part as the second position of the sound source in the preset coordinate system after the change in the attitude angle.
[0135] In one embodiment, the generation module 20 is further configured to:
[0136] The initial angle information is converted to a preset coordinate system to obtain the initial position information;
[0137] The initial position information is converted into a quaternion to obtain a first position vector, wherein the real part of the first position vector is zero, and the imaginary part of the first position vector is the initial position information.
[0138] In one embodiment, the generation module 20 is further configured to:
[0139] If the change in position is not zero and the change in attitude angle is zero, then determine the first position of the sound source after the change in position.
[0140] If the position change is zero and the attitude angle change is not zero, then determine the second position of the sound source after rotating by the attitude angle change.
[0141] If the change in position is not zero and the change in attitude angle is not zero, then determine the first position of the sound source after the change in position and the second position of the sound source after the change in attitude angle.
[0142] Furthermore, this application also proposes a surround sound generation device, which includes: a memory, a processor, and a computer program stored in the memory and executable on the processor, the computer program being configured to implement the steps of the surround sound generation method described above.
[0143] In addition, refer to Figure 5 The diagram illustrates a structural schematic suitable for implementing the surround sound generation device of the embodiments of this application. The surround sound generation device in the embodiments of this application may also include, but is not limited to, mobile terminals such as headphones, AR glasses, VR glasses, mobile phones, laptops, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Portable Application Description), PMPs (Portable Media Players), etc., as well as fixed terminals such as digital TVs, desktop computers, etc. Figure 5 The surround sound generation device shown is merely an example and should not impose any limitation on the functionality and scope of use of the embodiments of this application.
[0144] like Figure 5 As shown, the surround sound generation device may include a processing unit 1001 (e.g., a central processing unit, a graphics processing unit, etc.), which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 1002 or a program loaded from a storage device 1003 into a random access memory (RAM) 1004. The RAM 1004 also stores various programs and data required for the operation of the surround sound generation device. The processing unit 1001, ROM 1002, and RAM 1004 are interconnected via a bus 1005. An input / output (I / O) interface 1006 is also connected to the bus. Typically, the following systems can be connected to the I / O interface 1006: input devices 1007 including, for example, touchscreens, touchpads, keyboards, mice, image sensors, microphones, accelerometers, gyroscopes, etc.; output devices 1008 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices 1003 including, for example, magnetic tapes, hard disks, etc.; and communication devices 1009. Communication device 1009 allows the surround sound generating device to communicate wirelessly or wiredly with other devices to exchange data. Although the figures show surround sound generating devices with various systems, it should be understood that implementation or possession of all the systems shown is not required. More or fewer systems may be implemented alternatively.
[0145] Specifically, according to the embodiments disclosed in this application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments disclosed in this application include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device, or installed from storage device 1003, or installed from ROM 1002. When the computer program is executed by processing device 1001, it performs the functions defined in the methods of the embodiments disclosed in this application.
[0146] The surround sound generation device provided in this application, employing the surround sound generation method described in the above embodiments, can solve the technical problem of ensuring the spatial and directional sense of surround sound when the listener's position changes, thereby guaranteeing surround sound effects. Compared with the prior art, the beneficial effects of the surround sound generation device provided in this application are the same as those of the surround sound generation method provided in the above embodiments, and other technical features of this surround sound generation device are the same as those disclosed in the previous embodiment method, and will not be repeated here.
[0147] It should be understood that the various parts disclosed in this application can be implemented using hardware, software, firmware, or a combination thereof. In the description of the above embodiments, specific features, structures, materials, or characteristics can be combined in any suitable manner in one or more embodiments or examples.
[0148] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
[0149] In addition, to achieve the above objectives, this application also provides a readable storage medium having computer-readable program instructions (i.e., a computer program) stored thereon, which are used to execute the surround sound generation method in the above embodiments.
[0150] The computer-readable storage medium provided in this application embodiment may be, for example, a USB flash drive, but is not limited to, electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices, or any combination thereof. More specific examples of computer-readable storage media may include, but are not limited to: electrical connections with one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof. In this embodiment, the computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, system, or device. The program code contained on the computer-readable storage medium may be transmitted using any suitable medium, including but not limited to: wires, optical cables, RF (Radio Frequency), etc., or any suitable combination thereof.
[0151] The aforementioned computer-readable storage medium may be included in the surround sound generating device; or it may exist independently and not assembled into the surround sound generating device.
[0152] The aforementioned computer-readable storage medium carries one or more programs that, when executed by the surround sound generation device, cause the surround sound generation device to: upon detecting a change in the user's pose, acquire the user's current pose, wherein the current pose includes the current position and the current attitude angle; regenerate a surround sound signal based on the current pose, and output the regenerated surround sound signal through a preset sound source array, so that the current pose is located in the sweet spot of the sound field formed by the regenerated surround sound signal.
[0153] Computer program code for performing the operations of this application can be written in one or more programming languages or a combination thereof, including object-oriented programming languages such as Java, Smalltalk, and C++, and conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a Local Area Network (LAN) or a Wide Area Network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0154] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this application. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.
[0155] The modules described in the embodiments of this application can be implemented in software or hardware. The names of the modules do not necessarily limit the functionality of the unit itself.
[0156] The readable storage medium provided in this application is a computer-readable storage medium that stores computer-readable program instructions (i.e., a computer program) for executing the above-described surround sound generation method. This solves the technical problem of ensuring the spatial and directional sense of surround sound when the listener's position changes, thereby guaranteeing surround sound effects. Compared with the prior art, the beneficial effects of the computer-readable storage medium provided in this application are the same as those of the surround sound generation method provided in the above embodiments, and will not be repeated here.
[0157] Furthermore, embodiments of this application also propose a computer program product, including a surround sound generation program, which, when executed by a processor, implements the steps of the surround sound generation method as described above.
[0158] The specific implementation of the computer program product in this application is basically the same as the embodiments of the above-described surround sound generation method, and will not be repeated here.
[0159] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or system that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or system. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or system that includes that element.
[0160] The sequence numbers of the embodiments in this application are for descriptive purposes only and do not represent the superiority or inferiority of the embodiments.
[0161] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software sensor. This computer software sensor is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) as described above, and includes several instructions to cause a terminal device (which may be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in the various embodiments of this application.
[0162] The above are merely preferred embodiments of this application and do not limit the patent scope of this application. Any equivalent structural or procedural transformations made using the content of this application's specification and drawings, or direct or indirect applications in other related technical fields, are similarly included within the patent protection scope of this application.
Claims
1. A method for generating surround sound, characterized in that, The surround sound generation method includes the following steps: When a change in the user's pose is detected, the user's current pose is obtained, wherein the current pose includes the current position and the current posture angle; The current pose is used to regenerate a surround sound signal, and the regenerated surround sound signal is output through a preset sound source array so that the current pose is located in the sweet spot of the sound field formed by the regenerated surround sound signal.
2. The surround sound generation method as described in claim 1, characterized in that, The step of regenerating the surround sound signal based on the current pose includes: Obtain the user's reference pose, wherein the reference pose includes a reference position and a reference attitude angle; Calculate the pose change between the reference pose and the current pose; The virtual position of the sound source in the preset sound source array is determined based on the pose change amount, wherein the virtual position is the position of the sound source in the preset coordinate system after the pose change amount is adjusted. The virtual location is input into a preset rendering system to obtain a surround sound signal.
3. The surround sound generation method as described in claim 2, characterized in that, The pose change includes position change and attitude angle change. The step of determining the virtual position of the sound source in the preset sound source array based on the pose change includes: Determine the first position in the preset coordinate system after the position change of the sound source in the preset sound source array is moved by the amount of the position change, and / or determine the second position in the preset coordinate system after the attitude angle change of the sound source in the preset sound source array is rotated by the amount of the attitude angle change; The virtual positions of the sound sources in the preset sound source array are determined based on the first position and / or the second position.
4. The surround sound generation method as described in claim 3, characterized in that, The attitude angle change includes pitch angle change, roll angle change, and yaw angle change. The step of determining the second position of the sound sources in the preset sound source array in the preset coordinate system after rotating by the attitude angle change includes: After determining the pitch angle change, roll angle change, and yaw angle change of the sound sources in the preset sound source array around the horizontal axis of the preset coordinate system, the second position in the preset coordinate system is determined.
5. The surround sound generation method as described in claim 3, characterized in that, The step of determining the second position in the preset coordinate system after the change in the attitude angle of the sound source in the preset sound source array is rotated includes: Obtain the initial angle information of the sound sources in the preset sound source array, wherein the initial angle information includes the azimuth angle and the elevation angle; The initial angle information is converted into a quaternion to obtain a first position vector, and the current pose is converted into a quaternion to obtain a second position vector; The target position vector is obtained by multiplying the first position vector and the second position vector. Extract the imaginary part of the target position vector, and determine the extracted imaginary part as the second position of the sound source in the preset coordinate system after the change in the attitude angle.
6. The surround sound generation method as described in claim 5, characterized in that, The step of converting the initial angle information into a quaternion to obtain the first position vector includes: The initial angle information is converted to a preset coordinate system to obtain the initial position information; The initial position information is converted into a quaternion to obtain a first position vector, wherein the real part of the first position vector is zero, and the imaginary part of the first position vector is the initial position information.
7. The surround sound generation method as described in claim 3, characterized in that, The steps of determining the first position in the preset coordinate system after the position change of the sound source in the preset sound source array, and / or determining the second position in the preset coordinate system after the attitude angle change of the sound source in the preset sound source array, include: If the change in position is not zero and the change in attitude angle is zero, then determine the first position of the sound source after the change in position. If the position change is zero and the attitude angle change is not zero, then determine the second position of the sound source after rotating by the attitude angle change. If the change in position is not zero and the change in attitude angle is not zero, then determine the first position of the sound source after the change in position and the second position of the sound source after the change in attitude angle.
8. A surround sound generating device, characterized in that, The surround sound generation device includes: a memory, a processor, and a computer program stored in the memory and executable on the processor, the computer program being configured to implement the steps of the surround sound generation method as described in any one of claims 1 to 7.
9. A readable storage medium, characterized in that, The readable storage medium is a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, it implements the steps of the surround sound generation method as described in any one of claims 1 to 7.
10. A computer program product, characterized in that, The computer program product includes a computer program that, when executed by a processor, implements the steps of the surround sound generation method as described in any one of claims 1 to 7.