Audio playing method and related apparatus
By acquiring the spatial trajectory of audio elements and controlling the movement of speakers, the problem of insufficient spatial adaptability in traditional audio playback solutions is solved, achieving an immersive panoramic sound experience.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- IFLYTEK (SUZHOU) TECH CO LTD
- Filing Date
- 2026-03-28
- Publication Date
- 2026-06-16
AI Technical Summary
Traditional audio playback solutions fail to meet users' needs for an immersive listening experience and neglect the compatibility between space and audio system.
By acquiring the sound image space trajectory of audio elements in the audio to be played, the target motion trajectory of the speaker is determined, and the speaker is controlled to move according to the trajectory to restore the sound image space trajectory of the audio elements, thereby achieving accurate restoration of the audio elements in space.
It achieves the consistency between the sound image movement trajectory of the audio elements perceived by the human ear in the target playback space and the sound image spatial trajectory of the audio elements in the audio to be played, creating an immersive panoramic sound experience.
Smart Images

Figure CN122227174A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of audio technology, and in particular to an audio playback method and related apparatus. Background Technology
[0002] In recent years, the audio industry has been rapidly iterating and upgrading. With the optimization of audio systems, users' listening needs have continued to upgrade. They are no longer satisfied with basic sound quality performance, but instead have higher requirements for spatial listening immersion, and they yearn to obtain an immersive listening experience that is enveloped by the sound field and feels like they are there in various scenarios. Current traditional listening solutions focus solely on optimizing device sound quality, neglecting the compatibility with the space and audio system, thus failing to meet users' needs for immersive listening. Therefore, enabling users to achieve an immersive listening experience is of paramount importance. Summary of the Invention
[0003] The main technical problem addressed by this application is to provide an audio playback method and related apparatus that enable users to have an immersive listening experience.
[0004] To solve the above-mentioned technical problems, one technical solution adopted in this application is: to provide an audio playback method, the method comprising: acquiring the sound image spatial trajectory of audio elements in the audio to be played; using the sound image spatial trajectory to determine the target motion trajectory of a speaker in a target playback space; controlling the speaker to move according to the target motion trajectory; and controlling the speaker to play the audio to be played while moving; wherein, the sound image motion trajectory of the audio to be played played when the speaker moves according to the target motion trajectory matches the sound image spatial trajectory of the audio elements.
[0005] To solve the aforementioned technical problems, another technical solution adopted in this application is to provide an audio playback device, which includes an acquisition module, a determination module, and a control module. The acquisition module is used to acquire the sound image spatial trajectory of audio elements in the audio to be played; the determination module is used to determine the target motion trajectory of a speaker in the target playback space using the sound image spatial trajectory; the control module is used to control the speaker to move according to the target motion trajectory, and to control the speaker to play the audio while moving; wherein, the sound image motion trajectory of the audio to be played played when the speaker moves according to the target motion trajectory matches the sound image spatial trajectory of the audio elements.
[0006] To solve the above-mentioned technical problems, another technical solution adopted in this application is: to provide an electronic device, including a memory and a processor coupled to each other, wherein the memory stores program instructions; and the processor is used to execute the program instructions stored in the memory to implement the above-mentioned method.
[0007] To solve the above-mentioned technical problems, another technical solution adopted in this application is: to provide an audio playback system, including a speaker rotation assembly and the above-mentioned electronic device; the speaker rotation assembly is communicatively connected to the electronic device, and the speaker rotation assembly is used to control the movement of the speaker.
[0008] To solve the above-mentioned technical problems, another technical solution adopted in this application is to provide a computer-readable storage medium for storing program instructions that can be executed to implement the above-mentioned method.
[0009] The above solution can control the speaker to move along the target motion trajectory, and the sound image motion trajectory of the audio to be played when the speaker moves along the target motion trajectory matches the sound image spatial trajectory of the audio element. Therefore, this method can make the sound image motion trajectory of the audio element perceived by the human ear in the target playback space consistent with the sound image spatial trajectory of the audio element in the audio to be played, thereby achieving accurate restoration of the sound image spatial trajectory of the audio element in the audio to be played, and giving the user an immersive panoramic sound experience. Attached Figure Description
[0010] Figure 1 This is a flowchart illustrating an embodiment of the audio playback method provided in this application; Figure 2 This is a schematic diagram of the audio element track division provided in this application; Figure 3 This is a schematic diagram of the sound image subspace trajectory corresponding to the audio element provided in this application; Figure 4 yes Figure 1 The flowchart of step S12 shown is a schematic diagram of one embodiment; Figure 5 This is a schematic diagram of an embodiment of the target motion trajectory of the loudspeaker provided in this application; Figure 6 This is a flowchart illustrating an embodiment of determining the target pose of a loudspeaker at various times provided in this application; Figure 7 This is a schematic diagram of an embodiment of the in-vehicle sound field center provided in this application; Figure 8 This is a schematic diagram of the framework of an embodiment of the audio playback device provided in this application; Figure 9 This is a schematic diagram of the framework of an embodiment of the electronic device provided in this application; Figure 10 This is a schematic diagram of the framework of the audio playback system provided in this application; Figure 11 This is a schematic diagram of the framework of the computer-readable storage medium provided in this application. Detailed Implementation
[0011] To make the purpose, technical solution and effects of this application clearer and more explicit, the following describes this application in further detail with reference to the accompanying drawings and embodiments.
[0012] Furthermore, if the embodiments of this application involve descriptions such as "first" or "second," these descriptions are for descriptive purposes only and should not be construed as indicating or implying their relative importance or implicitly specifying the number of technical features indicated. Therefore, features defined with "first" or "second" may explicitly or implicitly include at least one of those features. Additionally, the technical solutions of various embodiments can be combined with each other, but this must be based on the ability of those skilled in the art to implement them. If the combination of technical solutions is contradictory or impossible to implement, it should be considered that such a combination of technical solutions does not exist and is not within the scope of protection claimed in this application.
[0013] The audio playback method provided in this application is used to determine the target motion trajectory of the speaker when playing the audio by utilizing the sound image spatial trajectory of the audio elements in the audio to be played, and to control the speaker to play the audio when moving according to the target motion trajectory, so as to restore the sound image spatial trajectory of the audio elements in the audio to be played through the speaker movement, thereby creating an immersive panoramic sound listening experience.
[0014] It should be noted that the sound emitted by a loudspeaker has a distinct directional characteristic, and the direction of its sound emission directly affects the human ear's perception of the sound source's location. The human ear can perceive the spatial orientation of the sound source based on the direction or angle from which the sound is transmitted. This application, by controlling the movement of the loudspeaker, can actively change its sound emission direction, allowing the sound to radiate to the listener from different angles and directions, giving the listener the auditory sensation of the sound source's location moving. Therefore, by controlling the loudspeaker to move along a predetermined target trajectory, this application can continuously change the sound emission direction, simulating the movement / flow of a sound source from one location to another in space, thereby reproducing the location and dynamic changes of sound in space.
[0015] The audio to be played can be, but is not limited to, pre-stored audio (such as a song), or it can be the user's singing voice and background music collected in a karaoke setting. The audio playback method provided in this application can play audio in any preset target playback space, which can be, but is not limited to, the interior of a vehicle, a karaoke room, a movie theater, a bedroom, etc. Taking a scenario where the target playback space is the interior of a vehicle as an example, the user can select to play the corresponding audio file via Bluetooth, online, or other means. The vehicle's infotainment system inputs the audio signal of the audio to be played to the audio playback device or system, enabling the audio playback device or system to play the audio using the method provided in this application.
[0016] An audio element refers to an independently identifiable audio unit with independent acoustic properties within the audio to be played. In simple terms, vocals, guitar sounds, drum sounds, bass sounds, ambient sound effects, and special effects sounds in a song or a movie sound effect are all independent audio elements; in multi-channel and immersive audio, the various tracks and independent sound source signals are also classified as audio elements.
[0017] Sound image is a virtual sound source or sensory sound source. It can be simply understood as the virtual sound source location corresponding to the audio element perceived by the human ear within the listening space. It belongs to the virtual sound source localization result at the auditory level, and is not the actual physical placement of the speaker. The spatial trajectory of a sound image is the complete path of the virtual sound source corresponding to an audio element as it moves and changes in space over time. For example, the sound effect of a car speeding by in a movie will have its virtual sound source (sound image) move from the left rear, through the middle, to the right front, forming a continuous trajectory. The audio playback method provided in this application is mainly used to restore the movement path of the virtual sound source corresponding to the audio element in space over time by controlling the movement of the speaker when playing audio. For example, it can restore the sound effect of a car moving from the left rear, through the middle, to the right front, creating a panoramic sound experience for users.
[0018] Please see Figure 1 , Figure 1 This is a schematic flowchart of an embodiment of the audio playback method provided in this application. It should be noted that if substantially the same result is achieved, this embodiment does not necessarily replace it. Figure 1 The illustrated process sequence is limited. For example... Figure 1 As shown, this embodiment includes: S11: Obtain the sound image space trajectory of the audio elements in the audio to be played.
[0019] This embodiment is used to determine the target motion trajectory of the speaker by utilizing the sound image space trajectory of the audio elements in the audio to be played, and then control the speaker to move according to the target motion trajectory and control the speaker to play the audio while moving, so that when the speaker plays the audio, it restores the sound image space trajectory of the audio elements, that is, restores the movement path of the virtual sound source corresponding to the audio elements in space over time, creating an immersive panoramic sound listening experience for the listener.
[0020] The audio to be played includes sub-audio elements of at least one audio element, and the sound image space trajectory includes the sound image sub-space trajectory of each audio element. For example, a song includes corresponding sub-audio elements such as vocals, guitar sounds, and drum sounds. The sound image space trajectory includes: the movement trajectory of the vocals in space over time (i.e., the sound image sub-space trajectory of the vocals), the movement trajectory of the guitar sounds in space over time, and the movement trajectory of the drum sounds in space over time, etc.
[0021] In one embodiment, the sub-audio corresponding to each audio element can be separated from the audio to be played first; then, the sound image subspace trajectory of the corresponding audio element can be parsed from each sub-audio.
[0022] This process utilizes relevant track-segmentation models or algorithms (such as AI track-segmentation algorithms) to separate the audio elements in the audio to be played into several tracks representing different audio elements, such as piano, vocals, bass, drums, etc. These tracks are then arranged according to the playback time order of each audio element and converted into independent audio materials (sub-audio). For an example, please refer to [link to example]. Figure 2 , Figure 2 This is a schematic diagram of the audio element track division provided in this application, such as... Figure 2 As shown, the sub-audio elements (voice, drums, violin, guitar, etc.) corresponding to each audio element are separated from the audio to be played.
[0023] Furthermore, for each separated sub-audio element, the spatial metadata of the corresponding audio element can be parsed, and then the spatial metadata of the audio element can be projected onto a coordinate system pre-constructed in the target playback space, such as projecting it onto a spatial coordinate system pre-constructed inside the vehicle; using the spatial metadata of the audio element, the sound image position P of the audio element in the spatial coordinate system at each time point can be determined. Δt (x Δt ,y Δt , z Δt This is used to obtain the sound image spatial trajectory of the audio element as a function of time t. First, a trajectory function f(t) can be fitted to the sound image position of the audio element at each time point. Then, this trajectory function f(t) is used to characterize the sound image subspace trajectory of the audio element. For an example, please refer to [link to example]. Figure 3 , Figure 3 This is a schematic diagram of the sound image subspace trajectory corresponding to the audio element provided in this application, such as... Figure 3 The image subspace trajectory f1(t) of the piano and the image subspace trajectory f2(t) of the human voice are shown.
[0024] The spatial metadata of the audio to be played consists of attribute data describing the spatial location, sound image range, and motion trajectory of each audio element. During the mixing stage, spatial positioning information is written into the audio file and encapsulated, allowing it to be directly parsed and mapped to a spatial coordinate system constructed in the target playback space. This facilitates accurate calculation of the speaker's target motion trajectory, accurately reconstructing the sound image spatial trajectory of the audio to be played, and achieving precise sound image positioning and speaker matching.
[0025] Of course, in some embodiments, other methods can also be used to separate the sub-audio corresponding to each audio element and parse the sound image subspace trajectory of the corresponding audio element.
[0026] S12: Using the sound image space trajectory, determine the target motion trajectory of the speaker in the target playback space.
[0027] The audio to be played includes at least one sub-audio element, and the sound image space trajectory of the audio to be played includes the sound image sub-space trajectory of each audio element.
[0028] In some embodiments, the number of speakers disposed in the target playback space is multiple. In scenarios where the audio to be played contains at least two audio elements, in order to accurately reconstruct the sound image sub-space trajectory of each audio element in the target playback space, a corresponding speaker can be determined for each audio element. The movement of each speaker is only used to reconstruct the sound image sub-space trajectory of its corresponding audio element. For example, the speaker corresponding to the piano is speaker A, and the speaker corresponding to the human voice is speaker B. The movement of speaker A is used to reconstruct the movement trajectory of the piano in space, and the movement of speaker B is used to reconstruct the movement trajectory of the human voice in space.
[0029] Based on this, in one embodiment, the speaker corresponding to each audio element can be determined from the target playback space by first using the sound image subspace trajectory of each audio element; wherein, different audio elements correspond to different speakers; then, the target motion trajectory of the speaker corresponding to each audio element can be determined by using the sound image subspace trajectory of each audio element.
[0030] For example, the target motion trajectory of speaker A corresponding to the piano is determined using the sound image subspace trajectory of the piano, and the target motion trajectory of speaker B corresponding to the human voice is determined using the sound image subspace trajectory of the human voice. This allows for subsequent control of speaker A to move according to its target motion trajectory to reconstruct the sound image subspace trajectory of the piano (i.e., reconstruct the piano's motion trajectory in the target playback space), and control of speaker B to move according to its target motion trajectory to reconstruct the sound image subspace trajectory of the human voice (i.e., reconstruct the human voice's motion trajectory in the target playback space). For details on how to determine the target motion trajectory of the speakers, please refer to the following text. Figure 4 The relevant description of the corresponding embodiments.
[0031] Understandably, since different audio elements correspond to different speakers, the movement of the speakers for each audio element is unaffected by other audio elements. This method is advantageous for utilizing the speakers corresponding to different audio elements to reconstruct the sound image subspace trajectory of each audio element, creating a panoramic sound experience for the listener. For an example, please refer to [link to example]. Figure 3C represents the position of the sound field center (which can be understood as the optimal listening position). The sound image subspace trajectory f1(t) of the piano can be reproduced by controlling the movement of speaker A, and the sound image subspace trajectory f2(t) of the human voice can be reproduced by controlling the movement of speaker B. The movement of speaker A is not affected by the human voice, and speaker B is not affected by the piano. Therefore, it can create the sound effect of each audio element flowing in the target playback space according to the corresponding sound image subspace.
[0032] S13: Control the speaker to move according to the target motion trajectory, and control the speaker to play the audio to be played while moving; wherein, when the speaker moves according to the target motion trajectory, the sound image motion trajectory of the audio to be played matches the sound image spatial trajectory of the audio element.
[0033] It should be noted that this embodiment does not arbitrarily control the speaker's movement, but rather controls the speaker to move according to a defined target trajectory. When the speaker moves according to the corresponding target trajectory, the sound image trajectory of the audio being played matches the sound image spatial trajectory of the audio element. This method allows the sound image trajectory of the audio element perceived by the human ear in the target playback space to be consistent with the sound image spatial trajectory of the audio element in the audio being played, thereby achieving accurate restoration of the sound image spatial trajectory of the audio element in the audio being played and providing the user with an immersive panoramic sound experience.
[0034] In some embodiments, in scenarios where it is necessary to use different speakers to reconstruct the sound image subspace trajectory of different audio elements, the sub-audio of each audio element can be sent to its corresponding speaker before the speaker is controlled to play the audio to be played in motion; then the speaker corresponding to each audio element is controlled to play the sub-audio of the corresponding audio element in motion.
[0035] For example, audio element a corresponds to speaker A, audio element b corresponds to speaker B, and audio element c corresponds to speaker C. Before playing the audio to be played, sub-audio 1 corresponding to audio element a is sent to speaker A to control speaker A to play sub-audio 1 while in motion. Similarly, sub-audio 2 corresponding to audio element b is sent to speaker B to control speaker B to play sub-audio 2 while in motion, and sub-audio 3 corresponding to audio element c is sent to speaker C to control speaker C to play sub-audio 3 while in motion.
[0036] This method is beneficial for restoring the sound image subspace trajectory of different audio elements by controlling the movement of different speakers, thereby enabling users to simultaneously perceive the movement trajectory of each audio element in the target playback space.
[0037] In one implementation scenario, the speaker is a rotatable speaker with a fixed installation position (e.g., a speaker rotation component connected to the speaker). This means the speaker can only be rotated, not its installation position adjusted. Therefore, only the speaker's pointing direction can be adjusted, not its position in space. Thus, rotating the speaker only makes it point at a certain angle (e.g., 45 degrees), allowing the human ear to perceive the direction of the sound source, but not its specific location within that direction.
[0038] The sound image subspace trajectory of each audio element includes: the sound image position of the corresponding audio element at each moment in the audio to be played (this position refers to the specific position of the audio element in the constructed spatial coordinate system). In order to accurately restore the sound source position, so that the human ear can not only perceive the direction of the sound source, but also the distance between the sound source and the human ear, the playback sound pressure of the speaker can be adjusted according to the relative distance between the sound source position and the speaker installation position, so as to restore the position of the sound source by the sound pressure level of the audio.
[0039] Specifically, controlling the speakers corresponding to each audio element to play the sub-audio of the corresponding audio element while in motion includes the following steps: Step 1: For each audio element, obtain the speaker distance of that audio element at each time, where the speaker distance is the relative distance between the sound image position at the corresponding time and the corresponding speaker installation position.
[0040] That is, first obtain the relative distance between the sound image position of the audio element at each moment and the corresponding speaker installation position. Here, the sound image position of the audio element refers to the sound image position of the audio element in the spatial coordinate system constructed in the target playback space, and the speaker installation position is the position of the speaker in the spatial coordinate system. Therefore, the relative distance between the two can be directly calculated using their positions.
[0041] Step 2: Utilize the speaker distance at each moment to determine the playback sound pressure level of the sub-audio at the corresponding moment; control the corresponding speaker to play the sub-audio of the audio element according to each playback sound pressure level when it moves.
[0042] In one specific implementation, the playback sound pressure level corresponding to an audio element can be determined using the following formula:
[0043] Where d0 is the reference distance (e.g., 1m), L0 is the target sound pressure level marked at d0 during tuning, d represents the relative distance, and L(d) represents the playback sound pressure at a relative distance of d.
[0044] This formula shows that the playback sound pressure is negatively correlated with the relative distance; that is, the greater the distance, the lower the playback sound pressure. In this way, users can perceive the distance of the sound source based on the magnitude of the playback sound pressure they hear.
[0045] Of course, in other implementation scenarios, the speaker can also be a rotatable and movable speaker. In this scenario, the position of the sound source in the spatial coordinate system can be restored directly by moving the speaker.
[0046] Please see Figure 4 , Figure 4 yes Figure 1 The flowchart of one embodiment of step S12 is shown. In this embodiment, step S12 utilizes the sound image spatial trajectory to determine the target motion trajectory of the speaker in the target playback space, and further includes: S41: Using the sound image subspace trajectory of each audio element, determine the speaker corresponding to each audio element from the target playback space.
[0047] This embodiment is used to determine the speaker corresponding to each audio element by utilizing the sound image subspace trajectory of each audio element, and further determine the target motion trajectory of the corresponding speaker, so that the sound image space trajectory of the corresponding audio element can be restored by moving the speaker according to the corresponding target motion trajectory.
[0048] In this embodiment, different audio elements correspond to different speakers. Furthermore, it should be noted that the sound waves from a speaker are directional. For the speaker corresponding to an audio element to be able to reconstruct the sound image subspace trajectory of that audio element while in motion, the sound radiation range of the speaker corresponding to that audio element during its movement must be able to cover the sound image subspace trajectory of that audio element.
[0049] In one embodiment, for each audio element, it is necessary to identify, from the target playback space, a loudspeaker whose sound radiation range matches the sound image subspace trajectory corresponding to that audio element, and use it as the loudspeaker corresponding to the audio element. The sound image subspace trajectory includes the sound image position of the corresponding audio element at each moment in the audio to be played; therefore, it is necessary to identify loudspeakers whose sound radiation range during movement can cover each sound image position of the audio element.
[0050] In one implementation scenario, each speaker is a rotatable speaker installed at a preset position in the target playback space. Each speaker has a corresponding radiation range in each posture. The aforementioned sound radiation range is the sum of the radiation ranges corresponding to each posture of the speaker.
[0051] In one embodiment, the total radiation range (the union of the radiation ranges of all postures) of each loudspeaker can be determined first, and then the set of sound image positions of the audio element at each moment can be extracted. All loudspeakers are traversed one by one for verification: if the total radiation range of a loudspeaker can contain all the sound image positions of the audio element, it is determined to be a matched loudspeaker.
[0052] When no single loudspeaker can cover the entire sound image trajectory, the combination of two or more loudspeakers is traversed, the total radiation range of each combination is calculated (the union of the radiation ranges of all loudspeakers), and it is verified whether the total range of the combination can cover all sound image positions. Finally, the combination with the fewest loudspeakers and the highest overlap is selected as the matching cooperative loudspeaker group (i.e., the loudspeaker corresponding to the audio element).
[0053] S42: Determine the target motion trajectory of the loudspeaker corresponding to each audio element by using the sound image subspace trajectory of each audio element.
[0054] It should be noted that the sound image subspace trajectory includes the sound image position of the corresponding audio element at each moment in the audio to be played. Furthermore, for each audio element, the target motion trajectory of its corresponding speaker includes the target posture of the speaker when pointing to the sound image position at that moment. This target posture can be calculated based on the sound image position of the audio element at the corresponding moment and the installation position of the corresponding speaker.
[0055] It is understandable that each audio segment has its own duration, and the time taken to play the audio is consistent with the duration of the audio itself, and the playback sequence of the audio corresponds perfectly to the chronological order of the audio segments. Therefore, in order to ensure that the speaker can accurately reproduce the sound image position corresponding to each playback moment, the target pose of the speaker pointing to the sound image position at each moment can be used sequentially according to the playback sequence of the audio, forming a motion trajectory that changes continuously with the playback time, and this trajectory is used as the target motion trajectory of the speaker corresponding to the audio element.
[0056] For example, please refer to Figure 5 , Figure 5 This is a schematic diagram of an embodiment of the target motion trajectory of the loudspeaker provided in this application. Figure 5 The orange area corresponds to the target movement trajectory of the speaker corresponding to the human voice.
[0057] In one implementation scenario, each speaker is a rotatable speaker installed at a preset position in the target playback space. The target pose of the speaker at each moment includes: the azimuth angle on the horizontal plane and the polar angle on the vertical plane at the corresponding moment. Here, the azimuth angle and polar angle are the angles to which the speaker should rotate at the corresponding moment, not the angles that need to be changed.
[0058] Specifically, please refer to Figure 6 , Figure 6 This is a flowchart illustrating an embodiment of determining the target pose of a loudspeaker at various times, as provided in this application. This embodiment includes: S61: For the sound image position at each moment, obtain the components of the relative vector between the sound image position and the corresponding speaker installation position in each target direction; each target direction includes: a first direction, a second direction and a third direction, the first direction and the second direction form a horizontal plane, and the plane containing the third direction is the vertical plane.
[0059] In this embodiment, please continue to refer to Figure 5 Since each speaker is a rotatable speaker with a fixed installation position, the rotation of the speaker is centered on that installation position. Therefore, to facilitate determining the movement trajectory of the speaker, for each speaker, the installation position S(x) of the speaker in the target playback space can be used as the reference point. s , y s , z s Establish a spherical coordinate system with as the origin. ), where r is a constant (the distance from the speaker's rotation center to the edge of the speaker housing, which does not change during rotation). It is the polar angle (the angle between the polar angle and the positive Z-axis). It is the azimuth angle (the angle between the azimuth and the positive X-axis in the XY plane. The first direction represents the Y-axis of the XY plane, the second direction represents the X-axis of the XY plane, and the third direction represents the Z-axis).
[0060] It should be noted that the sound image position P at each moment can be used first. Δt (x Δt ,y Δt , z Δt ) and the corresponding speaker mounting position S(x) s , y s , z s Determine the relative vector d between the two:
[0061] Wherein, the component of the relative vector d in the first direction is dy, the component in the second direction is dx, the component in the third direction is dz, and the relative distance is the magnitude of the relative vector.
[0062] S62: Obtain the ratio of the components of the relative vector in the first and second directions, perform an arctangent operation on the component ratio to obtain the azimuth angle of the speaker on the horizontal plane; and obtain the ratio of the component in the third direction to the relative distance, perform an inverse cosine operation on the ratio to obtain the polar angle of the speaker on the vertical plane.
[0063] Among these, determining the azimuth angle of the loudspeaker on the horizontal plane. and polar angle The corresponding formulas are as follows:
[0064] Where, relative distance is the magnitude of the relative vector, and relative distance... The formula for determining it is as follows:
[0065] It should be noted that the target motion trajectory of the loudspeaker corresponds to the target pose of the loudspeaker when pointing to the sound image position at each moment, according to the sequence of audio playback, forming a continuously changing motion trajectory over time. Therefore, the target motion trajectory of the loudspeaker is composed of a series of target poses. In scenarios where the loudspeaker is a rotatable loudspeaker with a fixed installation position, to determine how to control the loudspeaker's rotation to reach the corresponding target pose at a given moment, it is necessary to know the loudspeaker's pose before rotation. Therefore, before controlling the loudspeaker to move along the target motion trajectory, it is necessary to first determine the loudspeaker's initial pose. Then, in controlling the loudspeaker to move along the target motion trajectory, the loudspeaker's initial pose is used as the starting point for movement, controlling the loudspeaker to move along the target motion trajectory.
[0066] In order to reduce the amplitude of speaker posture adjustment at the initial moment of movement, the speaker's posture when pointing to the center of the sound field can be used as the speaker's initial posture, and the speaker can be rotated to the corresponding initial posture before controlling the speaker to move according to the target motion trajectory.
[0067] Furthermore, it should be noted that in scenarios where different speakers are used to reproduce the sound image subspace trajectories of different audio elements, the movements of different speakers are independent of each other and there is no mutual interference. Moreover, the playback sub-audio corresponding to different audio elements is different. Therefore, multiple independent and non-interfering auditory focal points can be formed simultaneously in the target playback space, where each focal point corresponds to a sound field center.
[0068] Based on this, in one specific embodiment, before controlling the loudspeakers to move according to the target trajectory, at least one sound field center can be determined in the target playback space; then, the initial pose of the loudspeakers associated with each sound field center is determined; for each sound field center, the loudspeakers associated with the sound field center are controlled to rotate to the initial pose; wherein, when the loudspeakers associated with the sound field centers rotate to the initial pose, they point towards the sound field center. Each sound field center can be pre-associated with its corresponding loudspeaker so that the corresponding loudspeaker can be quickly located based on the position of the sound field center.
[0069] For example, please refer to Figure 7 , Figure 7 This is a schematic diagram of an embodiment of the in-vehicle sound field center provided in this application. For example... Figure 7As shown, the car interior contains two sound field centers: C1, which corresponds to the front row, and C2, which corresponds to the rear left rear position. Before the speakers are controlled to move according to the target motion trajectory, all the speakers in the front row point to the sound field center C1, and all the speakers in the rear row point to the sound field center C2.
[0070] In one embodiment, at least one sound field center can be determined in the target playback space by utilizing the listener's listening needs. Taking the target playback space as an in-vehicle cabin as an example, multiple sound field modes can be provided to passengers. The sound field mode is used to indicate the area where the sound field space is to be constructed, including but not limited to the entire vehicle, front row, rear row, driver's seat, front passenger seat, left rear, right rear, and other vehicle cabin positions. When a passenger selects a sound field mode that includes multiple seats, such as the entire vehicle, front row, or rear row, the sound field center can be defined using pre-calibrated cabin space center point data. When a passenger selects a sound field mode that includes only a single seat, such as the driver's seat or front passenger seat, the three-dimensional coordinate P of the current passenger's head is defined as the sound field center.
[0071] Among them, different passengers can choose the corresponding sound field mode according to their own listening needs: if the areas corresponding to the sound field modes selected by different passengers overlap (for example, one passenger selects the whole car and another passenger selects the front row), then the pre-marked position in the larger area is used as the sound field center; if the areas corresponding to the sound field modes selected by different passengers do not overlap (for example, one passenger selects the front row and another passenger selects the left rear), then the front row marked position and the left rear passenger's head position are used as their respective sound field center coordinates.
[0072] Of course, users can also select the spatial type of audio playback, where the spatial type refers to the panoramic surround sound space that the user expects; for example, near-field surround, recording studio, concert, concert hall, etc., and the audio will be played according to the selected spatial type when controlling the speakers to play audio.
[0073] In another embodiment, each seat in the vehicle can be identified to determine whether there are passengers in each seat (this can be done using cameras in the vehicle cabin and / or seat pressure sensors). Then, based on the identification results of each seat, at least one sound field center can be determined. The sound field center can be determined based on the area where the passengers are located. For example, if only the driver's seat has a passenger, the driver's head position is used as the sound field center; if both the driver's seat and the left rear seat have passengers, the driver's head position and the left rear seat passenger's head position are used as sound field centers respectively; or the midpoint between the driver's and the left rear seat passenger's head positions can be used as the sound field center.
[0074] In some embodiments, the listening area corresponding to each sound field center includes at least one listener; in controlling the loudspeaker to play the audio to be played while in motion, the method further includes: for each listening area corresponding to each sound field center, monitoring the change information of the head position of each listener in the listening area; in response to the change information of each head position indicating that the head displacement of each listener is greater than a preset threshold (e.g., 30 cm), re-executing the step of determining at least one sound field center in the target playback space.
[0075] This method is beneficial when the position of the listener corresponding to the listening area changes significantly, so that the center of each sound field can be re-determined based on the change in the listener's position, ensuring that the listener is always in the optimal listening position.
[0076] In other embodiments, the process of controlling the loudspeaker to play the audio to be played while in motion further includes: for each listening area corresponding to the center of each sound field, obtaining the difference between the sound pressure received by each listening position in the listening area and the sound pressure output by the loudspeaker; in response to the difference being greater than a preset difference (e.g., 3dB), returning to the step of using the sound image space trajectory to determine the target motion trajectory of the loudspeaker in the target playback space.
[0077] This method is beneficial for re-determining the speaker's pointing position when there is a large deviation in the speaker's pointing position, so as to calibrate the corresponding positional deviation.
[0078] Please see Figure 8 , Figure 8 This is a schematic diagram of a framework of an embodiment of the audio playback device provided in this application. In this embodiment, the audio playback device 80 includes: an acquisition module 81, a determination module 82, and a control module 83. The acquisition module 81 is used to acquire the sound image spatial trajectory of audio elements in the audio to be played; the determination module 82 is used to determine the target motion trajectory of the speaker in the target playback space using the sound image spatial trajectory; the control module 83 is used to control the speaker to move according to the target motion trajectory, and to control the speaker to play the audio to be played while moving; wherein, the sound image motion trajectory of the audio to be played played when the speaker moves according to the target motion trajectory matches the sound image spatial trajectory of the audio elements.
[0079] In some embodiments, the audio to be played includes a sub-audio of at least one audio element; the acquisition module 81 acquires the sound image space trajectory of the audio element in the audio to be played, including: separating the sub-audio corresponding to each audio element from the audio to be played; parsing the sound image sub-space trajectory of the corresponding audio element from each sub-audio; the sound image space trajectory includes the sound image sub-space trajectory of each audio element.
[0080] In some embodiments, the number of audio elements in the audio to be played is at least one, the sound image space trajectory includes the sound image subspace trajectory of each audio element, and the number of speakers in the target playback space is multiple; the determining module 82 uses the sound image space trajectory to determine the target motion trajectory of the speakers in the target playback space, including: using the sound image subspace trajectory of each audio element to determine the speaker corresponding to each audio element in the target playback space; and using the sound image subspace trajectory of each audio element to determine the target motion trajectory of the speaker corresponding to each audio element.
[0081] In some embodiments, the loudspeaker corresponding to each audio element is determined from the target playback space using the sound image subspace trajectory of each audio element, including: for each audio element, finding loudspeakers whose sound radiation range matches the sound image subspace trajectory of the audio element from the target playback space, and using them as the loudspeakers corresponding to the audio element.
[0082] In some embodiments, the audio to be played includes sub-audio of each audio element; before the control module 83 controls the speaker to play the audio to be played in motion, the audio playback device 80 is further configured to: for each audio element, send the sub-audio of the audio element to its corresponding speaker; the control module 83 controls the speaker to play the audio to be played in motion, including: controlling the speaker corresponding to each audio element to play the sub-audio of the corresponding audio element in motion.
[0083] In some embodiments, the sound image subspace trajectory includes: the sound image position of the corresponding audio element at each moment in the audio to be played; controlling the speaker corresponding to each audio element to play the sub-audio of the corresponding audio element while in motion, including: for each audio element, obtaining the speaker distance of the audio element at each moment, wherein the speaker distance is the relative distance between the sound image position at the corresponding moment and the installation position of the corresponding speaker; using the speaker distance at each moment to determine the playback sound pressure of the sub-audio at the corresponding moment; controlling the corresponding speaker to play the sub-audio of the audio element according to each playback sound pressure while in motion.
[0084] In some embodiments, each speaker is a rotatable speaker, and the sound image subspace trajectory includes: the sound image position of the corresponding audio element at each moment in the audio to be played; using the sound image subspace trajectory of each audio element, determining the target motion trajectory of the speaker corresponding to each audio element includes: for each audio element, calculating the target posture of the speaker when pointing to the sound image position at each moment by using the sound image position of the audio element at each moment and the installation position of the corresponding speaker; the target motion trajectory of the speaker includes: the target posture of the speaker when pointing to the sound image position at each moment.
[0085] In some embodiments, the target posture includes the azimuth angle of the speaker on the horizontal plane and the polar angle on the vertical plane; the target posture corresponding to the speaker pointing to the sound image position at each moment is determined by calculating using the sound image position of the audio element at each moment and the corresponding speaker installation position, including: for the sound image position at each moment, obtaining the components of the relative vector between the sound image position and the corresponding speaker installation position in each target direction; each target direction includes: a first direction, a second direction and a third direction, the first direction and the second direction forming the horizontal plane, and the plane containing the third direction being the vertical plane; obtaining the ratio of the components of the relative vector in the first direction and the second direction, performing an arctangent operation on the component ratio to obtain the azimuth angle of the speaker on the horizontal plane; and obtaining the ratio of the component in the third direction to the relative distance, performing an inverse cosine operation on the ratio to obtain the polar angle of the speaker on the vertical plane; wherein, the relative distance is the magnitude of the relative vector.
[0086] In some embodiments, before the control module 83 controls the speaker to move according to the target motion trajectory, the audio playback device 80 is further configured to: determine at least one sound field center in the target playback space; determine the initial pose of the speaker associated with each sound field center; for each sound field center, control the speaker associated with the sound field center to rotate to the initial pose; wherein the speaker associated with the sound field center points to the sound field center when it rotates to the initial pose; the control module 83 controls the speaker to move according to the target motion trajectory, including: using the initial pose of the speaker as the starting point of the motion, and controlling the speaker to move according to the target motion trajectory.
[0087] In some embodiments, the listening area corresponding to each sound field center includes at least one listener; while the control module 83 controls the loudspeaker to play the audio to be played in motion, the audio playback device 80 is further configured to: monitor the change information of the head position of each listener in the listening area corresponding to each sound field center; in response to each change information indicating that the head displacement of the corresponding listener is greater than a preset threshold, re-execute the step of determining at least one sound field center in the target playback space; and / or, while the control module 83 controls the loudspeaker to play the audio to be played in motion, the audio playback device 80 is further configured to: obtain the difference between the sound pressure received by each listening position in the listening area and the sound pressure output by the loudspeaker for each listening area corresponding to each sound field center; in response to the difference being greater than a preset difference, return to the step of determining the target motion trajectory of the loudspeaker in the target playback space using the sound image space trajectory.
[0088] Please see Figure 9 , Figure 9 This is a schematic diagram of a framework of an embodiment of the electronic device provided in this application. In this embodiment, the electronic device 90 includes a memory 91 and a processor 92 coupled to each other.
[0089] The memory 91 stores program instructions, and the processor 92 executes the program instructions stored in the memory 91 to implement the steps of any of the above-described method implementations. In a specific implementation scenario, the electronic device 90 may include, but is not limited to, a microcomputer or a server. In addition, the electronic device 90 may also include mobile devices such as laptops and tablets, which are not limited here.
[0090] Specifically, processor 92 controls itself and memory 91 to implement the steps of any of the above embodiments. Processor 92 can also be referred to as a CPU (Central Processing Unit). Processor 92 may be an integrated circuit chip with signal processing capabilities. Processor 92 can also be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. A general-purpose processor can be a microprocessor or any conventional processor. Furthermore, processor 92 can be implemented using integrated circuit chips.
[0091] Please see Figure 10 , Figure 10 This is a schematic diagram of the audio playback system provided in this application. In this embodiment, the audio playback system 100 includes a speaker rotation assembly 10 and the aforementioned electronic device 90. The speaker rotation assembly 10 is communicatively connected to the electronic device 90, and the speaker rotation assembly 10 is used to control the movement of the speaker.
[0092] Please see Figure 11 , Figure 11This is a schematic diagram of the framework of the computer-readable storage medium provided in this application. The computer-readable storage medium 110 of this application embodiment stores program instructions 111, which, when executed, implement the methods provided in any embodiment or any non-conflicting combination of the above methods. The program instructions 111 can form a program file and be stored in the computer-readable storage medium 110 in the form of a software product, so that a computer device (which may be a personal computer, server, or network device, etc.) can execute all or part of the steps of the methods of various embodiments of this application. The aforementioned computer-readable storage medium 110 includes various media capable of storing program code, such as a USB flash drive, portable hard drive, read-only memory (ROM), random access memory (RAM), magnetic disk, or optical disk, or terminal devices such as computers, servers, mobile phones, and tablets.
[0093] The above solution can control the speaker to move along the target motion trajectory, and the sound image motion trajectory of the audio to be played when the speaker moves along the target motion trajectory matches the sound image spatial trajectory of the audio element. Therefore, this method can make the sound image motion trajectory of the audio element perceived by the human ear in the target playback space consistent with the sound image spatial trajectory of the audio element in the audio to be played, thereby achieving accurate restoration of the sound image spatial trajectory of the audio element in the audio to be played, and giving the user an immersive panoramic sound experience.
[0094] In some embodiments, the functions or modules of the apparatus provided in this disclosure can be used to perform the methods described in the above method embodiments. The specific implementation can be referred to the description of the above method embodiments, and for the sake of brevity, it will not be repeated here.
[0095] The description of the various embodiments above tends to emphasize the differences between the various embodiments. The similarities or similarities between them can be referred to, and for the sake of brevity, they will not be repeated here.
[0096] In the several embodiments provided in this application, it should be understood that the disclosed methods and apparatus can be implemented in other ways. For example, the apparatus implementations described above are merely illustrative. For instance, the division of modules or units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between devices or units may be electrical, mechanical, or other forms.
[0097] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment, depending on actual needs.
[0098] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0099] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) or processor to execute all or part of the steps of the methods of various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0100] The above description is merely an embodiment of this application and does not limit the patent scope of this application. Any equivalent structural or procedural transformations made using the content of this application's specification and drawings, or direct or indirect applications in other related technical fields, are similarly included within the patent protection scope of this application.
[0101] If the technical solution of this application involves personal information, the product using this technical solution has clearly informed the user of the personal information processing rules and obtained the user's voluntary consent before processing the personal information. If the technical solution of this application involves sensitive personal information, the product using this technical solution has obtained the user's separate consent before processing the sensitive personal information, and also meets the requirement of "express consent". For example, at personal information collection devices such as cameras, clear and prominent signs are set up to inform users that they have entered the scope of personal information collection and that personal information will be collected. If an individual voluntarily enters the collection scope, it is deemed that they have agreed to the collection of their personal information; or on the personal information processing device, with clear signs / information informing users of the personal information processing rules, authorization is obtained from the individual through pop-up information or by asking the individual to upload their personal information; wherein, the personal information processing rules may include information such as the personal information processor, the purpose of personal information processing, the processing method, and the types of personal information processed.
Claims
1. An audio playback method, characterized in that, The method includes: Obtain the sound image space trajectory of audio elements in the audio to be played; Using the aforementioned sound image spatial trajectory, the target motion trajectory of the speaker in the target playback space is determined; The speaker is controlled to move according to the target motion trajectory, and the speaker is controlled to play the audio to be played during the movement; wherein the sound image motion trajectory of the audio to be played played when the speaker moves according to the target motion trajectory matches the sound image spatial trajectory of the audio element.
2. The method according to claim 1, characterized in that, The audio to be played includes at least one sub-audio element; The step of obtaining the sonic spatial trajectory of audio elements in the audio to be played includes: Extract the sub-audio elements corresponding to each audio element from the audio to be played; From each of the sub-audio elements, the sound image subspace trajectory of the corresponding audio element is parsed out; the sound image subspace trajectory includes the sound image subspace trajectory of each of the audio elements.
3. The method according to claim 1, characterized in that, The number of audio elements in the audio to be played is at least one, the sound image space trajectory includes the sound image subspace trajectory of each audio element, and the number of speakers in the target playback space is multiple. The step of determining the target motion trajectory of the speaker in the target playback space using the sound image spatial trajectory includes: Using the sound image subspace trajectory of each audio element, determine the speaker corresponding to each audio element from the target playback space; The target motion trajectory of the loudspeaker corresponding to each audio element is determined by using the sound image subspace trajectory of each audio element.
4. The method according to claim 3, characterized in that, The step of determining the speaker corresponding to each audio element from the target playback space using the sound image subspace trajectory of each audio element includes: For each audio element, find the loudspeaker whose sound radiation range matches the sound image subspace trajectory of the audio element from the target playback space, and use it as the loudspeaker corresponding to the audio element.
5. The method according to claim 3, characterized in that, The audio to be played includes sub-audio of each audio element; Before controlling the speaker to play the audio to be played during movement, the method further includes: For each of the aforementioned audio elements, the sub-audio of the audio element is sent to its corresponding speaker; The control of the speaker to play the audio during movement includes: Control the speakers corresponding to each audio element to play the sub-audio of the corresponding audio element while it is in motion.
6. The method according to claim 5, characterized in that, The sound image subspace trajectory includes: the sound image position of the corresponding audio element at each moment in the audio to be played; The control of the speakers corresponding to each audio element to play the sub-audio of the corresponding audio element during movement includes: For each audio element, the speaker distance of the audio element at each time moment is obtained, wherein the speaker distance is the relative distance between the sound image position and the corresponding speaker installation position at the corresponding time moment; Using the speaker distance at each moment, the playback sound pressure level of the sub-audio at the corresponding moment is determined; When the corresponding speaker is in motion, it plays the sub-audio of the audio element according to the specified sound pressure level.
7. The method according to claim 3, characterized in that, Each of the aforementioned speakers is a rotatable speaker, and the sound image subspace trajectory includes: the sound image position of the corresponding audio element at each moment in the audio to be played; The step of determining the target motion trajectory of the loudspeaker corresponding to each audio element by utilizing the sound image subspace trajectory of each audio element includes: For each audio element, the target posture of the speaker when it points to the audio image position at each time is determined by using the audio image position of the audio element at each time and the installation position of the corresponding speaker. The target motion trajectory of the speaker includes the target posture of the speaker when it points to the audio image position at each time.
8. The method according to claim 7, characterized in that, The target attitude includes the azimuth angle of the speaker in the horizontal plane and the polar angle in the vertical plane; The step of calculating and determining the target posture corresponding to the speaker pointing to the sound image position at each moment by using the sound image position of the audio element at each moment and the installation position of the corresponding speaker includes: For the sound image position at each time moment, obtain the components of the relative vector between the sound image position and the corresponding speaker installation position in each target direction; each target direction includes: a first direction, a second direction and a third direction, the first direction and the direction constitute the horizontal plane, and the plane containing the third direction is the vertical plane; The ratio of the components of the relative vector in the first direction and the second direction is obtained, and the arctangent operation is performed on the component ratio to obtain the azimuth angle of the speaker in the horizontal plane; and the ratio of the component in the third direction to the relative distance is obtained, and the arccosine operation is performed on the ratio to obtain the polar angle of the speaker in the vertical plane; wherein, the relative distance is the magnitude of the relative vector.
9. The method according to claim 1, characterized in that, Before controlling the speaker to move according to the target motion trajectory, the method further includes: Determine at least one sound field center in the target playback space; Determine the initial pose of the loudspeakers associated with each of the sound field centers; For each of the sound field centers, the loudspeaker associated with the sound field center is controlled to rotate to the initial pose; wherein, when the loudspeaker associated with the sound field center rotates to the initial pose, it points towards the sound field center; Controlling the speaker to move according to the target motion trajectory includes: Using the initial pose of the speaker as the starting point of motion, the speaker is controlled to move according to the target motion trajectory.
10. The method according to claim 9, characterized in that, Each of the aforementioned sound field centers corresponds to a listening area that includes at least one listener; In controlling the speaker to play the audio to be played while in motion, the method further includes: for each listening area corresponding to the sound field center, monitoring the change information of the head position of each listener in the listening area; in response to each of the change information indicating that the head displacement of the corresponding listener is greater than a preset threshold, re-executing the step of determining at least one sound field center in the target playback space; And / or, during the process of controlling the loudspeaker to play the audio to be played while in motion, the method further includes: for each listening area corresponding to the center of the sound field, obtaining the difference between the sound pressure received by each listening position in the listening area and the sound pressure output by the loudspeaker; in response to the difference being greater than a preset difference, returning to the step of determining the target motion trajectory of the loudspeaker in the target playback space using the sound image space trajectory.
11. An audio playback device, characterized in that, The device includes: The acquisition module is used to acquire the sound image spatial trajectory of audio elements in the audio to be played; The determining module is used to determine the target motion trajectory of the speaker in the target playback space using the sound image space trajectory; A control module is used to control the speaker to move according to the target motion trajectory, and to control the speaker to play the audio to be played during the movement; wherein, the sound image motion trajectory of the audio to be played played when the speaker moves according to the target motion trajectory matches the sound image spatial trajectory of the audio element.
12. An electronic device, characterized in that, Including interconnected memory and processor, The memory stores program instructions; The processor is used to execute program instructions stored in the memory to implement the method according to any one of claims 1-10.
13. An audio playback system, characterized in that, It includes a speaker rotating assembly and an electronic device as described in claim 12, wherein the speaker rotating assembly is communicatively connected to the electronic device and is used to control the movement of the speaker.
14. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores program instructions that can be executed by a processor to implement the method according to any one of claims 1-10.