An audio playing method and device, electronic equipment, computer program product
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-05-27
- Publication Date
- 2026-08-11
AI Technical Summary
[0003]本申请提供一种音频播放方法和装置、电子设备、计算机程序产品,用于解决如何使得音频始终对着用户播放的问题
[0005] The audio playback method provided in this application updates the real-time playback orientation of the target sub-audio based on the user's real-time location. Since the updated real-time playback orientation maintains the preset requirements for the playback position and beam pointing angle of the target sub-audio relative to the user, the relative playback position and relative beam pointing angle of the target sub-audio remain unchanged during the user's movement, thereby ensuring that the target sub-audio is always played towards the user, thus improving the user's auditory experience.
Smart Images

Figure CN122554770A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of audio processing technology, specifically to an audio playback method and apparatus, electronic device, and computer program product. Background Technology
[0002] Immersive sound technology's audio renderer calculates how the sound of each audio object should be played based on the actual speaker layout, thus creating a highly immersive three-dimensional sound field. In related technologies, the playback parameters of audio objects do not change with the user's state during playback, resulting in a poor auditory experience. Summary of the Invention
[0003] This application provides an audio playback method and apparatus, electronic device, and computer program product to solve the problem of how to ensure that audio is always played to the user.
[0004] To achieve the above objectives, the technical solution adopted by the present invention is as follows: In a first aspect, embodiments of this application provide an audio playback method, the method comprising: acquiring a target sub-audio in an audio to be played, and collecting the real-time orientation of a user in a current area; updating the real-time playback orientation of the target sub-audio based on the real-time orientation to obtain an updated real-time playback orientation; controlling the playback position and beam pointing angle of the target sub-audio relative to the user to maintain preset requirements based on the updated real-time playback orientation; and controlling a speaker array set in the current area to play the target sub-audio during the playback of the audio to be played.
[0005] The audio playback method provided in this application updates the real-time playback orientation of the target sub-audio based on the user's real-time location. Since the updated real-time playback orientation maintains the preset requirements for the playback position and beam pointing angle of the target sub-audio relative to the user, the relative playback position and relative beam pointing angle of the target sub-audio remain unchanged during the user's movement, thereby ensuring that the target sub-audio is always played towards the user, thus improving the user's auditory experience.
[0006] In some embodiments, the real-time orientation includes the real-time orientation angle; the real-time playback orientation includes the real-time beam pointing angle; based on the real-time orientation, the real-time playback orientation of the target sub-audio is updated to obtain the updated real-time playback orientation, including: smoothing the real-time orientation angle to obtain the processed orientation angle; updating the real-time beam pointing angle based on the positional relationship between the beam pointing angle and the user and the processed orientation angle to obtain the updated real-time playback orientation.
[0007] In this embodiment, by smoothing the real-time orientation angle, instantaneous jitter and random noise during data acquisition can be eliminated, and the beam pointing angle can be smoothly changed, thereby ensuring stable playback of the target sub-audio and improving the user's auditory experience. Positional relationships are also considered when updating the beam pointing angle to ensure that the sound-emitting surface of the target sub-audio meets preset requirements, thus guaranteeing that the target sub-audio is always played towards the user.
[0008] In some embodiments, the method further includes: acquiring the user's real-time movement speed in the current area; smoothing the real-time orientation angle to obtain the processed orientation angle, including: determining a smoothing coefficient based on the real-time movement speed; the smoothing coefficient being positively correlated with the real-time movement speed; smoothing the real-time orientation angle based on the smoothing coefficient to obtain the processed orientation angle.
[0009] In this embodiment, the real-time orientation angle is smoothed by a smoothing coefficient that is positively correlated with the user's real-time movement speed. This allows the real-time orientation angle to change stably when the user is moving at low speed, thereby ensuring a stable update of the real-time sound beam pointing angle and a stable sound field. When the user is moving at high speed, the real-time orientation angle changes rapidly, thereby ensuring that the real-time sound beam pointing angle quickly follows the user and that the user can receive the target sub-audio without lag.
[0010] In some embodiments, real-time orientation includes real-time location; real-time playback orientation includes real-time playback location; the method further includes: acquiring the user's real-time movement speed and real-time movement direction angle in the current area; updating the real-time playback orientation of the target sub-audio based on the real-time orientation to obtain the updated real-time playback orientation, including: performing position prediction based on real-time location, real-time movement speed, real-time movement direction angle and data processing delay to obtain the user's predicted location at the next moment; updating the real-time playback location based on the predicted location to obtain the updated real-time playback orientation.
[0011] In this embodiment, the user's real-time position, real-time movement speed, real-time movement direction angle, and data processing delay are used to predict the position at the next moment, and the predicted position is used to update the real-time playback position of the target sub-audio. This can avoid the sound lag caused by data processing delay, thereby improving the user's auditory experience.
[0012] In some embodiments, the method further includes: calculating the real-time distance from the user to the speaker array based on the real-time location; and controlling the speaker array located in the current area to play target sub-audio based on the updated real-time playback orientation during the playback of audio to be played, including: adjusting the gain of the speaker array based on the real-time distance using a distance attenuation model; and controlling the speaker array located in the current area to play target sub-audio based on the adjusted gain during the playback of audio to be played, based on the updated real-time playback orientation.
[0013] In this embodiment, by adjusting the gain of the speaker array based on the real-time distance between the user and the speaker array, and playing audio based on the speaker array with the adjusted gain, the volume of the target sub-audio heard by the user can be consistent in different directions.
[0014] In some embodiments, obtaining a target sub-audio in an audio to be played includes: generating a sub-audio based on the audio to be played using an audio separation model; and determining the target sub-audio from the sub-audio based on the type of the sub-audio and a preset type priority.
[0015] In this embodiment of the application, the target sub-audio is determined from the sub-audio by the type and type priority of the sub-audio. This allows the high-priority sub-audio to always follow the user's changes and automatically switches to the second-highest priority sub-audio after the high-priority sub-audio disappears, thereby ensuring that there is always a sub-audio following the user.
[0016] In some embodiments, the gesture action type is determined based on the user's gesture changes within a preset time period; if the gesture action type matches the preset action type, the real-time playback position of the target controlled sub-audio is updated based on the user's hand termination position within the preset time period to obtain the updated real-time playback position; the target controlled sub-audio is the audio in the sub-audio; during the playback of the audio to be played, the speaker array is controlled to play the sub-audio based on the updated real-time playback position.
[0017] In this embodiment of the application, by recognizing the user's gesture action type and updating the playback position of the target controlled sub-audio based on the user's gesture action, the user can actively control the playback position of the audio, thereby improving the user's sense of participation.
[0018] Secondly, embodiments of this application provide an audio playback device, comprising: an acquisition unit for acquiring a target sub-audio in an audio to be played; a collection unit for collecting the real-time location of a user in a current area; an update unit for updating the real-time playback location of the target sub-audio based on the real-time location to obtain an updated real-time playback location; the updated real-time playback location controls the playback position of the target sub-audio relative to the user and the beam pointing angle to maintain preset requirements; and a control unit for controlling a speaker array located in the current area to play the target sub-audio based on the updated real-time playback location during the playback of the audio to be played.
[0019] Thirdly, embodiments of this application provide an electronic device, including a memory and a processor. The memory stores a computer program that can run on the processor, and the processor executes the program to implement the steps in the method of the first aspect.
[0020] Fourthly, embodiments of this application provide a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the method in the first aspect.
[0021] Fifthly, embodiments of this application provide a computer program product including instructions, comprising a computer program or instructions that, when executed by a processor, implement the steps of the method in the first aspect. Attached Figure Description
[0022] Figure 1 A schematic diagram of the implementation process of an audio playback method provided in this application embodiment. Figure 1 ; Figure 2 A schematic diagram of the implementation process of an audio playback method provided in this application embodiment. Figure 2 ; Figure 3 A schematic diagram of the implementation process of an audio playback method provided in this application embodiment. Figure 3 ; Figure 4 A schematic diagram of the implementation process of an audio playback method provided in this application embodiment. Figure 4 ; Figure 5 A schematic diagram of the implementation process of an audio playback method provided in this application embodiment. Figure 5 ; Figure 6 A schematic diagram of the implementation process of an audio playback method provided in this application embodiment. Figure 6 ; Figure 7 A schematic diagram of the implementation process of an audio playback method provided in this application embodiment. Figure 7 ; Figure 8 This application provides a schematic diagram of the structure of an interactive system. Figure 9 A schematic diagram of the implementation process of an audio playback method provided in this application embodiment. Figure 8 ; Figure 10 This is a schematic diagram of the structure of the audio playback device proposed in the embodiments of this application; Figure 11 This is a schematic diagram of the structure of the electronic device proposed in the embodiments of this application. Detailed Implementation
[0023] The technical solutions in the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings.
[0024] To make the objectives, technical solutions, and advantages of this application clearer, the application will be further described below in conjunction with the accompanying drawings. The embodiments described below are only some embodiments of this application, not all embodiments. Therefore, the described embodiments should not be regarded as limitations on this application. All other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0025] In the following description, references to "some embodiments" or "other embodiments" describe a subset of all possible embodiments, but "some embodiments" or "other embodiments" may be the same subset or different subsets of all possible embodiments and may be combined with each other without conflict.
[0026] In the following description, the terms "first" and "second" are used merely to distinguish similar objects and do not represent a specific ordering of objects. It is understood that "first" and "second" may be interchanged in a specific order or sequence where permitted, so that the embodiments of this application described herein can be implemented in an order other than that illustrated or described herein.
[0027] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terminology used herein is for the purpose of describing embodiments of this application only and is not intended to limit this application.
[0028] In related technologies, during audio playback, the audio's position and angle relative to the user do not move with the user's movement. In other words, the audio's playback position relative to the user cannot maintain the preset requirements as the user moves, resulting in a poor listening experience for the user.
[0029] To address the aforementioned issues, this application proposes an audio playback method. In this method, the real-time playback orientation of the target sub-audio is updated by acquiring the user's real-time location within the current area, resulting in an updated real-time playback orientation. Then, the updated real-time playback orientation is used to control a speaker array within the current area to play the target sub-audio. Since the updated real-time playback orientation is used to maintain the target sub-audio's playback position and beam pointing angle relative to the user within preset requirements, the target sub-audio can always be played towards the user as the user moves, thereby improving the user's auditory experience.
[0030] The technical solutions in the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings.
[0031] Embodiments of this application propose an audio playback method, such as... Figure 1 As shown, the method includes the following steps: Step 101: Obtain the target sub-audio in the audio to be played, and collect the user's real-time location in the current area.
[0032] In this embodiment of the application, the audio to be played refers to the original mixed audio that needs to be played, which includes audio from one or more different sound sources.
[0033] For example, the audio to be played is a musical performance, a film clip, or a conference recording.
[0034] In this embodiment of the application, the target sub-audio refers to the audio that is separated from the audio to be played and needs to be played dynamically following the user.
[0035] For example, the target sub-audio is the lead vocals in a musical performance or the narration in a film or television dialogue.
[0036] In some embodiments, after separating the audio from different sound sources from the audio to be played, any one of the audio sources can be used as the target sub-audio, or the target sub-audio can be determined from the audio from different sound sources based on preset rules.
[0037] In this embodiment of the application, the current area refers to the physical space where the current user is located.
[0038] For example, the current area is a living room, a meeting room, or a car.
[0039] In this embodiment of the application, real-time location refers to the spatial status information of the user within the current area.
[0040] In some embodiments, real-time orientation includes real-time position and real-time orientation angle.
[0041] In some embodiments, the user's real-time orientation can be the real-time position and orientation angle of the user's body, or the real-time position and orientation angle of the user's head.
[0042] Step 102: Based on the real-time orientation, update the real-time playback orientation of the target sub-audio to obtain the updated real-time playback orientation; the updated real-time playback orientation controls the playback position of the target sub-audio relative to the user and the beam pointing angle to maintain the preset requirements.
[0043] In this embodiment of the application, the real-time playback orientation refers to the set of control parameters when the target sub-audio is played in the current area.
[0044] In some embodiments, real-time playback orientation includes real-time playback position and real-time sound beam pointing angle. The real-time playback position refers to the coordinates of the target sub-audio as a virtual audio object within a three-dimensional sound field constructed based on the current region. These coordinates are used to calculate the signal distribution weight of each speaker in the speaker array based on the layout of the speaker array within the current region and in conjunction with immersive sound rendering technology, thereby audibly simulating the effect that the target sub-audio is emitting sound from these coordinates. The real-time sound beam pointing angle refers to the direction in which the sound beam of the target sub-audio as a virtual audio object propagates concentratedly in the three-dimensional sound field. This direction is used to determine the angular parameters of each speaker in the speaker array when playing the target sub-audio, based on the layout of the speaker array within the current region and in conjunction with immersive sound rendering technology, thereby audibly simulating that the target sub-audio is emanating from this direction. The sound beam pointing angle refers to the main radiation direction of the audio.
[0045] In this embodiment of the application, the playback position refers to the relative spatial relationship between the real-time playback position and the user's real-time position.
[0046] For example, the playback position is 0.5m directly in front.
[0047] In this embodiment of the application, the beam pointing angle refers to the relative angular relationship between the real-time beam pointing angle and the user's real-time orientation angle.
[0048] For example, the sound beam pointing angle is directly towards the user's ear.
[0049] In this embodiment of the application, the preset requirement refers to the pre-set auditory conditions that the target sub-audio should meet relative to the user's real-time orientation.
[0050] In some embodiments, in order to create automatically following directional audio, preset requirements may include the target sub-audio being played at a distance within a set distance range relative to the user in a set relative direction, and the sound beam pointing angle being directed towards the user.
[0051] For example, the preset requirement can be that the playback is 0.5-1m directly in front of the user's head and pointing towards the user's face; or it can be 0.3-0.5m directly to the left of the user's body and pointing towards the user's left ear.
[0052] In this embodiment of the application, since the user's position may change and the orientation angle may change within the current area, the fixed playback position and beam pointing angle of the audio cannot adapt to the user's dynamic changes, thus causing the target sub-audio to fail to meet the preset requirements during playback. Therefore, it is necessary to update the real-time playback orientation of the target sub-audio so that the playback position and beam pointing angle of the target sub-audio relative to the user maintain the preset requirements.
[0053] In some embodiments, when updating the real-time playback position of the target sub-audio, the user's position and orientation angle are taken into account. Therefore, to ensure that the playback position of the target sub-audio relative to the user is always within a set distance range in a set relative direction, the real-time playback position of the target sub-audio is updated based on the user's real-time position and real-time orientation angle to obtain the updated real-time playback position.
[0054] For example, the real-time position of the user's head is (U X U Y U Z ); The orientation angle includes the horizontal orientation angle β1 and the vertical orientation angle β2, and the real-time playback position (A X A Y A Z The update can be performed using the following formula: (1) (2) (3) Among them, U X U Y U Z These refer to the coordinates of the user's head on the X, Y, and Z axes, respectively; A X A Y A Z These refer to the coordinates of the user's head on the X, Y, and Z axes, respectively; M refers to the pre-set playback distance of the target sub-audio relative to the user, for example, M is 0.5m.
[0055] In some embodiments, in order to ensure that the target sub-audio is always played towards the user, the real-time beam pointing angle of the target sub-audio is updated based on the user's real-time orientation angle to obtain the updated real-time beam pointing angle.
[0056] In this embodiment, after obtaining the updated real-time playback position and the updated real-time beam pointing angle, the playback position and beam pointing angle of the target sub-audio relative to the user are calculated based on the user's real-time position and real-time orientation angle. Due to noise during data acquisition and delays during data calculation, the calculated playback position and beam pointing angle may have deviations. Therefore, it is necessary to determine whether the playback position and beam pointing angle of the target sub-audio relative to the user meet the preset requirements. If not, the updated real-time playback position and the updated real-time beam pointing angle are adjusted to maintain the preset requirements for the playback position and beam pointing angle of the target sub-audio relative to the user.
[0057] In some embodiments, maintaining the updated real-time playback orientation control target sub-audio's playback position and beam pointing angle relative to the user as preset requirements includes: maintaining the updated real-time playback orientation control target sub-audio's playback position and beam pointing angle relative to the user as preset.
[0058] In this embodiment of the application, the preset playback position refers to the preset relative position between the target sub-audio and the user.
[0059] For example, the preset playback position is that the target sub-audio is fixed 0.5m directly in front of the user's head.
[0060] In this embodiment of the application, the preset beam pointing angle refers to the preset main radiation direction of the target sub-audio relative to the user.
[0061] For example, the preset beam pointing angle is directed at the user's face.
[0062] In this embodiment of the application, after directly updating the real-time playback position and real-time beam pointing angle of the target sub-audio, the relative position of the target sub-audio with respect to the user and the main radiation direction with respect to the user can be determined, thereby achieving the purpose of controlling the target sub-audio to maintain a preset playback position and beam pointing angle with respect to the user.
[0063] Step 103: During the playback of the audio to be played, based on the updated real-time playback orientation, control the speaker array set in the current area to play the target sub-audio.
[0064] In this embodiment of the application, a loudspeaker array refers to a plurality of loudspeakers arranged in the current area.
[0065] In some embodiments, based on the updated real-time playback orientation, the signal amplitude of the target sub-audio that each speaker in the speaker array should output is calculated using spatial audio rendering technology, and the target sub-audio is played.
[0066] In this embodiment, the real-time playback position of the target sub-audio is updated based on the user's real-time location. Since the updated real-time playback position ensures that the playback position of the target sub-audio relative to the user and the beam pointing angle maintain preset requirements, no matter how the user moves or rotates, the playback position of the target sub-audio relative to the user is always within a set distance range in the set relative direction, and the beam pointing angle always points towards the user, thereby ensuring that the target sub-audio always maintains the preset requirements for playback, thereby improving the user's auditory experience.
[0067] In some embodiments, such as Figure 2 As shown, when updating the real-time playback orientation of the target sub-audio based on its real-time location to obtain the updated real-time playback orientation, the following steps may be included: Step 201: Smooth the real-time orientation angle to obtain the processed orientation angle.
[0068] In the embodiments of this application, smoothing refers to using a filtering algorithm to reduce noise and stabilize continuous real-time orientation angles.
[0069] In some embodiments, multiple real-time orientation angles are continuously acquired to form an orientation angle sequence, and then a corresponding filtering algorithm is used to smooth the orientation angle sequence to obtain the processed orientation angle.
[0070] For example, a first-order inertial filtering algorithm and a Kalman filtering algorithm are used to smooth the orientation angle sequence.
[0071] Step 202: Based on the positional relationship between the sound beam pointing angle and the user, as well as the processed orientation angle, update the real-time sound beam pointing angle to obtain the updated real-time playback orientation.
[0072] In this embodiment of the application, the positional relationship between the sound beam pointing angle and the user refers to the preset positional relationship between the main radiation direction of the target sub-audio and the user.
[0073] In some embodiments, to ensure that the beam pointing angle of the target sub-audio is always pointing towards the user, it is necessary to update the real-time beam pointing angle of the target sub-audio based on the positional relationship between the beam pointing angle and the user. First, the positional relationship between the beam pointing angle of the target sub-audio and the user is preset, and then the real-time beam pointing angle is updated based on the real-time orientation angle to obtain the updated real-time playback orientation.
[0074] For example, when the positional relationship between the beam pointing angle and the user is such that the target sub-audio is directly facing the user's ears, the reverse real-time orientation angle is used as the real-time beam pointing angle; when the positional relationship between the beam pointing angle and the user is such that the target sub-audio is directly facing the user's back, the real-time orientation angle is directly used as the real-time beam pointing angle.
[0075] In this embodiment, the real-time orientation angle is first smoothed to effectively suppress noise during data acquisition and data errors caused by slight shaking of the user. Then, based on the positional relationship between the sound beam pointing angle and the user, the real-time sound beam pointing angle is updated so that the main radiation direction of the target sub-audio accurately meets the preset requirements.
[0076] In some embodiments, the audio playback method provided in this application further includes the following steps: Step 300: Collect the user's real-time movement speed in the current area.
[0077] In this embodiment of the application, real-time motion speed refers to the instantaneous speed of the user in the current area.
[0078] In some embodiments, the user's real-time movement speed is obtained by measuring the real-time movement speed of the target part of the user (e.g., head, body) using sensors.
[0079] In some embodiments, such as Figure 3 As shown, the process of smoothing the real-time orientation angle to obtain the processed orientation angle may include the following steps: Step 301: Determine the smoothing coefficient based on the real-time motion speed; the smoothing coefficient is positively correlated with the real-time motion speed.
[0080] In this embodiment, the smoothing coefficient refers to the weighting coefficient in the filtering algorithm, used to filter the real-time orientation angle. The magnitude of this weighting coefficient affects the response speed and smoothness of the filtered real-time orientation angle to changes in the real-time orientation angle, thereby affecting the update degree of the real-time playback orientation of the target sub-audio. The larger the smoothing coefficient, the faster the filtered real-time orientation angle can follow changes in the real-time orientation angle; the smaller the smoothing coefficient, the slower the filtered real-time orientation angle follows changes in the real-time orientation angle.
[0081] In some embodiments, considering that the user's real-time movement speed may vary, using a fixed smoothing coefficient could result in audio lag when the user moves at high speeds and significant audio variations when moving at low speeds. Therefore, it is necessary to adjust the smoothing coefficient based on the user's real-time movement speed. Specifically, based on the real-time movement speed, a corresponding smoothing coefficient is determined through a preset mapping relationship between movement speed and coefficient. The coefficient in this mapping relationship is positively correlated with the movement speed.
[0082] For example, the mapping relationship is a monotonically increasing function or a piecewise increasing function.
[0083] Step 302: Smooth the real-time orientation angle based on the smoothing coefficient to obtain the processed orientation angle.
[0084] In some embodiments, a smoothing coefficient dynamically determined based on real-time motion speed is incorporated into the corresponding filtering algorithm, and then the filtering algorithm is used to smooth the real-time orientation angle to obtain the processed orientation angle.
[0085] In this embodiment, a smoothing coefficient determined based on the user's real-time movement speed is used to smooth the real-time orientation angle. This can suppress sudden changes in the real-time orientation angle when the user is moving at low speed, thereby ensuring stable updates of the real-time sound beam pointing angle and thus guaranteeing stable playback of the target sub-audio. When the user is moving at high speed, the real-time orientation angle changes rapidly, thereby ensuring that the real-time sound beam pointing angle quickly follows the user and guarantees that the user can receive the target sub-audio without lag.
[0086] In some embodiments, the audio playback method provided in this application further includes the following steps: Step 400: Collect the user's real-time movement direction angle in the current area.
[0087] In this embodiment, the real-time motion direction angle is used to characterize the direction in which the user moves in the current area.
[0088] In some embodiments, such as Figure 4 As shown, when updating the real-time playback orientation of the target sub-audio based on its real-time location to obtain the updated real-time playback orientation, the following steps may be included: Step 401: Based on real-time location, real-time motion speed, real-time motion direction angle and data processing delay, position prediction is performed to obtain the user's predicted position at the next moment.
[0089] In this embodiment of the application, data processing latency refers to the time required from data acquisition to the playback of the target sub-audio.
[0090] In some embodiments, the predicted position of the user at the next moment can be obtained by superimposing the predicted value on the real-time position. The predicted value is determined based on the real-time motion speed, real-time motion direction angle, and data processing delay.
[0091] Step 402: Based on the predicted position, update the real-time playback position to obtain the updated real-time playback orientation.
[0092] In some embodiments, due to data processing delays and the user's continuous movement, directly updating the real-time playback position of the target sub-audio using the user's real-time location results in audio lag because the user has already moved a distance by the time they actually hear the audio. Therefore, it is necessary to predict the location and use the predicted location to directly update the real-time playback position in the real-time playback orientation to obtain the updated real-time playback orientation.
[0093] In this embodiment of the application, by updating the real-time playback position using the user's future position, the audio lag caused by time delay can be avoided, thereby further ensuring that the playback orientation and beam pointing angle of the target sub-audio relative to the user maintain the preset requirements, and thus improving the user's listening experience.
[0094] In some embodiments, the audio playback method provided in this application further includes the following steps: Step 500: Calculate the real-time distance from the user to the speaker array based on the real-time location information.
[0095] In this embodiment of the application, real-time distance refers to the distance from the user to each speaker in the speaker array.
[0096] In some embodiments, as the user moves continuously, the distance between the user and each speaker changes synchronously. This distance directly affects the volume of the audio the user hears. If the speaker gain is not compensated, the volume of the target sub-audio will constantly change, causing auditory discomfort. Therefore, it is necessary to compensate the speaker gain using real-time distance.
[0097] In some embodiments, the coordinates of each speaker in the speaker array in the current area are pre-stored, and then the real-time distance is calculated based on the user's real-time location.
[0098] In some embodiments, such as Figure 5 As shown, when controlling the speaker array set in the current area to play the target sub-audio during the playback of the audio to be played, based on the updated real-time playback orientation, the following steps may be included: Step 501: Adjust the gain of the speaker array based on the real-time distance using a distance attenuation model.
[0099] In this embodiment of the application, the distance attenuation model refers to a preset mathematical operation model, which is used to describe the negative correlation between gain attenuation and distance increase.
[0100] For example, the distance decay model can be an inverse square decay model, a linear decay model, or a logarithmic decay model.
[0101] In some embodiments, after obtaining the real-time distance, the gain of each speaker is determined by a preset distance attenuation model, and then the gain of the speaker array is adjusted using the gain of each speaker.
[0102] Step 502: During the playback of the audio to be played, based on the updated real-time playback orientation, control the speaker array set in the current area to play the target sub-audio based on the adjusted gain.
[0103] In some embodiments, the target sub-audio is played using the gain-adjusted speaker array after gain adjustment of the speaker array.
[0104] In this embodiment, the gain of the speaker array is adjusted based on the real-time distance between the user and the speaker array, and the target sub-audio is played using the gain-adjusted speaker array. No matter where the user moves in the current area, the volume of the target sub-audio heard by the user remains unchanged, thereby enhancing the realism of audio tracking.
[0105] In some embodiments, such as Figure 6 As shown, when acquiring the target sub-audio in the audio to be played, the following steps may be included: Step 601: Based on the audio to be played, generate sub-audio using an audio separation model.
[0106] In this embodiment of the application, the audio separation model refers to a pre-trained neural network model used for audio separation.
[0107] In some embodiments, a large amount of mixed audio is pre-acquired, each mixed audio comprising audio signals from multiple different sound sources, and audio signals from individual sound sources within each mixed audio are also acquired, with corresponding labels (e.g., human voices or birdsong) added to these individual sound source audio signals. When training the audio separation model, the mixed audio is used as input to obtain predicted sub-audio, which includes predicted audio signals from multiple different sound sources. Then, a loss value is calculated between each predicted sub-audio and its corresponding label, and finally, the parameters of the audio separation model are updated based on the calculated loss values.
[0108] For example, the audio separation model is Hybrid Demucs or Spleeter.
[0109] In the embodiments of this application, sub-audio refers to an audio signal extracted from the audio to be played, which has only one sound source.
[0110] In some embodiments, the audio to be played can be input into a trained audio separation model to obtain multiple independent sub-audio files with a single sound source.
[0111] For example, the sub-audio includes lead vocals, narrator vocals, backing vocals, bird sounds, and rain sounds.
[0112] Step 602: Determine the target sub-audio from the sub-audio based on the type of the sub-audio and the preset type priority.
[0113] In this embodiment of the application, the type of sub-audio refers to the audio category to which the sub-audio belongs.
[0114] For example, the types of sub-audio include: lead vocals, narrator voices, ambient sounds, and alarm sounds.
[0115] In some embodiments, acoustic features are extracted for each sub-audio, and the type is determined based on the extracted acoustic features.
[0116] In this embodiment of the application, the preset type priority refers to the preset importance of different types of sub-audio.
[0117] For example, the preset type priority includes: lead vocals > narrator vocals > ambient sounds > alarm sounds.
[0118] In some embodiments, users may prioritize different types of audio depending on the audio to be played. For example, when the audio is pop music, users may focus on the lead vocals, followed by backing vocals; when the audio is a movie clip, users may focus on the dialogue with the main character, followed by narration. Therefore, it is necessary to determine the audio that users most want to hear based on their different levels of importance.
[0119] In some embodiments, after determining the type of each sub-audio, the priority of each sub-audio is determined by a preset type priority, and the sub-audio with the highest priority is selected as the target sub-audio.
[0120] In this embodiment, the target sub-audio is filtered based on the preset type priority and the type of sub-audio, which can determine the audio that best meets the user's needs and automatically switch to the next target sub-audio after the current target sub-audio has finished playing.
[0121] In some embodiments, such as Figure 7 As shown in the embodiments of this application, the audio playback method further includes the following steps: Step 701: Determine the type of gesture based on the user's gesture changes within a preset time period.
[0122] In this embodiment of the application, gesture change refers to the user's hand posture change information and movement trajectory information within a preset time period.
[0123] In this embodiment of the application, the gesture action type refers to the type of hand gestures performed within a continuous time period.
[0124] For example, the gesture type can be grabbing and dragging, waving, raising the hand, etc.
[0125] In some embodiments, considering that adjusting the audio playback position solely based on the user's real-time location lacks a means for the user to actively adjust the audio playback position, and since gestures are actively initiated by the user and are the most intuitive, the audio playback position is also adjusted using the user's gestures. Based on this, within a set time period, the continuously collected positions and postures of the user's hands are used as gesture changes, and then the corresponding gesture action type is determined using a gesture type recognition model.
[0126] Step 702: If the gesture action type matches the preset action type, update the real-time playback position of the target controlled sub-audio based on the user's hand termination position within a preset time period to obtain the updated real-time playback position; the target controlled sub-audio is the audio in the sub-audio.
[0127] In this embodiment of the application, the preset action type refers to a preset, valid action type that allows manipulation of audio.
[0128] For example, the preset action type could be grab and drag, swing, etc.
[0129] In some embodiments, to avoid accidental control, it is necessary to further identify the type of the user's gestures. The position will only be updated based on the user's gestures when a valid action type is identified. Based on this, a type table including preset action types is pre-stored. After obtaining a gesture action type, it is determined whether the gesture action type exists in the type table. If it exists, the gesture action type is determined to conform to the preset action type; if it does not exist, the gesture action type is determined to not conform to the preset action type.
[0130] In this embodiment of the application, the hand termination position refers to the position of the hand at the end of a preset time period.
[0131] In the embodiments of this application, the target controlled sub-audio refers to the audio selected from the sub-audio and controlled by the user's gesture.
[0132] In some embodiments, the target sub-audio is determined by the user's gesture.
[0133] For example, when the user's gesture is to grab and drag, the target sub-audio is the sub-audio that the user is grabbing; when the user's gesture is to wave, the target sub-audio is the sub-audio that the finger is facing.
[0134] In some embodiments, it is considered that different types of user gestures represent corresponding behavioral intentions. For example, when the gesture type is grasping and dragging, the playback position of the grasped audio falls precisely at the end of the drag; when the gesture type is waving, the playback position of the waved audio falls near the end of the wave. Based on this, corresponding coordinate update rules are set according to different action types. After determining that the gesture type conforms to the preset action type, the real-time playback position of the target sub-audio is updated based on the coordinate update rules and the hand termination position.
[0135] In some embodiments, considering that the user may control the target sub-audio, and still wanting the playback orientation and beam pointing angle of the target sub-audio to remain unchanged relative to the user even when the user moves, when the target controlled sub-audio is the target sub-audio, the real-time playback position of the target sub-audio is updated based on the coordinate update rules and the hand termination position, while the preset requirements are also updated.
[0136] For example, the preset requirement is that the target sub-audio is 0.5m directly in front of the user and pointing towards the user's face; the user makes a gesture to grab the target sub-audio 0.5m directly in front of them, and simultaneously rotates it 90° to the left and drags it to a position of 0.6m. The position of 0.6m is used as the real-time playback position of the target sub-audio, and the preset requirement is updated based on the position of 0.6m. The updated preset requirement becomes 0.6m directly to the left of the user and pointing towards the user's face.
[0137] Step 703: During the playback of the audio to be played, control the speaker array to play sub-audio based on the updated real-time playback position.
[0138] In some embodiments, based on the updated real-time playback orientation, spatial audio rendering technology is used to calculate the signal amplitude of the target controlled sub-audio that each speaker in the speaker array should output and to play the target controlled sub-audio.
[0139] In summary, the audio playback method provided in this application first updates the real-time playback orientation of the target sub-audio based on the collected real-time orientation of the user. Since the updated real-time playback orientation maintains the target sub-audio's playback position and beam pointing angle relative to the user within preset requirements, the relative playback position and beam pointing angle of the target sub-audio remain unchanged during the user's movement, thus ensuring that the target sub-audio is always played towards the user. By updating the user's real-time orientation angle using a smoothing coefficient determined based on the user's real-time movement speed, the real-time beam pointing angle of the target sub-audio changes smoothly, resulting in stable playback. By predicting the user's position and using the predicted position to update the real-time playback position of the target sub-audio, sound lag caused by data processing delays can be avoided, thereby improving the user's auditory experience. Adjusting the gain of the speaker array based on the user's real-time position ensures that the volume of the target sub-audio heard by the user remains consistent during movement.
[0140] Based on the above embodiments, another embodiment of this application proposes an audio playback method, including a multi-scenario application method based on AI (Artificial Intelligence) panoramic sound technology.
[0141] For example, the following is an exemplary description of possible implementations of the multi-scenario application method based on AI panoramic sound technology proposed in the embodiments of this application.
[0142] With the rapid development of audio technology, immersive sound (e.g., Dolby Atmos, DTS:X, MPEG-H 3D Audio) technology has gradually moved from professional cinemas to home entertainment and personal consumer electronics. Unlike traditional channel-based stereo or surround sound, immersive sound uses an object-based audio model, treating sound elements (e.g., vocals, instruments, sound effects) as independent "audio objects" and assigning them precise coordinates (X, Y, Z) in three-dimensional space. During playback, the audio renderer calculates in real time how the sound of each object should be distributed based on the actual speaker layout, thereby creating an extremely accurate and immersive three-dimensional sound field. Currently, some technologies and products on the market have emerged that overmix traditional stereo audio into immersive sound formats. The core AI immersive sound technology upon which these technologies are based typically uses deep learning models to separate the audio tracks, identify different sound elements, and assign them a reasonable static spatial position based on algorithms or training data. For example, vocals can be fixed in the center, drums placed in the back, and harmonies distributed on both sides. This technology has been initially applied in in-vehicle audio systems, providing passengers with a more immersive and vivid auditory experience than traditional stereo through a fixed speaker layout within the cabin.
[0143] However, the relevant technologies, including their current applications in the automotive environment, have the following significant limitations: 1. The contradiction between static sound fields and dynamic environments: The panoramic sound fields produced by upmixing techniques in related technologies are usually static. That is, once the position of the sound object is determined by the algorithm, it does not change during audio playback. However, in many real-world scenarios, the user, environment, or content itself is dynamically changing. For example, in smart homes, when a user moves around a room, a static sound field can cause the optimal listening position to disappear, resulting in a degraded experience. In AR (Augmented Reality) scenarios, the sound of virtual objects cannot remain stable as the user's head moves, which severely undermines immersion.
[0144] 2. Lack of interaction with the environment and users: Related technologies primarily focus on the conversion and playback of audio itself, failing to correlate and interact with sound fields in relation to physical space information, user status (e.g., location, orientation, posture, heart rate), or external digital events. Sound, as a powerful sensory dimension, has not had its interactive potential fully explored. For example, it cannot achieve intelligent interactions such as "when a user looks at a lamp, the volume of the ambient sound effects emitted by the lamp automatically increases."
[0145] 3. Limited Application Areas: Currently, the commercial application of this technology is mainly concentrated in high-end home theaters and in-vehicle entertainment systems, with its design philosophy essentially focused on "serving music listening / movie watching." For broader fields such as healthcare, education and training, and retail marketing, there is a lack of targeted, scenario-integrated interactive audio solutions. Therefore, there is an urgent need for an interactive panoramic sound system and method that can break through the limitations of static sound fields and achieve intelligent linkage and dynamic mapping between sound fields and space, users, and content.
[0146] In order to overcome the shortcomings of AI panoramic sound technology in related technologies, such as static sound field, lack of interaction, and narrow application fields, this application provides a dynamic sound field mapping and interaction system and method based on AI panoramic sound that is highly versatile, immersive, and interactive.
[0147] This application upgrades AI panoramic sound technology from a mere audio post-processing tool into an "intelligent sensory hub" connecting the physical world, the digital world, and the user themselves. Specifically, this application introduces an "environmental perception layer," a "user perception layer," and an "interaction logic layer" on top of the core AI panoramic sound conversion engine.
[0148] This application provides a multi-scene interaction method based on AI panoramic sound technology, which mainly includes the following steps: S1, Audio Objectification.
[0149] Utilizing core AI panoramic sound technology, an audio separation model is used to separate and convert the input ordinary stereo audio stream into multiple independent audio objects (sub-audio) with metadata (e.g., sound source type, loudness, spectral characteristics) in real time. Simultaneously, initial spatial coordinates are assigned to these independent audio objects; these initial spatial coordinates can be randomly assigned or assigned based on the type of the independent audio object. For example, the ordinary stereo audio stream can be music, podcasts, white noise, etc.
[0150] S2, Environment and User Perception.
[0151] A single integrated sensing module captures real-time spatial information of the physical environment and user status information. This integrated sensing module may include a camera, depth sensor, LiDAR (Light Detection and Ranging), UWB (Ultra Wide Band) positioning base station, and infrared sensor. For example, the spatial information of the physical environment may include one or more of the following: the three-dimensional structure of the room, furniture layout, and specific object markers; the user status information may include one or more of the following: the real-time three-dimensional coordinates of the user's head within the room, head orientation angle, gaze focus, gestures, and physiological data (e.g., heart rate).
[0152] S3, Dynamic Sound Field Mapping.
[0153] According to the preset mapping strategy, the spatial coordinates of the independent audio objects generated in step S1 are associated with the spatial information of the physical environment / user state information obtained in step S2. The association process is constantly changing.
[0154] Mapping strategies include: First mapping strategy: Real-time triggered location binding.
[0155] Anchoring a standalone audio object (e.g., birdsong) to a target physical entity (e.g., greenery) in the physical environment makes the standalone audio object appear to play on that target object, regardless of how the user moves.
[0156] The implementation process includes: S311, Physical environment calibration.
[0157] By using lidar, vision sensors, millimeter-wave radar, or sound field modeling, a three-dimensional coordinate system is established inside the vehicle or indoors to record the three-dimensional coordinates, orientation, identification, and effective sound range of the target physical entity inside the vehicle or indoors.
[0158] S312, Audio object anchoring.
[0159] Assign binding identifiers and target physical coordinates to individual audio objects (e.g., ambient sounds, sound effects, cue sounds, etc.). Through a pre-established mapping relationship between the identifiers of individual audio objects and target physical entities, after acquiring an audio object, the corresponding identifier is obtained through this mapping relationship. Then, the initial spatial coordinates of the individual audio object are modified to the three-dimensional coordinates of the target physical entity corresponding to that identifier, and the sound source radiation characteristics of the individual audio object are set. These sound source radiation characteristics include: directivity, diffusion angle, height, and attenuation curves at different distances.
[0160] S313, real-time spatial alignment.
[0161] Sound field rendering is performed at a fixed frequency (e.g., 30~60Hz). Regardless of how the user moves or turns, the spatial coordinates of individual audio objects remain locked to their corresponding target physical entities.
[0162] S314, rendering output.
[0163] Based on the target physical coordinates of the anchored independent audio object, audio with spatial positioning is output through HRTF (Head Related Transfer Function) or multi-channel rendering technology.
[0164] Second mapping strategy: User follow.
[0165] When a target user is present in a vehicle or indoors, user following is triggered. This means the core audio object (target sub-audio) is set to always follow the target user. If multiple users are present, the target user can be any user or a user in a preset location (e.g., the driver's seat). If the independent audio object includes the lead vocal, that vocal is the core audio object. If the opposing audio object does not include the lead vocal, any independent audio object can be used as the core audio object. When the second mapping strategy is triggered, the core audio object always faces the target user as they move, creating the effect of a personal mobile stage.
[0166] The implementation process includes: S321, Real-time Header Data Acquisition and Preprocessing.
[0167] Using a depth camera, millimeter-wave radar, UWB positioning base station, and head posture calculation algorithm, the system acquires the user's three-dimensional head coordinates (real-time position), real-time orientation angles (including horizontal orientation angle θ and vertical orientation angle φ), real-time movement velocity v, and real-time movement direction angle α in real time at a fixed frequency (e.g., 30~60Hz). This data is low-pass filtered and outlier removed to output stable data, and position prediction is performed based on the current motion parameters to compensate for data processing latency. Specifically, the real-time acquired three-dimensional head coordinates and orientation angle constitute the user's real-time orientation; the movement direction angle refers to the angle of the user's head's direction of movement on the horizontal plane. Position prediction refers to predicting the head position at the next moment based on the user's three-dimensional head coordinates, real-time movement direction angle, real-time movement velocity, and the system's fixed latency (data processing latency).
[0168] S322, Establish a local positive coordinate system for the head center.
[0169] A local relative coordinate system is constructed with the user's head as the dynamic origin. The playback position of the core audio object is fixed at a relative coordinate of a preset standard distance directly in front of the head, for example, the relative coordinate is (0, 0.5, 0), to ensure that the sound source of the core audio object is always in the user's auditory front area.
[0170] S323, real-time binding of global coordinates based on head position. Using the global spatial coordinate system constructed indoors / in the vehicle as a reference, calculate the global three-dimensional coordinates of the core audio object according to the above formulas (1), (2) and (3), obtain the updated global three-dimensional coordinates (updated real-time playback position), and realize the synchronous movement of the core audio object and the user's position.
[0171] S324, vocal direction alignment based on head orientation.
[0172] The sound source pointing angle (beam pointing angle) of the core audio object is synchronized in real time with the real-time horizontal orientation angle θ of the head, ensuring that the sound source's emitting surface is always directly facing the user's ears. A first-order inertial filter is used to achieve a smooth transition of the horizontal orientation angle, avoiding abrupt changes in the sound field. When the core audio object is played facing the user, the sound source pointing angle takes the opposite value of the real-time horizontal orientation angle θ; when the audio object is played facing away from the user, the sound source pointing angle is consistent with the real-time horizontal orientation angle θ. The formula for the first-order inertial filter can be: (4) Where k is the smoothing coefficient and t is the current time.
[0173] S325, based on real-time motion speed and direction tracking smooth compensation.
[0174] The smoothing coefficient k in the first-order inertial filter is adapted according to the user's real-time movement speed. Different levels of smoothing coefficient k are assigned based on the user's real-time movement speed. At low speeds, the smoothing coefficient k is reduced to improve positioning response speed, while at high speeds, the smoothing coefficient k is increased to suppress sound field jitter, thus ensuring that the movement direction of the core audio object is consistent with the movement direction angle α. Combined with the predicted position from step S321, the real-time spatial coordinates of the user's head are updated to achieve advanced tracking and eliminate sound lag.
[0175] S326, distance and loudness adaptive calibration.
[0176] The system calculates the straight-line distance between the user's head and each speaker in real time (real-time distance) and dynamically adjusts the speaker gain based on the acoustic distance attenuation model (distance attenuation model). The greater the distance, the higher the gain, and the closer the distance, the lower the gain, ensuring consistent loudness and balanced timbre for the core audio.
[0177] S327, real-time panoramic sound rendering output.
[0178] The corrected coordinates of the core audio object and the sound source pointing angle are input into the immersive sound rendering engine. Through head-related transformation functions or vector base amplitude panning (VBAP) algorithms, combined with the layout of in-vehicle / indoor speakers, multi-channel rendering is completed, driving the speaker array to play back, so that the core audio object is always facing the user.
[0179] In this embodiment, the system has various pre-built or downloaded interactive logics. When the user activates the interactive function, the system enters a real-time monitoring state: sensor modules such as eye tracking, depth cameras, millimeter-wave radar, human presence sensors, and spatial positioning continuously collect user status, location, gaze, gestures, and environmental information. Once the sensor data meets a specific trigger condition defined in a rule, the system will immediately match and invoke the corresponding interactive logic without manual intervention. During the execution phase, the system dynamically modifies the attributes and behaviors of spatial audio objects based on the invoked interactive logic, including but not limited to: sound source location, volume gain, sound field range, playback status, channel weights, spatial orientation, etc., and can also link with IoT devices such as lights and displays to output feedback synchronously, ultimately achieving a natural, seamless, and immersive spatial audio interactive experience.
[0180] In some embodiments, the interaction logic includes: Interaction Logic 1: When the user's gaze (determined by eye tracking) is focused on a physical object bound to an audio object for more than a set time threshold (e.g., 2 seconds), the volume of the audio object corresponding to that physical object is increased (e.g., increased by 6dB), and it emits a faint light (if the physical object is a smart lamp). The focus time, volume increase value, and light emission level can be defined by the system user.
[0181] Interaction Logic 2: When the user makes a "grab and drag" gesture, the user is allowed to "grab" an audio object (target controlled sub-audio) from one location in space and "place" it to another location.
[0182] Interaction Logic 3: When the sensor detects that the user has entered room B from room A, the system automatically and smoothly transitions the sound field of the music being played from room A to room B, creating a "music follow" experience.
[0183] The beneficial effects of this application include: a) Breaking the limitations of optimal listening position: Through user-following mode, accurate reproduction of the core sound image is achieved from any location, greatly enhancing the auditory experience while on the move. b) Creating a deeply immersive AR audio experience: Seamlessly anchoring virtual sound into the real world provides a lower-cost and more immersive audio solution for AR applications. c) Improving the efficiency and fun of information delivery: In museums, exhibition halls, or retail stores, each exhibit is equipped with an independent spatial audio narration that plays automatically when users approach or look at it, avoiding the clutter of loudspeakers. d) Providing a completely new form of entertainment interaction: Turning sound into an interactive, tangible, and movable entity opens up new application possibilities such as "sound toys" and "spatial music creation."
[0184] In some embodiments, such as Figure 8 As shown, the AI panoramic sound dynamic sound field mapping and interaction system provided in this application includes the following core modules: audio input and objectification conversion module 801, environment and user perception module 802, interaction logic and control module 803, sound field rendering and synthesis module 804, and audio output module 805.
[0185] In this embodiment, the audio input and objectification conversion module 801 receives mixed audio objects (e.g., ordinary stereo or multi-channel audio signals) from various sound sources (such as local storage, network streaming media, and Bluetooth transmission). The built-in AI panoramic sound conversion engine (audio separation model) analyzes and processes the input signal, using, for example, source separation models based on deep neural networks, such as Demucs, and spatial location allocation models. This engine first separates the mixed audio signal into several individual sound sources (such as vocals, bass, drums, ambient sounds, etc.), and then generates an independent audio object for each sound source. Each audio object contains its own waveform data and metadata. The metadata includes at least the audio object's identifier, a sound source type label, and an initial three-dimensional spatial coordinate (X, Y, Z), which can be preset based on the audio type or be random.
[0186] In this embodiment, the environment and user perception module 802 is responsible for constructing a dynamic digital model of the physical space and the user, including: an environment perception submodule: integrating multiple sensors, such as wide-angle cameras, depth cameras (e.g., Intel RealSense), LiDAR, or millimeter-wave radar. Through simultaneous localization and mapping (SLAM) technology, it generates a 3D point cloud map of the environment in real time, identifying and marking the spatial locations and boundaries of key objects (e.g., sofas, tables, windows, specific markers).
[0187] User perception submodule: Also using the above sensors, and combined with specific algorithms (such as skeletal key point detection, face recognition, eye tracking algorithm), it tracks the user's head in space in real time in three-dimensional coordinates (Ux, Uy, Uz), head orientation angle (θ_head), and gaze direction (Φ_gaze).
[0188] In this embodiment, the interaction logic and control module 803 serves as the system's brain, comprising a strategy library and a logic executor. The strategy library stores various sound field mapping strategies and interaction rules. These strategies and rules can be customized by users or developers through a graphical interface. Examples of mapping strategies include: "Lead vocals follow the user," "Drum sounds are fixed at spatial coordinates (0, 0, -2)," and "Ambient sound effects are randomly distributed at ceiling height." Interaction rules include specific rules such as "eye focus trigger," "proximity trigger," and "gesture trigger." The logic executor continuously receives data from the user perception submodule and determines whether the trigger conditions of the preset rules in the strategy library are met. Once met, it sends a control command to the sound field rendering and synthesis module 804. Proximity trigger refers to automatically playing audio (e.g., an introduction) set on a physical object when the user is within a set distance of that object.
[0189] In this embodiment, the sound field rendering and synthesis module 804 receives an audio object stream from the audio input and objectification conversion module 801, and control commands from the interaction logic and control module 803. Its core is a dynamic panoramic sound renderer (such as a renderer based on VBAP vector basis amplitude panning or Ambisonics). The panoramic sound renderer updates the target spatial coordinates of one or more audio objects in real time and smoothly according to the control commands. For example, when the command requires "object A to follow the user," the renderer will set the coordinates of object A to (Ux, Uy+0.5, Uz) (i.e., 0.5 meters in front of the user) in each frame calculation. Finally, based on the latest coordinates of all audio objects and the actual layout of the current speakers, the renderer determines the signal allocation weights using the VBAP / Ambisonics positioning algorithm, combines distance attenuation and acoustic simulation to calculate the sound signal that each speaker should emit, and outputs it.
[0190] In this embodiment, the audio output module 805 uses a power amplifier to drive a group of speaker arrays (such as distributed speakers in smart homes or multi-speaker systems in vehicles) to play back the multi-channel audio signal output by the sound field rendering and synthesis module 804, thereby achieving the final immersive and interactive sound field.
[0191] In some embodiments, the overall workflow of the system is as follows: Figure 9 As shown, it includes the following steps: Step 901, audio object conversion.
[0192] The audio input and object conversion module 801 continuously converts the input stereo sound into an audio object set.
[0193] Step 902, Real-time Environment and User Perception.
[0194] The Environment and User Perception Module 802 continuously monitors the environment and users, and updates the space and user status model.
[0195] Step 903, interaction logic matching.
[0196] The interaction logic and control module 803 matches the latest environment / user data with preset strategies.
[0197] Step 904: Determine whether the interaction logic is triggered.
[0198] The interaction logic and control module 803 determines whether the interaction logic is triggered. If it is triggered, proceed to step 905; if it is not triggered, proceed to step 906.
[0199] Step 905: Generate sound field control commands.
[0200] The interaction logic and control module 803 generates audio object coordinate update instructions or other attribute (such as volume, filtering) modification instructions based on the triggered interaction logic.
[0201] Step 906, Dynamic sound field rendering.
[0202] The sound field rendering and compositing module 804 performs real-time panoramic sound rendering based on the latest control commands and audio object data. Rendering continues regardless of whether the coordinates are updated.
[0203] Step 907, audio output.
[0204] The audio output module 805 plays the rendered audio signal through the speaker.
[0205] Then, step 901 continues to be executed, forming a continuous, dynamic interactive loop until the audio playback ends or the system shuts down.
[0206] Based on the above embodiments, Figure 10 This is a schematic diagram of an apparatus provided in an embodiment of this application, such as... Figure 10 As shown, the device 1000 includes: an acquisition unit 1001, a data collection unit 1002, an update unit 1003, and a control unit 1004. Wherein: Acquisition unit 1001 is used to acquire the target sub-audio in the audio to be played.
[0207] The acquisition unit 1002 is used to acquire the user's real-time location in the current area.
[0208] The update unit 1003 is used to update the real-time playback orientation of the target sub-audio based on the real-time orientation, so as to obtain the updated real-time playback orientation; the updated real-time playback orientation controls the playback position and beam pointing angle of the target sub-audio relative to the user to maintain preset requirements.
[0209] The control unit 1004, during the playback of the audio to be played, controls the speaker array set in the current area to play the target sub-audio based on the updated real-time playback orientation.
[0210] The descriptions of the above device embodiments are similar to those of the above method embodiments, and have similar beneficial effects. For technical details not disclosed in the device embodiments of this application, please refer to the descriptions of the method embodiments of this application for understanding.
[0211] It should be noted that the module division in the embodiments of this application is illustrative and only represents one logical functional division. In actual implementation, there may be other division methods. Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, exist as separate physical units, or have two or more units integrated into one unit. The integrated units can be implemented in hardware, as software functional units, or a combination of software and hardware.
[0212] It should be noted that, in the embodiments of this application, if the above methods are implemented as software functional modules and sold or used as independent products, they can also be stored in a computer-readable storage medium. Based on this understanding, the technical solutions of the embodiments of this application, or the parts that contribute to related technologies, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause an electronic device to execute all or part of the methods of the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), magnetic disks, or optical disks. Thus, the embodiments of this application are not limited to any specific hardware and software combination.
[0213] This application provides an electronic device. Figure 11 This is a schematic diagram of the structure of the electronic device provided in the embodiments of this application, such as... Figure 11 As shown, the electronic device 1100 includes a memory 1101 and a processor 1102. The memory 1101 stores a computer program that can run on the processor 1102. When the processor 1102 executes the program, it implements the steps in the method provided in the above embodiments.
[0214] It should be noted that the memory 1101 is configured to store instructions and applications executable by the processor 1102, and can also cache data to be processed or already processed (e.g., image data, audio data, voice communication data and video communication data) in the processor 1102 and various modules in the electronic device 1100. It can be implemented by flash memory or random access memory (RAM).
[0215] This application provides a computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements the steps of the method provided in the above embodiments.
[0216] This application provides a computer program product containing instructions that, when run on a computer, cause the computer to perform the steps in the method provided in the above-described method embodiments.
[0217] It should be noted that the descriptions of the above storage medium and electronic device embodiments are similar to the descriptions of the above method embodiments, and have similar beneficial effects. For technical details not disclosed in the storage medium and electronic device embodiments of this application, please refer to the descriptions of the method embodiments of this application for understanding.
[0218] It should be understood that the phrases "one embodiment," "an embodiment," or "some embodiments" mentioned throughout the specification mean that a specific feature, structure, or characteristic related to an embodiment is included in at least one embodiment of this application. Therefore, "in one embodiment," "in one embodiment," or "in some embodiments" appearing throughout the specification do not necessarily refer to the same embodiment. Furthermore, these specific features, structures, or characteristics can be combined in any suitable manner in one or more embodiments. It should be understood that in the various embodiments of this application, the sequence numbers of the above-described processes do not imply a sequential order of execution; the execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application. The sequence numbers of the above-described embodiments are merely for descriptive purposes and do not represent the superiority or inferiority of the embodiments. The descriptions of the various embodiments above tend to emphasize the differences between the various embodiments; their similarities or commonalities can be referred to mutually, and for the sake of brevity, they will not be repeated here.
[0219] In this article, the term "and / or" is merely a description of the relationship between related objects, indicating that there can be three kinds of relationships. For example, object A and / or object B can represent three situations: object A exists alone, object A and object B exist simultaneously, and object B exists alone.
[0220] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or electronic device. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element.
[0221] In the several embodiments provided in this application, it should be understood that the disclosed electronic devices and methods can be implemented in other ways. The embodiments described above are merely illustrative. For example, the division of modules is only a logical functional division, and in actual implementation, there may be other division methods, such as: multiple modules or components can be combined, or integrated into another system, or some features can be ignored or not executed. In addition, the coupling, direct coupling, or communication connection between the various components shown or discussed can be through some interfaces, and the indirect coupling or communication connection of electronic devices or modules can be electrical, mechanical, or other forms.
[0222] The modules described above as separate components may or may not be physically separate. The components shown as modules may or may not be physical modules. They may be located in one place or distributed across multiple network units. Some or all of the modules may be selected to achieve the purpose of this embodiment according to actual needs.
[0223] In addition, each functional module in the various embodiments of this application can be integrated into one processing unit, or each module can be a separate unit, or two or more modules can be integrated into one unit; the integrated modules can be implemented in hardware or in the form of hardware plus software functional units.
[0224] Those skilled in the art will understand that all or part of the steps of the above method embodiments can be implemented by hardware related to program instructions. The aforementioned program can be stored in a computer-readable storage medium. When the program is executed, it performs the steps of the above method embodiments. The aforementioned storage medium includes various media that can store program code, such as mobile storage devices, read-only memory (ROM), magnetic disks, or optical disks.
[0225] Alternatively, if the integrated units described above are implemented as software functional modules and sold or used as independent products, they can also be stored in a computer-readable storage medium. Based on this understanding, the technical solutions of the embodiments of this application, or the parts that contribute to related technologies, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause an electronic device to execute all or part of the methods of the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as mobile storage devices, ROMs, magnetic disks, or optical disks.
[0226] The methods disclosed in the several method embodiments provided in this application can be arbitrarily combined without conflict to obtain new method embodiments.
[0227] The features disclosed in the several product embodiments provided in this application can be arbitrarily combined without conflict to obtain new product embodiments.
[0228] The features disclosed in the several method or electronic device embodiments provided in this application can be arbitrarily combined without conflict to obtain new method or electronic device embodiments.
[0229] The above description is merely an embodiment of this application, but the scope of protection of this application is not limited thereto. Any changes or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application.
[0230] It should be understood that if this disclosure references any user data and personal information (including but not limited to device information, behavioral data, location information, etc.) and before applying the technical solutions described in the embodiments of this disclosure, the relevant products or services should comply with the laws and regulations concerning the protection of user data and personal information, strictly process users' personal information and data in accordance with the provisions of applicable laws and regulations throughout the entire data processing lifecycle, follow the principles of legality, legitimacy, necessity, good faith, openness, and transparency, and adopt reasonable privacy design schemes and technical measures to ensure the security of user data and personal information, protect users' legitimate rights and interests, and prevent the risks of leakage, theft, or tampering of user data and personal information.
[0231] Specifically, the company must publish and display its privacy policy in a prominent position on the user interface, clearly informing users of the types, purposes, uses, and methods of processing personal information, as well as other matters that should be disclosed as required by laws and regulations; obtain users' prior informed consent or explicit authorization regarding data processing through user-initiated interaction (such as confirmation pop-ups); process or store user data securely within the legally required timeframe; adopt a series of security technologies and management measures, including but not limited to data encryption and access control; share and transfer user data within the scope permitted by law and in a legally required manner; and process user rights, including the rights to query, access, correct, delete, withdraw authorization and consent, cancel registration, and obtain copies of personal information, within the legally required timeframe.
Claims
1. An audio playback method, characterized by, The method includes: Obtain the target sub-audio from the audio to be played, and collect the user's real-time location in the current area; Based on the real-time orientation, the real-time playback orientation of the target sub-audio is updated to obtain the updated real-time playback orientation; the updated real-time playback orientation controls the playback position and sound beam pointing angle of the target sub-audio relative to the user to maintain preset requirements; During the playback of the audio to be played, based on the updated real-time playback orientation, the speaker array set in the current area is controlled to play the target sub-audio.
2. The method of claim 1, wherein, The real-time orientation includes the real-time orientation angle; the real-time playback orientation includes the real-time sound beam pointing angle. The step of updating the real-time playback orientation of the target sub-audio based on the real-time orientation to obtain the updated real-time playback orientation includes: The real-time orientation angle is smoothed to obtain the processed orientation angle; Based on the positional relationship between the sound beam pointing angle and the user, as well as the processed orientation angle, the real-time sound beam pointing angle is updated to obtain the updated real-time playback orientation.
3. The method of claim 2, wherein, The method further includes: Collect the user's real-time movement speed in the current area; The step of smoothing the real-time orientation angle to obtain the processed orientation angle includes: A smoothing coefficient is determined based on the real-time motion speed; the smoothing coefficient is positively correlated with the real-time motion speed. The real-time orientation angle is smoothed based on the smoothing coefficient to obtain the processed orientation angle.
4. The method of claim 1, wherein, The real-time orientation includes the real-time location; the real-time playback orientation includes the real-time playback location; the method further includes: Collect the user's real-time movement speed and real-time movement direction angle in the current area; The step of updating the real-time playback orientation of the target sub-audio based on the real-time orientation to obtain the updated real-time playback orientation includes: Based on the real-time location, the real-time movement speed, the real-time movement direction angle, and the data processing delay, position prediction is performed to obtain the predicted location of the user at the next moment. Based on the predicted position, the real-time playback position is updated to obtain the updated real-time playback orientation.
5. The method of claim 4, wherein, The method further includes: Based on the real-time location, calculate the real-time distance from the user to the speaker array; During the playback of the audio to be played, based on the updated real-time playback orientation, controlling the speaker array located in the current area to play the target sub-audio includes: Based on the real-time distance, the gain of the speaker array is adjusted using a distance attenuation model; During the playback of the audio to be played, based on the updated real-time playback orientation, the speaker array located in the current area is controlled to play the target sub-audio based on the adjusted gain.
6. The method according to any one of claims 1 to 5, characterized in that, The step of obtaining the target sub-audio in the audio to be played includes: Based on the audio to be played, sub-audio is generated using an audio separation model; The target sub-audio is determined from the sub-audio based on the type of the sub-audio and the preset type priority.
7. The method according to claim 6, characterized in that, The method further includes: The type of gesture action is determined based on the user's gesture changes within a preset time period; If the gesture type matches a preset gesture type, the real-time playback position of the target controlled sub-audio is updated based on the user's hand termination position within the preset time period to obtain the updated real-time playback position; the target controlled sub-audio is the audio within the sub-audio. During the playback of the audio to be played, the speaker array is controlled to play the target controlled sub-audio based on the updated real-time playback position.
8. An audio playback device, characterized in that, The device includes: The acquisition unit is used to acquire the target sub-audio in the audio to be played. The acquisition unit is used to acquire the user's real-time location in the current area; The update unit is used to update the real-time playback orientation of the target sub-audio based on the real-time orientation, so as to obtain the updated real-time playback orientation; the updated real-time playback orientation controls the playback position and sound beam pointing angle of the target sub-audio relative to the user to maintain preset requirements; The control unit, during the playback of the audio to be played, controls the speaker array set in the current area to play the target sub-audio based on the updated real-time playback orientation.
9. An electronic device comprising a memory and a processor, the memory storing a computer program executable on the processor, characterized in that, When the processor executes the program, it implements the method described in any one of claims 1 to 7.
10. A computer program product, comprising a computer program or instructions, characterized in that, When the computer program or instructions are executed by a processor, they implement the method described in any one of claims 1 to 7.