Audio processing method, audio playback device, and computer-readable storage medium

The audio processing method dynamically adjusts virtual speaker positions and angles based on user motion to enhance spatial audio effects, addressing the inflexibility of current systems and improving the audio experience.

JP7760571B2Active Publication Date: 2025-10-27ANKER INNOVATIONS TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
JP2023184880
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2022-10-27
Filing Date
2023-10-27
Publication Date
2025-10-27
Estimated Expiration
2043-10-27

AI Technical Summary

Technical Problem

Current sound effect processing systems fail to adjust spatial audio effects flexibly according to user movements, resulting in an in-head sound effect experience.

Method used

An audio processing method that acquires motion information from a user, calculates virtual speaker positions and angles based on this information, and processes audio data to create spatial audio data dynamically, adjusting virtual speaker positions and angles in response to user movements.

Benefits of technology

This method enhances the sound effect tracking effect by dynamically adjusting spatial audio playback to match user movements, improving the realism and interaction of the audio experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007760571000002
    Figure 0007760571000002
  • Figure 0007760571000003
    Figure 0007760571000003
  • Figure 0007760571000004
    Figure 0007760571000004
Patent Text Reader

Abstract

To provide an audio processing method, an audio reproduction apparatus, and a computer readable storage medium for improving a sound effect following effect under a movement state of a user.SOLUTION: A method comprises: a step S11 of acquiring movement information on an audio reproduction apparatus moving with user's movement; a step S12 of calculating and acquiring position and angle information of each of at least two virtual speakers with respect to a user based on acquired movement trajectory of the user, real time movement speed, real time acceleration, and a preset sound effect function; a step S13 of acquiring audio data to be processed of the audio reproduction apparatus, and calculating and acquiring processed spatial audio data based on the preset sound effect function and the acquired position and angle information of each of at least two virtual speakers; and a step S14 of reproducing the spatial audio data by using the audio reproduction apparatus.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present application relates to the field of audio processing, and in particular to an audio processing method, an audio playback device and a computer-readable storage medium. [Background technology]

[0002] Signals processed by sound effect localization algorithms can simulate a variety of different spatial auditory effects. A virtual speaker is a virtual sound source processed by a sound effect function, and the position of the virtual speaker is the virtual sound source position processed by the sound effect function. Audio that is not processed by a sound effect function does not exhibit the spatial sound effect provided by a virtual speaker, but exhibits an in-head sound effect, i.e., the effect in which the listener feels that the audio is always being played inside the ear. Current sound effect processing cannot flexibly adjust according to the user's movements. Summary of the Invention [Problem to be solved by the invention]

[0003] The present application mainly provides an audio processing method, an audio playback device and a computer-readable storage medium, which solve the problem that the sound effect processing in the prior art cannot be flexibly adjusted according to the user's movements. [Means for solving the problem]

[0004] In order to solve the above technical problems, a first aspect of the present application provides an audio processing method, including the steps of: acquiring, by an audio playback device, motion information moving in accordance with the movement of a user, the motion information including at least a motion trajectory, a real-time motion speed, and a real-time acceleration of the user; calculating and acquiring position and angle information of at least two virtual speakers for the user based on the acquired motion trajectory, real-time motion speed, real-time acceleration of the user, and a preset sound effect function; acquiring audio data to be processed by the audio playback device, and calculating and acquiring processed spatial audio data based on the preset sound effect function and the acquired position and angle information of the at least two virtual speakers; and playing the spatial audio data using the audio playback device.

[0005] In order to solve the above technical problem, a second aspect of the present application provides an audio playback device, the audio playback device including a processor and a memory interconnected, a computer program stored in the memory, and the processor is used to execute the computer program to realize the steps of the audio processing method provided in the first aspect.

[0006] In order to solve the above technical problem, a third aspect of the present application provides a computer-readable storage medium having program data stored therein, the program data realizing the audio processing method provided by the first aspect when executed by a processor. [Effects of the Invention]

[0007] The advantageous effects of the present application are as follows: Unlike the prior art, the present application first acquires motion information of an audio playback device moving in accordance with a user's movement, the motion information including at least the user's motion trajectory, real-time motion speed, and real-time acceleration; then, based on the acquired user's motion trajectory, real-time motion speed, real-time acceleration, and a preset sound effect function, calculates and obtains position and angle information of at least two virtual speakers relative to the user; obtains audio data to be processed by the audio playback device; calculates and obtains processed spatial audio data based on the preset sound effect function and the acquired position and angle information of the at least two virtual speakers; and finally, plays the spatial audio data using the audio playback device. The above method calculates and obtains position and angle information of at least two virtual speakers using the motion information of the audio playback device moving in accordance with the user and the preset sound effect function; performs sound effect processing on the audio data to be processed by the audio playback device using the at least two virtual speakers to obtain spatial audio data; and realizes a spatial sound effect playback effect after playing the spatial audio data, thereby improving the sound effect tracking effect under a moving state.

[0008] In order to more clearly describe the technical solutions of the embodiments of the present application, the following will briefly describe the drawings that need to be used in the description of the embodiments. It is obvious that the drawings described below are only some embodiments of the present application, and those skilled in the art can obtain other drawings based on these drawings without any creative efforts. [Brief explanation of the drawings]

[0009] [Figure 1] 1 is a flowchart of an embodiment of an audio processing method according to the present application. [Figure 2] 1 is a schematic diagram illustrating a positional relationship between an audio playback device and a virtual speaker according to an embodiment of the present application; [Figure 3]FIG. 10 is a schematic diagram illustrating the positional relationship between an audio playback device and a virtual speaker according to another embodiment of the present invention. [Figure 4] FIG. 10 is a schematic diagram illustrating the positional relationship between an audio playback device and a virtual speaker according to yet another embodiment of the present application. [Figure 5] 1 is a schematic diagram illustrating a positional relationship between an audio playback device and a virtual speaker during an accelerated linear movement process according to an embodiment of the present application. FIG. [Figure 6] 1 is a schematic diagram illustrating a positional relationship between an audio playback device and a virtual speaker during a decelerating linear movement process according to an embodiment of the present application; [Figure 7] 1 is a flowchart of an embodiment of determining curve information according to the present application. [Figure 8] 2 is a schematic diagram of an embodiment of the movement direction of an audio playback device and road direction in a turning situation according to the present application; FIG. [Figure 9] 1 is a schematic diagram of an embodiment of an audio playback device orientation change in a bending situation according to the present application. FIG. [Figure 10] 1 is a schematic diagram illustrating the positional relationship between an audio playback device and a virtual speaker according to an embodiment of the present invention during an acceleration turn. [Figure 11] 1 is a schematic diagram illustrating the positional relationship between an audio playback device and a virtual speaker according to an embodiment of the present invention during a deceleration turn. [Figure 12] 1 is a schematic diagram of a positional relationship of an example of a user's head rotation according to the present application. FIG. [Figure 13] 10A and 10B are schematic diagrams illustrating positional relationships of another example of the user's head rotation according to the present application. [Figure 14] 1 is a structural schematic block diagram of an embodiment of an audio playback device according to the present application; [Figure 15] FIG. 2 is a structural schematic block diagram of another embodiment of an audio playback device according to the present application; [Figure 16] 1 is a structural schematic block diagram of an embodiment of a computer-readable storage medium according to the present application; DETAILED DESCRIPTION OF THE INVENTION

[0010] Hereinafter, the technical solutions of the embodiments of the present application will be described clearly and completely with reference to the drawings of the embodiments of the present application, but it should be clear that the described embodiments are only some of the embodiments of the present application, and not all of the embodiments, and all other embodiments obtained by those skilled in the art based on the embodiments of the present application without any creative efforts fall within the scope of protection of the present application.

[0011] The terms "first" and "second" used herein are merely descriptive and should not be understood as indicating or implying relative importance or the number of technical features being indicated. Thus, a feature qualified by "first" or "second" may explicitly or implicitly include at least one of the feature. In the present description, unless otherwise clearly and specifically limited, "plurality" means at least two, e.g., two, three, etc. Furthermore, the terms "comprise" and "have," and any variations thereof, are intended to cover non-exclusive inclusions. For example, a process, method, system, product, or apparatus comprising a series of steps or units is not limited to the listed steps or units, but may optionally further include unlisted steps or units, or may optionally further include other steps or units inherent to the process, method, product, or apparatus.

[0012] The term "embodiment" as used herein means that a particular feature, structure, or characteristic described with reference to the embodiment may be included in at least one embodiment of the present application. Appearances of the phrase in various places in the specification do not necessarily refer to the same embodiment, nor are they exclusive independent or alternative embodiments to other embodiments. Those skilled in the art will understand, both explicitly and implicitly, that the embodiments described herein can be combined with other embodiments.

[0013] Referring to Fig. 1, Fig. 1 is a flowchart of an embodiment of an audio processing method according to the present application. Note that this embodiment is not limited to the flow order shown in Fig. 1 as long as substantially the same results are obtained. This embodiment includes the following steps S11 to S14.

[0014] Step S11: Obtain motion information indicating that the audio playback device moves in accordance with the user's movements.

[0015] The audio playback device referred to in this specification includes, but is not limited to, wired earphones and wireless wearable devices such as, for example, wireless earphones (headsets, half in-ear earphones, in-ear earphones, etc.) and wireless audio glasses, and the audio playback device can establish a wired or wireless communication connection with the audio source device, thereby receiving audio data to be processed from the audio source device.

[0016] For example, the audio source device may be a mobile phone, a tablet computer, or a wearable audio source device such as a wristwatch or a band. The audio source device may store local audio data or obtain audio data via a network as audio data to be processed by an application program or a web page. The audio data to be processed may be, for example, audio data of music, audio data of e-books, audio from television / movies, etc.

[0017] The audio playback device moves in accordance with the user's movements, for example, in an exercise scene, the user wears the audio playback device, and the audio playback device is configured to move along with the user's movements.

[0018] In one embodiment, the motion information is acquired in real time using a positioning device and an acceleration sensor, at least one of which is provided in the audio playback device or in a smart mobile device communicatively connected to the audio playback device, such as a mobile phone, a smart wearable device such as a wristwatch, etc.

[0019] The positioning device uses radio frequency communication technology (such as UWB or Bluetooth technology) and GPS positioning technology to obtain information such as the user's angle, speed, acceleration, and trajectory, and then realizes spatial audio tracking under the scene. UWB (Ultra Wide Band) technology measures distance using the principle of TOF (Time of Flight). UWB is an ultra-wideband technology with advantages such as strong penetration, good multipath resistance, and accurate positioning accuracy, and is applicable to the positioning tracking and navigation of stationary or moving objects indoors.

[0020] The motion information includes at least the user's motion trajectory, real-time motion speed, and real-time acceleration, and more specifically, for example, whether there is acceleration or deceleration under the motion scene, acceleration under acceleration or deceleration state, turning information, etc.

[0021] Step S12: Based on the acquired user's movement trajectory, real-time movement speed, real-time acceleration, and preset sound effect function, calculate and obtain position and angle information of at least two virtual speakers for the user.

[0022] A virtual speaker is a virtual sound source processed by a sound effect function, and the position of the virtual speaker is a virtual sound source position processed by a sound effect function. Audio that has not been processed by a sound effect function does not represent the sound effect provided by the virtual speaker, but directly represents the original audio.

[0023] The sound effect functions referred to here are, for example, Head Related Transfer Functions (abbreviated as HRTF), also known as ATF (anatomical transfer function), which are personalized spatial sound effect algorithms.

[0024] Specifically, the head-related transfer function describes the transmission process of sound waves from a sound source to both ears, and it comprehensively takes into account factors such as the time difference in propagation of sound waves from the sound source to the two ears, the interaural level difference due to the shadow and scattering effect of the head on sound waves when the sound source is not on the perpendicular bisecting plane, the scattering and diffraction effect on sound waves by human physiological structures (e.g., head, pinna, and torso, etc.), the dynamic and psychological factors that cause localization confusion when the sound source is positioned above or below or in front or behind the mirror surface and on the perpendicular bisecting plane, etc. In practical applications, various different spatial auditory effects can be simulated by using earphones or speakers to retransmit signals processed by HRTFs.

[0025] The position information includes at least the distance in the horizontal direction between the audio reproduction device and the virtual speaker, and the angle information includes at least the angular relationship in the horizontal direction between the audio reproduction device and the virtual speaker.

[0026] For example, a head-related transfer function can be simply expressed as HRTF(L, θ1, θ2), where θ1 represents the horizontal angle parameter between the user and the virtual speaker, θ2 represents the depression / elevation angle between the audio playback device and the virtual speaker (i.e., the vertical angle between the audio playback device and the virtual speaker), and L is a distance parameter between the audio playback device and the virtual speaker, where L, θ1, and θ2 may be fixed values ​​or may be changed to different values ​​based on the motion position information and angle information of the virtual speaker relative to the user. Each virtual speaker can correspond to one head-related transfer function.

[0027] The angle parameter referred to in this specification refers to the angle between the virtual speaker and the front of the audio playback device. Specifically, refer to FIG. 2, which is a schematic diagram of the positional relationship between an audio playback device and a virtual speaker according to an embodiment of the present application. All of FIGS. 2 to 4 in this specification show the positional relationship between the audio playback device and the virtual speaker in a planar view. The position of the audio playback device in this embodiment is indicated as O. As can be understood, the audio playback device is worn by a person and moves along with them, and O can also indicate the user's position. Virtual speakers A and B are located on either side of audio playback device O, respectively. In this embodiment, the x-direction coordinate axis is defined with audio playback device O as the reference position. The x-axis is directly in front of the audio playback device, and the y-axis is directly to the right of the audio playback device. The xOy plane is the horizontal plane on which the audio playback device is located. When the audio playback device is worn correctly by the user, the x-axis direction is directly in front of the user, and the front-facing x-axis of the audio playback device overlaps with the center axis directly in front of the user. The angular parameter between virtual speaker A and audio playback device O can be expressed using the angle a formed by the connecting line between virtual speaker A and audio playback device O and the x-axis. Similarly, the angular parameter between virtual speaker B and audio playback device O can be expressed using the angle b formed by the connecting line between virtual speaker B and audio playback device O and the x-axis.

[0028] Step S13: Obtain the audio data to be processed by the audio playback device, and calculate and obtain processed spatial audio data according to the preset sound effect function and the position and angle information of each of the obtained at least two virtual speakers.

[0029] The audio data to be processed may be, for example, local audio data obtained from a sound source device, or audio data obtained via a network from an application program or a web page. The audio data to be processed may be, for example, music audio data, e-book audio data, television / movie audio, etc.

[0030] In this step, the position parameters L and θ1 in the corresponding sound effect functions are adjusted based on the position and angle information of the virtual speakers to obtain new sound effect functions, and the new sound effect functions are used to process the audio data to be processed to obtain processed spatial audio data.

[0031] In one implementation scenario, if the acquired user acceleration is greater than 0 (i.e., indicating that the audio playback device is accelerating and moving in response to the user), at least two virtual speakers are adjusted to be positioned in a direction opposite to the movement direction of the audio playback device (i.e., the angle between the connecting line between the virtual speakers and the audio playback device and the front of the audio playback device is greater than 90 degrees); if the acquired user acceleration is less than 0 (i.e., indicating that the audio playback device is decelerating and moving in response to the user), at least two virtual speakers are adjusted to be positioned in the same direction as the movement direction of the audio playback device (i.e., the angle between the connecting line between the virtual speakers and the audio playback device and the front of the audio playback device is less than 90 degrees).

[0032] The moving direction of the audio playback device is the direction in which the audio playback device moves in accordance with the user. Referring to Figures 2 and 3 in combination, the x-axis direction is directly ahead, and when the user's moving direction is in the x-axis direction, upon detecting accelerated movement, the virtual speakers are adjusted to be positioned in the direction opposite to the x-axis direction (i.e., adjusted to be behind the user), and the angles between the x-axis direction and the connecting lines between each of virtual speakers A and B and the audio playback device O are adjusted from the initial angle a to b. Therefore, when the user is currently moving in the x-axis direction, the virtual speakers are adjusted to be behind the user, which gives the user the auditory sensation of "swinging the virtual sound source backward."

[0033] As shown in Figures 2 and 4 in combination, the x-axis direction is directly ahead, and when the user's moving direction is in the x-axis direction, upon detecting deceleration, the virtual speakers are adjusted to be positioned in the direction indicated by x, and the angles between the connecting lines between each of virtual speakers A and B and the audio playback device O and the x-axis direction are adjusted from the initial angle a to c. When the user is currently moving in the direction indicated by x, the virtual speakers are adjusted to be in front of the user, which gives the user an auditory sensation of being "swung back" by the virtual sound source, causing the user to accelerate and chase the virtual sound source, thereby improving the interaction of sound effects during movement.

[0034] In one embodiment, adjusting the angle and distance information of the virtual speaker relative to the user based on the acceleration of the user specifically includes:

[0035] If the absolute value of the acquired acceleration of the user is equal to 0, the distance from the user among the position information of each of the at least two virtual speakers is set to 0, and the angle from the user among the angle information of each of the at least two virtual speakers is set to 0, i.e., the sound effect is adjusted to be returned to the ear.

[0036] If the absolute value of the acquired user acceleration is greater than a predetermined first threshold, the distance from the user among the position information of each of the at least two virtual speakers is set to a predetermined second threshold, and the angle from the user among the angle information of each of the at least two virtual speakers is set to a predetermined third threshold.

[0037] If the absolute value of the acquired acceleration of the user is greater than 0 and less than a first threshold, the distance from the user among the position information of each of the at least two virtual speakers is adjusted according to a predetermined first linear relationship, and the angle from the user among the angle information of each of the at least two virtual speakers is adjusted according to a predetermined second linear relationship.

[0038] A first linear relationship between the distance of the virtual speakers relative to the user and the user's acceleration, and a second linear relationship between the angle of the virtual speakers relative to the user and the user's acceleration can be preset. When the absolute value of the user's acceleration is detected to be greater than 0 and less than a first threshold, the angle and distance of each virtual speaker relative to the user can be adjusted based on the first and second linear relationships. In another embodiment, a corresponding relationship table between acceleration, angle, and distance can be determined based on the preset first and second linear relationships. After determining the current acceleration, the angle and distance corresponding to the current acceleration are looked up in the corresponding relationship table, and the looked-up angle and distance are used to adjust the angle and distance parameters in the sound effect function. The corresponding relationship table between acceleration, angle, and distance parameters is shown in the table below. The table divides acceleration into multiple acceleration ranges, and each acceleration value range corresponds to a corresponding angle and distance. The angle and distance values ​​corresponding to the looked-up acceleration range to which the current acceleration belongs are used as new angle and distance parameters in the sound effect function, thereby obtaining two virtual speakers with fixed positions relative to the audio playback device.

[0039] [Table 1]

[0040] As can be seen, the two virtual speakers in each embodiment of this specification are arranged symmetrically, so that under conditions of straight-line acceleration or deceleration, the virtual speakers are symmetrical with respect to the direction of movement of the audio playback device, and their angles and distances remain the same. Each embodiment of this specification uses a dual-channel sound effect as an example, but a similar approach can also be applied to multi-channel sound sources. Due to limitations of the Bluetooth transmission protocol, currently, all audio transmitted through earphones is stereo audio. In the industry, upmix algorithms can be used to convert audio files from stereo to multi-channel (e.g., 5.1), and deep learning methods for instrument separation can also be used to decompose stereo music files into multi-channel files containing different instruments. As can be seen, a multi-channel sound source can correspond to two or more virtual speakers, and this approach, like the above approach, can also be configured according to actual needs, with the linear relationship between the angle and acceleration of each virtual speaker and the linear relationship between the distance and acceleration, and is not limited herein.

[0041] The first linear relationship is that the ratio of the first threshold to the second threshold is equal to the ratio of the currently acquired acceleration of the user to the distance of the virtual speaker relative to the user. The second linear relationship is that the ratio of the first threshold to the third threshold is equal to the ratio of the currently acquired acceleration of the user to the angle of the virtual speaker relative to the user.

[0042] The first linear relationship between acceleration and the distance from the user of the position information of each virtual speaker during accelerated movement can be roughly expressed as follows: When it is detected that the audio playback device is decelerating and its acceleration is increasing, the virtual speaker moves in front of the user and the distance between the audio playback device and the virtual speaker increases; when it is detected that the audio playback device is accelerating and its acceleration is decreasing, the distance between the audio playback device and the virtual speaker decreases, and if the acceleration is zero, the virtual speaker returns to the ear. The first linear relationship between acceleration and distance during decelerated movement can be roughly expressed as follows: When it is detected that the audio playback device is decelerating and its acceleration is increasing, the virtual speaker moves in front of the user and the distance between the audio playback device and the virtual speaker increases; when it is detected that the audio playback device is decelerating and its acceleration is decreasing, the distance between the audio playback device and the virtual speaker decreases, and if the acceleration is zero, the virtual speaker returns to the ear.

[0043] The second linear relationship between acceleration and the angle of each virtual speaker relative to the user in the position information of each virtual speaker during accelerated movement can be roughly expressed as follows: When the audio playback device detects that it is accelerating and its acceleration is increasing, the virtual speaker moves behind the user, and the angle between the connecting line between the audio playback device and the virtual speaker and the front of the audio playback device decreases but is still greater than 90 degrees; when the audio playback device detects that it is accelerating and its acceleration is decreasing, the angle between the connecting line between the audio playback device and the virtual speaker and the front of the audio playback device increases, and if the acceleration is 0, the virtual speaker returns to the ear. The second linear relationship between acceleration and the angle of each virtual speaker relative to the user in the position information of each virtual speaker during decelerated movement can be roughly expressed as follows: When it is detected that the audio playback device is moving slowly and its acceleration is increasing, the virtual speaker moves in front of the user, and the angle between the connecting line between the audio playback device and the virtual speaker and the front of the audio playback device increases but is still smaller than 90 degrees; when it is detected that the audio playback device is moving slowly and its acceleration is decreasing, the angle between the connecting line between the audio playback device and the virtual speaker and the front of the audio playback device decreases; and when the acceleration is 0, the virtual speaker returns to the ear.

[0044] Referring to Figure 5, Figure 5 shows the change in the positional relationship between the audio playback device and the virtual speaker during the complete accelerated motion process of the audio playback device in the x direction from the rest time t11 to the times t12-t13-t14-t15, where O indicates the center position of the audio playback device, and A and B are the two virtual speakers under the two-source sound effect, respectively. During t11-t12-t13, the velocity v increases from 0 to v1, and the acceleration a1 increases from 0 to the maximum acceleration a1. max, the virtual speakers move back from the ears, and the angle between the connecting line between virtual speaker A and audio playback device O and the front of audio playback device O, and the angle between the connecting line between virtual speaker B and audio playback device O and the front of audio playback device O, both change from large to small but are still larger than 90 degrees. At the same time, the distance L between virtual speakers A, B and audio playback device O changes from small to large, and L max Between t13-t14-t15, the velocity changes from v1 to v max The acceleration a1 increases to the maximum acceleration a1 max The angle between the connecting line between virtual speaker A and audio playback device O and the front of audio playback device O, and the angle between the connecting line between virtual speaker B and audio playback device O and the front of audio playback device O, increase from small to large. At the same time, the distance L between virtual speakers A, B and audio playback device O increases until the speed reaches a maximum speed v max until it increases to L max When the acceleration a1 changes from small to large and becomes 0, the virtual speaker returns to the ear.

[0045] Referring to FIG. 6, FIG. 6 shows the change in the positional relationship between the audio playback device and the virtual speaker during the complete deceleration process of the audio playback device in the x direction from the rest time t21 to the time t22-t23-t24-t25, where O indicates the center position of the audio playback device, and A and B are the two virtual speakers under the two-source sound effect, respectively. During the time t21-t22-t23, the velocity v reaches the maximum velocity v max to v2, and the acceleration a2 decreases from 0 to the maximum acceleration a2 max , the virtual speakers move forward from the ears, and the angle between the connecting line between virtual speaker A and audio playback device O and the front of audio playback device O, and the angle between the connecting line between virtual speaker B and audio playback device O and the front of audio playback device O, both change from small to large, but are still smaller than 90 degrees. Between t23-t24-t25, the velocity decreases from v2 to v3, and the acceleration a2 reaches a maximum acceleration a2 maxThe angle between the connecting line between virtual speaker A and audio playback device O and the front of audio playback device O, and the angle between the connecting line between virtual speaker B and audio playback device O and the front of audio playback device O, both decrease from large to small, and when acceleration a2 becomes 0, the virtual speakers return to the ears.

[0046] Continuing to refer to Figure 6, between t21-t22-t23, the velocity v reaches the maximum velocity v max to v2, and the acceleration a2 decreases from 0 to the maximum acceleration a2 max , the virtual speakers move forward from the ears, and the distance L between each of the virtual speakers A and B and the audio playback device O changes from small to large, and L max Between t23-t24-t25, the velocity decreases from v2 to v3, and the acceleration a2 reaches the maximum acceleration a2 max to 0, and the distance L between each of the virtual speakers A and B and the audio playback device O is L. max When the acceleration a2 becomes 0, the virtual speaker returns to the ear.

[0047] Optionally, the motion information may further include speed information, and in each of the above embodiments, the angle between the user and each virtual speaker may have a predetermined linear relationship with the acceleration and speed during accelerating or decelerating movement. The distance between the user and each virtual speaker may also have a predetermined linear relationship with the acceleration and speed during decelerating or decelerating movement, and corresponding angle parameters and distance parameters may be determined based on the acceleration, speed, and the predetermined linear relationship under the current motion situation, which will not be described again here.

[0048] In another embodiment, the motion trajectory includes trajectory information of acceleration and deceleration movements. That is, the turning information of the audio playback device and the information on whether acceleration and / or deceleration are present can be simultaneously obtained. Regarding the turning information, the motion trajectory of the current audio playback device can be identified based on a map positioned by a GPS (Global Positioning System), and the turning information of the audio playback device can be determined based on the turning information of the link on which the audio playback device is currently located. Furthermore, the turning information can also be obtained based on a sensor such as a gyroscope provided in the audio playback device or a portable mobile device that is communicatively connected to the audio playback device.

[0049] 7, which is a flowchart of an embodiment of determining curve information according to the present invention. Based on the above embodiment, this embodiment further includes the following steps S21 to S22.

[0050] Step S21: Determine whether the moving direction of the audio playback device is deviated.

[0051] GPS positioning technology can be used to identify the road on which the audio playback device is currently located, and the included angle between the extension direction of the road and the current movement direction of the audio playback device can be determined. If the included angle exceeds a set angle threshold, it can be determined that the movement direction of the audio playback device has shifted. As shown in FIG. 8, x is the current movement direction of the audio playback device, y is the extension direction of link R1, and the included angle between them can be denoted as γ. Furthermore, the orientation of the audio playback device can be collected at set intervals. If the included angle between the orientation of the audio playback device at the current time and the orientation at the previous time exceeds a set angle threshold, it can be determined that the movement direction of the audio playback device has shifted. As shown in FIG. 9, the orientation of the audio playback device at the previous time is denoted as w, the orientation of the audio playback device at the previous time is denoted as v, and the included angle can be denoted as φ.

[0052] If the included angle does not exceed the set angle threshold, it is determined that there is no deviation.

[0053] Step S22: Determine the direction and angle of displacement of the audio playback device.

[0054] The deviation angle can be determined based on the included angle determination method in the previous step, and will not be described again here.

[0055] The direction of the movement deviation of the audio playback device can be determined based on the deviation of the audio playback device's current movement direction relative to the extension direction of the road. As shown in Figure 8, when the audio playback device changes from an x-direction to traveling along link R1, the direction of the movement deviation of the audio playback device can be determined to be a deviation to the right. When the audio playback device changes from an x-direction to traveling along link R2, the direction of the movement deviation of the audio playback device can be determined to be a deviation to the left. Alternatively, the direction of the movement deviation of the audio playback device at the current time relative to the direction of the audio playback device at the previous time can be determined to be a deviation to the right, for example, as shown in Figure 9, the direction of the movement deviation of the audio playback device can be determined to be a deviation to the right based on the fact that the direction v of the audio playback device at the current time is a deviation to the right relative to the direction w of the audio playback device at the current time.

[0056] In a situation where the audio playback device moves in a curved manner following the user, When the movement trajectory is an accelerating turning movement, at least two virtual speakers are adjusted to be located in a direction opposite to the movement direction of the audio playback device and on the opposite side of the turning direction, and when the movement trajectory is a decelerating turning movement, at least two virtual speakers are adjusted to be located in the same direction as the movement direction of the audio playback device and on the opposite side of the turning direction. For example, when the audio playback device turns left and accelerates with the user, the virtual speakers are adjusted to be located to the right rear of the audio playback device, and when the audio playback device turns right and decelerates with the user, the virtual speakers are adjusted to be located to the left front of the audio playback device.

[0057] 10 and 11, O in the drawings indicates the center position of the audio playback device, and A and B are two virtual speakers under the dual-sound-source sound effect. Figure 10 is a schematic diagram of the relative positions between the audio playback device and the virtual speakers in an accelerated turning movement situation, in which the audio playback device O accelerates and turns along a turning path during times t31-t32-t33-t34, x is the orientation of the audio playback device at each time, and the direction pointed by x is the front of the audio playback device O. In this case, during the accelerating movement process, the audio playback device O turns left and accelerates, and the two virtual speakers A and B are located behind the audio playback device O (i.e., the angle between the connecting line between at least one of the two virtual speakers A and B and the audio playback device O and the front of the audio playback device O is greater than 90 degrees). Figure 11 is a schematic diagram of the relative positions between the audio playback device and the virtual speakers in a deceleration turning movement situation, in which the audio playback device O decelerates and turns along the turning path during the time period t41-t42-t43-t44, x is the orientation of the audio playback device at each time, and the direction pointed by x is the front of the audio playback device O. In this case, during the deceleration movement process, the audio playback device O turns to the right and decelerates, and the two virtual speakers A and B are located in front of the audio playback device O (i.e., the angle between the connecting line between at least one of the two virtual speakers A and B and the audio playback device O and the front of the audio playback device O is less than 90 degrees).

[0058] During the acceleration or deceleration of the turn, the angle and acceleration between each virtual speaker and the audio playback device are linearly related to the offset angle. As can be seen, the angle and acceleration between each virtual speaker and the audio playback device are linearly related to the offset angle, and the sound field formed by the virtual speaker during the turn is offset to the left or right from the user.

[0059] Alternatively, when it is detected that the user's head is rotating left or right, head rotation angle information detected in real time by a head tracking device provided in the audio playback device is acquired, and angle information of each of the at least two virtual speakers is adjusted based on the acquired head rotation angle information and a preset head rotation angle adjustment mechanism. Specifically, when it is detected that the user's head is rotating left, an angle between a horizontal connecting line between the left virtual speaker of the user's head and the user and a direct front of the user is adjusted to decrease, and an angle between a horizontal connecting line between the right virtual speaker of the user's head and the user and a direct front of the user is adjusted to increase. When it is detected that the user's head is rotating right, an angle between a horizontal connecting line between the right virtual speaker of the user's head and the user and a direct front of the user is adjusted to decrease, and an angle between a horizontal connecting line between the left virtual speaker of the user's head and the user and a direct front of the user is adjusted to increase.

[0060] 12 and 13, X1, X2, and X3 are directly in front of the user's head, and O is the position of the audio playback device and the user. Before the user's head is rotated, the direction directly in front of the user's head is X1. When the user's head is rotated to the right in the X2 direction, the angle between the horizontal connecting line between the virtual speaker B on the right of the user's head and the user O and the direction directly in front of the user X2 is adjusted to decrease to a2, and the angle between the horizontal connecting line between the virtual speaker A on the left of the user's head and the user and the direction directly in front of the user X2 is adjusted to increase to a1. When the user's head is rotated to the left in the X3 direction, the angle between the horizontal connecting line between the virtual speaker B on the right of the user's head and the user O and the direction directly in front of the user X3 is adjusted to increase to a4, and the angle between the horizontal connecting line between the virtual speaker A on the left of the user's head and the user and the direction directly in front of the user X3 is adjusted to decrease to a3.

[0061] As can be understood, the angle parameters in each of the above embodiments may be the angles formed by two or more connecting lines between each of the two or more virtual speakers and the user's coordinate center, and any parameter that can adjust the angle between the connecting line between the virtual speaker and the audio playback device and the direction directly in front of the audio playback device is considered to be an alternative to the angle parameters of the present application and is considered to be within the scope of the protection claimed by the present application.

[0062] Step S14: Play the spatial audio data using an audio playback device.

[0063] The previous step processes the target audio data to obtain processed spatial audio data; in this step, an audio device is used to play the spatial audio data, which is adjusted based on the user's motion information and has corresponding spatial characteristics. During the user's continuous motion process, the spatial characteristics of the played audio will change accordingly according to changes in the user's motion state.

[0064] Unlike the prior art, this embodiment adjusts the position parameters in the sound effect function based on the motion information sensed by the audio playback device as the user moves, and further adjusts the angle and distance between the virtual speaker and the audio playback device, i.e., adjusts the orientation of the virtual speaker relative to the user, ultimately achieving the purpose of adjusting the sound effect. The audio playback effect changes dynamically with changes in motion information, giving the audio a more vivid expression effect, improving the user's sense of realism, and satisfying the emotional needs of the user's "audio partner", which is beneficial to improving the exercise experience and can guide the user to better achieve their exercise goals.

[0065] Please refer to FIG. 14, which is a structural schematic block diagram of an embodiment of an audio playback device according to the present application.

[0066] The audio playback device 100 includes an acquisition module 110, a parameter adjustment module 120, and an audio playback module 130. The acquisition module 110 is used to acquire audio data to be processed by the audio playback device and to acquire motion information that the audio playback device follows the user. The parameter adjustment module 120 is used to adjust a position parameter between the audio playback device and a virtual speaker in a sound effect function based on the motion information. The position parameter includes at least a horizontal angle parameter between the audio playback device and the virtual speaker. The position of the virtual speaker is a virtual sound source position processed by the sound effect function. The audio playback module 130 uses the adjusted sound effect function to convert the audio data to be processed into data to be played, and the audio playback device outputs the data to be played.

[0067] The audio playback device 100 may further include a communication module (not shown), which is used to establish a wired or wireless communication connection with the audio source device and receive audio data to be processed from the audio source device.

[0068] For example, the audio source device may be a mobile phone, a tablet computer, or a wearable audio source device such as a wristwatch or a band. The audio source device may store local audio data or may obtain audio data to be processed from an application program or a web page via a network. The audio data to be processed may be, for example, audio data of music, audio data of e-books, audio from television / movies, etc.

[0069] For the specific form of each step executed by each process, please refer to the description of each step in the embodiment of the audio processing method according to the present application, and the description will not be repeated here.

[0070] 15, which is a structural schematic block diagram of another embodiment of an audio playback device according to the present application. The audio playback device 200 includes a processor 210 and a memory 220 interconnected, the memory 220 stores a computer program, and the processor 210 executes the computer program to implement the audio processing method described in each of the above embodiments.

[0071] For a description of each step performed by the process, please refer to the description of each step in the embodiment of the audio processing method according to the present application, and the description will not be repeated here.

[0072] The memory 220 can be used to store program data and modules, and the processor 210 executes various functional applications and data processing by operating the program data and modules stored in the memory 220. The memory 220 may mainly include a program storage area and a data storage area. The program storage area can store an operating system, an application program required for at least one function (e.g., a parameter adjustment function, etc.), etc., and the data storage area can store data generated based on the use of the audio playback device 200 (e.g., audio data to be processed, motion information data, etc.). The memory 220 may also include a high-speed random access memory and may further include at least one non-volatile memory such as a magnetic disk storage device, a flash device, or other volatile solid-state storage device. Correspondingly, the memory 220 may further include a memory controller for providing access to the memory 220 by the processor 210.

[0073] In each embodiment of the present application, the disclosed methods and devices may be implemented in other ways. For example, the above-described embodiments of the audio playback device 200 are merely exemplary. For example, the division of the modules or units is merely a division of logical functions. In actual implementation, other division methods may be used. For example, multiple units or components may be combined or integrated into other systems, or some features may be omitted or not implemented. Furthermore, the shown or discussed mutual couplings or direct couplings or communication connections may be indirect couplings or communication connections via several interfaces, devices, or units, and may be electrical, mechanical, or other forms.

[0074] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, i.e., they may be located in one place or distributed across multiple network units, and some or all of the units can be selected according to actual needs to achieve the objective of the solution of this embodiment.

[0075] Furthermore, each functional unit in each embodiment of the present application may be integrated into one processing unit, each unit may exist physically independently, or two or more units may be integrated into one unit. The integrated units may be realized in the form of hardware or in the form of software functional units.

[0076] The integrated unit may be realized in the form of a software functional unit and stored in a single computer-readable storage medium when sold or used as an independent product. Based on this understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or a part of the technical solution, may be embodied in the form of a software product, and the computer software product is stored in a single storage medium.

[0077] Referring to FIG. 16, FIG. 16 is a structural schematic block diagram of one embodiment of a computer-readable storage medium according to the present application, in which program data 310 is stored in a computer-readable storage medium 300, and when the program data 310 is executed, steps of each embodiment of the audio processing method described above are realized.

[0078] For a description of each step performed by the process, please refer to the description of each step in the embodiment of the audio processing method according to the present application, and the description will not be repeated here.

[0079] The computer-readable storage medium 300 may be various media that can store program codes, such as a U disk, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.

[0080] The above are merely examples of the present application and do not limit the scope of the patent of the present application. Any equivalent structure or equivalent flow conversion made by using the contents of the specification and drawings of the present application, or that can be used directly or indirectly in other related technical fields, is included within the scope of the claims of the present application.

Claims

1. a step of acquiring motion information of the audio playback device moving along with the user's movement, the motion information including at least the user's motion trajectory, real-time motion speed, and real-time acceleration; calculating and acquiring position and angle information of at least two virtual speakers for the user based on the acquired movement trajectory, real-time movement speed, real-time acceleration, and a preset sound effect function of the user; acquiring audio data to be processed from the audio playback device, and calculating and obtaining processed spatial audio data based on the preset sound effect function and the acquired position and angle information of each of the at least two virtual speakers; and reproducing the spatial audio data using the audio playback device.

2. In the step of calculating and obtaining position and angle information of each of at least two virtual speakers for the user, the method further comprises: When detecting that the user's head is rotating left or right, acquiring head rotation angle information detected in real time by a head tracking device provided in the audio playback device; and adjusting angle information of each of the at least two virtual speakers based on the acquired head rotation angle information and a preset head rotation angle adjustment mechanism.

3. The head rotation angle adjustment mechanism includes: When it is determined that the user's head is rotating to the left, adjusting the angle between the left virtual speaker of the user's head and a horizontal connecting line between the user and the front of the user to decrease, and adjusting the angle between the right virtual speaker of the user's head and a horizontal connecting line between the user and the front of the user to increase; 3. The method of claim 2, further comprising the step of: adjusting, when it is determined that the user's head is rotating to the right, an angle between a horizontal connecting line between the right virtual speaker of the user's head and the user and directly in front of the user to decrease; and adjusting, when it is determined that the user's head is rotating to the right, an angle between a horizontal connecting line between the right virtual speaker of the user's head and the user and directly in front of the user to increase.

4. When the sound effect function is executed, If the absolute value of the acquired acceleration of the user is greater than a predetermined first threshold, setting a distance from the user among the position information of each of the at least two virtual speakers to a predetermined second threshold, and setting an angle from the user among the angle information of each of the at least two virtual speakers to a predetermined third threshold; If the absolute value of the acquired acceleration of the user is equal to 0, setting a distance with respect to the user among the position information of each of the at least two virtual speakers to 0, and setting an angle with respect to the user among the angle information of each of the at least two virtual speakers to 0; and adjusting, when the absolute value of the acquired acceleration of the user is greater than 0 and less than the first threshold, a distance from the user to the position information of each of the at least two virtual speakers in accordance with a predetermined first linear relationship, and adjusting an angle from the user to the angle information of each of the at least two virtual speakers in accordance with a predetermined second linear relationship.

5. The first linear relationship is such that the ratio of the first threshold to the second threshold is equal to the ratio of the currently acquired acceleration of the user to the distance of the virtual speaker relative to the user; and / or 5. The method of claim 4, wherein the second linear relationship is such that the ratio of the first threshold to the third threshold is equal to the ratio of the currently acquired acceleration of the user to the angle of the virtual speaker relative to the user.

6. 2. The method of claim 1, wherein the movement information is acquired in real time using a positioning device and an acceleration sensor, and at least one of the positioning device and the acceleration sensor is provided in the audio playback device or a smart mobile device communicatively connected to the audio playback device.

7. If the acquired acceleration of the user is greater than 0, the at least two virtual speakers are positioned in a direction opposite to a moving direction of the audio playback device, respectively; The method of claim 1 , wherein when the acquired acceleration of the user is less than 0, the at least two virtual speakers are positioned in the same direction as the direction of movement of the audio playback device.

8. the motion trajectory includes an acceleration turning movement and a deceleration turning movement; If the movement trajectory is an accelerated turning movement, the at least two virtual speakers are located in a direction opposite to a movement direction of the audio reproduction device and on a side opposite to a turning direction; 2. The method of claim 1, wherein when the movement trajectory is a deceleration turning movement, the at least two virtual speakers are located in the same direction as the movement direction of the audio reproduction device and on the opposite side of the turning direction.

9. 10. An audio playback device comprising a processor and a memory interconnected, the memory storing a computer program, the processor being adapted to execute the computer program to implement the steps of the method according to any one of claims 1 to 8.

10. A computer-readable storage medium having stored thereon program data, the program data implementing the steps of the method according to any one of claims 1 to 8 when executed by a processor.

Citation Information

Patent Citations

  • Vehicle

    JP2021160421A

  • Sound signal reproduction device, sound signal reproduction method, program and recording medium

    WO2016140058A1