Audio processing, model training method and device, equipment and storage medium

CN116684777BActive Publication Date: 2026-09-08ALIBABA DAMO (HANGZHOU) TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310454751.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-04-21
Publication Date
2026-09-08
Estimated Expiration
2043-04-21

AI Technical Summary

Technical Problem

[0003]目前可通过机器学习模型来生成空间音频,但是,机器学习模型生成的空间音频的精度有限,导致用户无法感受到声源的运动

Benefits of technology

[0046] In a tenth aspect, embodiments of this disclosure provide a computer-readable storage medium having a computer program stored thereon, the computer program being executed by a processor to implement the methods described in the first to seventh aspects.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116684777B_ABST
    Figure CN116684777B_ABST
Patent Text Reader

Abstract

The present disclosure relates to an audio processing method and device, a model training method and device, equipment and a storage medium. The present disclosure obtains monaural audio corresponding to a first target object in motion, and calculates the radial velocity of the first target object relative to a second target object. Further, according to the position information of the first target object, the posture information of the second target object, the monaural audio, and the radial velocity of the first target object relative to the second target object, a spatial audio is generated. The generated spatial audio contains the motion information of the first target object, and the difference between the phase of the audio signal in the spatial audio and the phase of the actual binaural audio is small, improving the accuracy of the phase calculation of the left and right ear channels in the spatial audio, greatly improving the accuracy of the generated spatial audio. Thus, when a user hears the spatial audio, the user feels the motion of the first target object according to the phase difference between the left and right ear audio signals in the spatial audio.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of information technology, and in particular to an audio processing, model training method, apparatus, device and storage medium. Background Technology

[0002] Since humans are binaural, we can determine the exact location of a sound source by analyzing the phase difference between the sounds heard by the left and right ears. Therefore, providing users with immersive spatial audio is crucial in speech technology. Spatial audio, also known as binaural audio or stereo audio, specifically comprises two audio signals, one provided to the left ear and the other to the right ear.

[0003] Currently, spatial audio can be generated using machine learning models. However, the accuracy of spatial audio generated by machine learning models is limited, causing users to be unable to perceive the movement of the sound source. Summary of the Invention

[0004] To solve the above-mentioned technical problems, or at least partially solve them, this disclosure provides an audio processing, model training method, apparatus, device, and storage medium, enabling users to accurately determine the location of a first target object, such as a sound source, based on the phase difference between the left and right ear audio signals in the spatial audio, and to perceive the movement of the first target object, thereby providing users with an immersive experience and greatly improving the user experience.

[0005] In a first aspect, embodiments of this disclosure provide an audio processing method, including:

[0006] Obtain the mono audio corresponding to the first target object in motion;

[0007] Calculate the radial velocity of the first target object relative to the second target object;

[0008] Spatial audio is generated based on the position information of the first target object, the posture information of the second target object, the mono audio, and the radial velocity of the first target object relative to the second target object.

[0009] Secondly, embodiments of this disclosure provide an audio processing method, including:

[0010] Based on the position information of a movable first target object in the virtual space relative to a second target object in the virtual space, calculate the radial velocity of the first target object relative to the second target object;

[0011] Spatial audio is generated based on the position information of the first target object, the posture information of the second target object, the mono audio pre-configured for the first target object, and the radial velocity of the first target object relative to the second target object.

[0012] Play the spatial audio.

[0013] Thirdly, embodiments of this disclosure provide an audio processing method, including:

[0014] Based on the position information of the movable target relative to the target user displayed in the virtual reality display device, the radial velocity of the movable target relative to the target user is calculated, wherein the target user is wearing the virtual reality display device;

[0015] Spatial audio is generated based on the position information of the movable target, the posture information of the target user, the mono audio pre-configured for the movable target, and the radial velocity of the movable target relative to the target user;

[0016] The spatial audio is played through the virtual reality display device.

[0017] Fourthly, embodiments of this disclosure provide an audio processing method, including:

[0018] Acquire mono audio emitted by a moving target;

[0019] Based on the position information of the movable target relative to the target user, the radial velocity of the movable target relative to the target user is calculated, and the target user is wearing an augmented reality device;

[0020] Spatial audio is generated based on the position information of the movable target, the posture information of the target user, the mono audio, and the radial velocity of the movable target relative to the target user;

[0021] The spatial audio is played through the augmented reality device.

[0022] Fifthly, embodiments of this disclosure provide an audio processing method, including:

[0023] Collect mono audio emitted by the first user during movement;

[0024] Based on the position information of the first user relative to the second user in the preset space, calculate the radial velocity of the first user relative to the second user;

[0025] Spatial audio is generated based on the location information of the first user, the posture information of the second user, the mono audio, and the radial velocity of the first user relative to the second user.

[0026] The spatial audio is sent to the audio playback device worn by the second user.

[0027] Sixthly, embodiments of this disclosure provide an audio processing method, including:

[0028] Collect mono audio emitted by the first user during movement;

[0029] Based on the position information of the first user in the preset space relative to the target position in the preset space, calculate the radial velocity of the first user relative to the target position;

[0030] Spatial audio is generated based on the location information of the first user, the posture information of the second user, the mono audio, and the radial velocity of the first user relative to the target position.

[0031] The spatial audio is sent to the audio playback device worn by the second user, who is located in a different space from the first user.

[0032] Seventhly, embodiments of this disclosure provide a model training method, the method comprising:

[0033] Acquire mono audio emitted by a movable sound source;

[0034] Acquire the two-channel audio collected by the first and second pickup units of the target object;

[0035] The position information of the movable sound source, the posture information of the target object, the mono audio, and the radial velocity of the movable sound source relative to the target object are input into the machine learning model to be trained, so that the machine learning model outputs spatial audio.

[0036] The machine learning model is trained based on the spatial audio and the stereo audio, and the trained machine learning model is used to execute the audio processing method described above.

[0037] Eighthly, embodiments of this disclosure provide an audio processing apparatus, comprising:

[0038] The acquisition module is used to acquire the mono audio corresponding to the first target object in motion;

[0039] The calculation module is used to calculate the radial velocity of the first target object relative to the second target object;

[0040] The generation module is used to generate spatial audio based on the position information of the first target object, the posture information of the second target object, the mono audio, and the radial velocity of the first target object relative to the second target object.

[0041] Ninthly, embodiments of this disclosure provide an electronic device, including:

[0042] Memory;

[0043] Processor; and

[0044] Computer programs;

[0045] The computer program is stored in the memory and configured to be executed by the processor to implement the methods described in the first to seventh aspects.

[0046] In a tenth aspect, embodiments of this disclosure provide a computer-readable storage medium having a computer program stored thereon, the computer program being executed by a processor to implement the methods described in the first to seventh aspects.

[0047] The audio processing, model training method, apparatus, device, and storage medium provided in this disclosure acquire mono audio corresponding to a moving first target object and calculate the radial velocity of the first target object relative to a second target object. Further, spatial audio is generated based on the position information of the first target object, the posture information of the second target object, the mono audio, and the radial velocity of the first target object relative to the second target object. This ensures that the generated spatial audio contains the motion information of the first target object, thereby matching the generated spatial audio with the actual stereo audio perceived by the second target object when the mono audio corresponding to the first target object reaches it. It also minimizes the difference between the phase of the audio signal in the spatial audio and the phase of the actual stereo audio, improving the accuracy of phase calculation for the left and right ear channels in the spatial audio and significantly enhancing the accuracy of the generated spatial audio. Consequently, when a user hears the spatial audio, they can accurately determine the position of the first target object, such as the sound source, based on the phase difference between the left and right ear audio signals in the spatial audio, and perceive the movement of the first target object, thus providing an immersive experience and greatly improving the user experience. Attached Figure Description

[0048] The accompanying drawings, which are incorporated in and form a part of this specification, illustrate embodiments consistent with this disclosure and, together with the description, serve to explain the principles of this disclosure.

[0049] To more clearly illustrate the technical solutions in the embodiments of this disclosure or the prior art, the accompanying drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0050] Figure 1 A flowchart of the model training method provided in this embodiment of the disclosure;

[0051] Figure 2 A schematic diagram illustrating an application scenario provided by an embodiment of this disclosure;

[0052] Figure 3 A schematic diagram of radial velocity provided for an embodiment of this disclosure;

[0053] Figure 4 A schematic diagram of the structure of the twisted network provided in the embodiments of this disclosure;

[0054] Figure 5 This is a schematic diagram of the structure of a dual-channel gradient network provided in another embodiment of the present disclosure;

[0055] Figure 6 This is a flowchart of an audio processing method provided in another embodiment of the present disclosure;

[0056] Figure 7 This is a flowchart of an audio processing method provided in another embodiment of the present disclosure;

[0057] Figure 8 This is a flowchart of an audio processing method provided in another embodiment of the present disclosure;

[0058] Figure 9 This is a flowchart of an audio processing method provided in another embodiment of the present disclosure;

[0059] Figure 10 A schematic diagram illustrating an application scenario provided by another embodiment of this disclosure;

[0060] Figure 11 This is a flowchart of an audio processing method provided in another embodiment of the present disclosure;

[0061] Figure 12 A schematic diagram illustrating an application scenario provided by another embodiment of this disclosure;

[0062] Figure 13 This is a flowchart of an audio processing method provided in another embodiment of the present disclosure;

[0063] Figure 14 A schematic diagram illustrating an application scenario provided by another embodiment of this disclosure;

[0064] Figure 15This is a flowchart of an audio processing method provided in another embodiment of the present disclosure;

[0065] Figure 16 A schematic diagram illustrating an application scenario provided by another embodiment of this disclosure;

[0066] Figure 17 This is a schematic diagram of the structure of the audio processing apparatus provided in the embodiments of this disclosure;

[0067] Figure 18 A schematic diagram of the structure of an electronic device embodiment provided in this disclosure. Detailed Implementation

[0068] To better understand the above-mentioned objectives, features, and advantages of this disclosure, the solutions disclosed herein will be further described below. It should be noted that, unless otherwise specified, the embodiments and features described herein can be combined with each other.

[0069] Numerous specific details are set forth in the following description in order to provide a full understanding of this disclosure, but this disclosure may also be implemented in other ways different from those described herein; obviously, the embodiments in the specification are only some, and not all, of the embodiments of this disclosure.

[0070] It should be noted that the location information (including the location information of moving objects, etc.), posture information (such as the posture information of users) and mono audio (including but not limited to the sound emitted by users in motion, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties. Furthermore, the collection, use and processing of the relevant data must comply with the relevant laws, regulations and standards of the relevant countries and regions, and corresponding operation entry points are provided for users to choose to authorize or refuse.

[0071] In addition, the audio processing method provided in this application involves the following explanations of terms, detailed below:

[0072] Wave observation network (Wave Net): A common type of waveform synthesis neural network.

[0073] Stereo: Also called spatial audio or two-channel audio, because the distance between the sound source and the left and right ears is different, the two channels have different waveforms. For example, the sound emitted by the sound source is mono audio, but after propagation, it becomes two-channel audio when it reaches both ears.

[0074] Warp Net: A neural network used for synthesizing two-channel audio.

[0075] Binaural Grad: A neural network used to synthesize two-channel audio.

[0076] Since humans are binaural, the phase difference between the sounds heard by the left and right ears can be used to determine the specific location of a sound source. Therefore, providing users with immersive spatial audio is crucial in speech technology. Spatial audio, also called binaural audio or stereo sound, specifically includes two audio signals, one provided to the left ear and the other to the right ear. Currently, spatial audio can be generated using machine learning models; however, the accuracy of spatial audio generated by machine learning models is limited, causing users to be unable to perceive the movement of the sound source. To address this problem, this disclosure provides a model training method, which will be described below with reference to specific embodiments.

[0077] Figure 1 This is a flowchart illustrating a model training method provided in an embodiment of this disclosure. The method can be executed by a model training device, which can be implemented in software and / or hardware. This device can be configured in an electronic device, such as a server or terminal, where the terminal specifically includes a mobile phone, computer, or tablet computer. The server can specifically be a cloud server, and the model training method can be executed in the cloud. Several computing nodes (cloud servers) can be deployed in the cloud, each with computing, storage, and other processing resources. In the cloud, multiple computing nodes can be organized to provide a certain service; of course, a single computing node can also provide one or more services. The cloud can provide the service by providing a service interface, which users call to use the corresponding service. Service interfaces include Software Development Kits (SDKs), Application Programming Interfaces (APIs), etc. Furthermore, the model training method described in this embodiment is applicable to... Figure 2 The application scenarios shown are as follows. Figure 2 As shown, this application scenario includes terminal 21 and server 22. Server 22 can execute the model training method. For example, server 22 can train the machine learning model to be trained according to this method. When training is complete, the trained machine learning model can be stored locally on server 22 or deployed to terminal 21 or other servers, so that server 22, terminal 21, or other servers can convert mono audio into accurate spatial audio based on the trained machine learning model. The following section combines... Figure 2 This method will be described in detail, such as Figure 1 As shown, the specific steps of this method are as follows:

[0078] S101. Acquire mono audio emitted by a movable sound source.

[0079] like Figure 3As shown, the sound source 31 is movable, and 32 represents the mono audio emitted by the sound source 31. In this embodiment, a movable sound source is referred to as a movable sound source. Furthermore, this embodiment does not limit the specific form of the sound source 31. For example, the sound source 31 can be a user emitting sound, such as a user speaking or clapping. Alternatively, the sound source 31 can be a loudspeaker, a speaker, a robot, a car model, etc. Among these, a robot can be a humanoid robot, a robotic vacuum cleaner, etc. It is understood that the sound source 31 is not limited to the forms described above, and can also be other movable objects capable of producing sound or generating audio signals.

[0080] Specifically, the sound source 31 may be equipped with an audio acquisition device, which can acquire the mono audio emitted by the sound source 31 in real time or periodically, and send the mono audio emitted by the sound source 31 to the server 22. Alternatively, the sound source 31 may send the mono audio it generates to the server 22.

[0081] S102. Obtain the dual-channel audio collected by the first and second pickup units of the target object.

[0082] like Figure 3 As shown, a target object 33 exists around the sound source 31. The target object 33 can be a listener, an audience member, a sound pickup device, a human mannequin, etc. Specifically, the target object 33 includes a first sound pickup unit 34 and a second sound pickup unit 35. Specifically, when the target object 33 is a listener or audience member, that is, when the target object 33 is a person, the first sound pickup unit 34 can be an audio acquisition device worn on the right ear of the target object 33, and the second sound pickup unit 35 can be an audio acquisition device worn on the left ear of the target object 33.

[0083] When the target object 33 is a sound pickup device, the first sound pickup unit 34 and the second sound pickup unit 35 are two sound pickup units in that device. Assuming the center point of the sound pickup device is... Figure 3 As shown by the origin 36, the distance of the first pickup unit 34 relative to the origin 36 can be the distance of the person's right ear relative to the center of the person's head, or the distance of the first pickup unit 34 relative to the origin 36 is determined based on the distance of the person's right ear relative to the center of the person's head. Similarly, the distance of the second pickup unit 35 relative to the origin 36 can be the distance of the person's left ear relative to the center of the person's head, or the distance of the second pickup unit 35 relative to the origin 36 is determined based on the distance of the person's left ear relative to the center of the person's head.

[0084] When the target object 33 is a human body model, the human body model has parts or components that simulate human organs. For example, the human body model has a head model, a right ear model, and a left ear model. The distance of the right ear model relative to the center of the head model can be the distance of the right ear relative to the center of the head, or the distance of the right ear model relative to the center of the head model is determined based on the distance of the right ear relative to the center of the head. Similarly, the distance of the left ear model relative to the center of the head model can be the distance of the left ear relative to the center of the head, or the distance of the left ear model relative to the center of the head model is determined based on the distance of the left ear relative to the center of the head. In addition, the right ear model and the left ear model of the human body model can each be equipped with an audio acquisition device. For example, the audio acquisition device worn on the right ear model is denoted as the first pickup unit 34, and the audio acquisition device worn on the left ear model is denoted as the second pickup unit 35.

[0085] It is understandable that the implementation form of the target object 33 is not limited to the types mentioned above, and can also be other objects that can achieve the sound pickup function.

[0086] When the mono audio emitted by the sound source 31 reaches the target object 33 after being transmitted through a transmission medium, such as through air, the first pickup unit 34 and the second pickup unit 35 can each collect an audio signal. The audio signals collected by the first pickup unit 34 and the second pickup unit 35 are denoted as stereo audio.

[0087] For example, the audio signal acquired by the first pickup unit 34 is denoted as audio signal A, and the audio signal acquired by the second pickup unit 35 is denoted as audio signal B. Since the positions of the first pickup unit 34 and the second pickup unit 35 relative to the sound source 31 are different, and the directions of the first pickup unit 34 and the second pickup unit 35 relative to the sound source 31 are also different, audio signal A and audio signal B are not entirely identical. For example, the phase of audio signal A is different from the phase of audio signal B. Alternatively, the time at which audio signal A arrives at the first pickup unit 34 is different from the time at which audio signal B arrives at the second pickup unit 35.

[0088] When the target object 33 is a listener or viewer, if audio signal A is received by the right ear and audio signal B is received by the left ear, the brain of the target object 33 can determine the position and direction of the sound source 31 relative to the target object 33 based on the phase difference between audio signals A and B, or based on the time difference between audio signals A and B. It is understood that when the target object 33 is moving, audio signals A and B can change in real time, and the phase difference and time difference will also change accordingly. At this time, the target object 33 can perceive the movement of the sound source 31 based on the real-time changing phase difference and time difference. For example, the target object 33 can not only perceive that the sound source 31 is moving, but also perceive how the sound source 31 is moving, such as the sound source 31 circling around the target object 33, or the sound source 31 continuously approaching the target object 33, etc. Furthermore, in this embodiment, the target object 33 can be stationary.

[0089] S103. The position information of the movable sound source, the posture information of the target object, the mono audio, and the radial velocity of the movable sound source relative to the target object are input into the machine learning model to be trained, so that the machine learning model outputs spatial audio.

[0090] For example Figure 3 The coordinate system established by the x and y axes shown is a top view of a Cartesian coordinate system with the center of the target object 33 as the origin; that is, the height dimension of the Cartesian coordinate system is ignored. Specifically, the sound source 31 emits sound at a speed v. xy It moves on the plane formed by the x-axis and y-axis. Based on the velocity v xy The radial velocity of the sound source 31 relative to the target object 33 can be obtained. For example, the radial velocity of the sound source 31 relative to the target object 33 includes the radial velocity of the sound source 31 relative to the first pickup unit 34 and the radial velocity of the sound source 31 relative to the second pickup unit 35. For example, the direction of the line connecting the sound source 31 to the first pickup unit 34 is r, and the velocity is v. xy It can be decomposed into components v in the r direction. r and components perpendicular to the r direction Component v r This can be denoted as the radial velocity of the sound source 31 relative to the first pickup unit 34. Similarly, the radial velocity of the sound source 31 relative to the second pickup unit 35 can be calculated.

[0091] In addition, Figure 3In the coordinate system established by the x-axis and y-axis shown, the position information of the sound source 31 and the attitude information of the target object 33 can also be determined. For example, the position information of the sound source 31 can be the coordinates of the sound source 31 in this coordinate system. The attitude information of the target object 33 can be the head direction of the target object 33 in this coordinate system.

[0092] Furthermore, server 22 can input the mono audio emitted by sound source 31, the position information of sound source 31, the posture information of target object 33, and the radial velocity of sound source 31 relative to target object 33 into the machine learning model to be trained, so that the machine learning model outputs spatial audio. Specifically, the spatial audio output by the machine learning model is the stereo audio predicted by the machine learning model to be trained when the mono audio emitted by sound source 31 propagates to target object 33, given the mono audio emitted by sound source 31, the position information of sound source 31, the posture information of target object 33, and the radial velocity of sound source 31 relative to target object 33.

[0093] S104. The machine learning model is trained based on the spatial audio and the dual-channel audio, and the trained machine learning model is used to execute the audio processing method.

[0094] Since the stereo audio collected by the first pickup unit 34 and the second pickup unit 35 is the actual stereo audio when the mono audio propagates to the target object 33, and there will be a difference between the stereo audio predicted by the machine learning model to be trained when the mono audio propagates to the target object 33 and the actual stereo audio, the machine learning model is trained based on the spatial audio output by the machine learning model to be trained and the actual stereo audio.

[0095] For example, based on the similarity or difference between the spatial audio output by the machine learning model to be trained and the actual stereo audio, the parameters in the machine learning model can be adjusted so that the spatial audio output by the machine learning model and the actual stereo audio gradually become similar during subsequent iterative training. For instance, the parameter adjustment can be guided by the gradient data between the spatial audio output by the machine learning model and the actual stereo audio. Here, the gradient is essentially a vector representing the maximum value of the directional derivative of a function at a given point along that direction; that is, the function changes most rapidly and has the largest rate of change along that direction (the direction of the gradient) at that point. Based on this principle, the direction of parameter adjustment can be guided so that the spatial audio output by the machine learning model is closer to the actual stereo audio. This results in a trained machine learning model. This trained machine learning model can then be used to perform the audio processing methods described below.

[0096] Understandably, since the sound source 31 is mobile, the mono audio emitted by the sound source 31, the position information of the sound source 31, and the radial velocity of the sound source 31 relative to the target object 33 are variable at different times. Therefore, the mono audio, the position information of the sound source 31, the radial velocity of the sound source 31 relative to the target object 33, the stereo audio collected by the first pickup unit 34 and the second pickup unit 35, and the spatial audio output by the machine learning model to be trained can be aligned. For example, the position information of the sound source 31 indicates the location of the sound source 31 when it emits the mono audio. The radial velocity of the sound source 31 relative to the target object 33 is the radial velocity of the sound source 31 relative to the target object 33 when it emits the mono audio. The stereo audio collected by the first pickup unit 34 and the second pickup unit 35 is the stereo audio collected by the first pickup unit 34 and the second pickup unit 35 when the mono audio propagates to the target object 33. The spatial audio output by the machine learning model to be trained is the spatial audio predicted by the machine learning model based on the mono audio. Thus, during the training process, the machine learning model is trained based on the spatial audio corresponding to the same mono audio and the actual stereo audio.

[0097] In this embodiment, when the mono audio emitted by a movable sound source reaches the target object, dual-channel audio is acquired through the first and second pickup units of the target object. Furthermore, the position information of the movable sound source, the posture information of the target object, the mono audio, and the radial velocity of the movable sound source relative to the target object are input into a machine learning model to be trained. This allows the machine learning model to learn the motion information of the movable sound source based on its radial velocity relative to the target object, thereby enabling the spatial audio output by the machine learning model to contain the motion information of the movable sound source. Further, the machine learning model is trained based on the spatial audio and the dual-channel audio, so that the motion information of the movable sound source contained in the spatial audio gradually approaches the actual motion information of the movable sound source. This allows the trained machine learning model to output accurate spatial audio, that is, spatial audio that is infinitely close to the actual dual-channel audio. When a user hears this spatial audio, they can perceive the motion of the movable sound source, thus providing an immersive experience and greatly improving the user experience.

[0098] Specifically, this embodiment does not limit the structure of the machine learning model as described above. For example, the machine learning model can be a warp network or a binaural gradient network as described above.

[0099] like Figure 4As shown, the Warp Net comprises Neural Time Warping and Temporal Conv Net. Neural Time Warping includes a warp layer, a warp activation function, a neural warp layer, and a geometric warp layer. The Temporal Conv Net consists of N hyperconvolutional layers. C0 represents the position information of the sound source 31 and the pose information of the target object 33, as shown above. r-left This indicates the radial velocity of the sound source 31 relative to the second pickup unit 35. r-right This represents the radial velocity of the sound source 31 relative to the first pickup unit 34. 1:T This indicates that the sound source 31 emits a mono audio signal, which includes T-frame audio.

[0100] In the case of Figure 4 During the training of the twisted network shown, x 1:T C0, v r-left v r-right ρ can be used as input to this twisted network. 1:T This represents the output of the warp activation function in neural temporal warp, or the output of other intermediate layers above the warp activation function. This represents the audio signals in the two frequency domains of the neural time-warped output. This represents the two temporal audio signals output by the temporal convolutional network, i.e. This is the spatial audio output by the twisted network to be trained. Further, the twisted network is trained based on the actual two-channel audio collected by the first pickup unit 34 and the second pickup unit 35 and the spatial audio output by the twisted network.

[0101] like Figure 5 As shown, the Binaural Grad network consists of M residual blocks and a conditional network. The M residual blocks are sequentially connected; for example, the output of the next residual block is the input of the previous residual block. Each residual block includes a fully connected layer (FC), a dilated convolutional layer (Dilated Conv), and convolutional layers (Conv). The conditional network includes multiple convolutional layers (Conv).

[0102] When training a dual-channel gradient network, C0, v r-left v r-rightThis can be used as input to the two-channel gradient network, so that the output of the two-channel gradient network is the spatial audio predicted by the two-channel gradient network. Further, the two-channel gradient network is trained based on the actual two-channel audio collected by the first pickup unit 34 and the second pickup unit 35 and the spatial audio output by the two-channel gradient network.

[0103] It is understood that after the machine learning model described above is trained, the trained machine learning model can be stored on server 22 or deployed to terminal 21 or other servers. This allows server 22, terminal 21, or other servers to implement the audio processing method described in the following embodiments based on the trained machine learning model.

[0104] Figure 6 This is a flowchart illustrating an audio processing method according to another embodiment of the present disclosure. For example, this audio processing method can be executed by a cloud server, such as server 22. In this embodiment, the specific steps of the method are as follows:

[0105] S601. Obtain the mono audio corresponding to the first target object in motion.

[0106] Specifically, the first target object is a movable sound source, and the mono audio emitted by the sound source can be acquired by an audio acquisition device. Alternatively, in some embodiments, the system controlling the first target object can pre-configure a mono audio signal for the first target object and treat this mono audio signal as the audio emitted by the first target object.

[0107] For example, when the primary target is a user who is speaking, the user may wear an audio acquisition device that captures the mono audio emitted by the user and sends the mono audio to the server 22.

[0108] When the primary target is a game character displayed by a game application installed on an electronic device, the electronic device is considered a control system for that game character, or the control system for that game character can be a service platform established by the developer of the game application. This control system can configure mono audio for the game character and treat that mono audio as audio emitted by the primary target. Furthermore, the control system can send that mono audio to server 22.

[0109] S602. Calculate the radial velocity of the first target object relative to the second target object.

[0110] For example, the second target object is located around the first target object, and the second target object can be stationary. For example, the second target object can be a listener, an audience member, a sound pickup device, a mannequin, etc. The server 22 can calculate the radial velocity of the first target object relative to the second target object based on the position information of the first target object relative to the second target object.

[0111] S603. Generate spatial audio based on the position information of the first target object, the posture information of the second target object, the mono audio, and the radial velocity of the first target object relative to the second target object.

[0112] For example, server 22 can generate spatial audio based on the position information of the first target object, the posture information of the second target object, the mono audio, and the radial velocity of the first target object relative to the second target object.

[0113] Optionally, the position information of the first target object is the position information of the first target object in a first coordinate system, wherein the first coordinate system is a coordinate system with the second target object as the origin; the attitude information of the second target object is the orientation information of the second target object.

[0114] For example, in this embodiment, a Cartesian coordinate system can be established with the center of the first target object as the origin. This Cartesian coordinate system is denoted as the first coordinate system. The position information of the first target object can be its position information in the Cartesian coordinate system. The attitude information of the second target object is its orientation information in the Cartesian coordinate system, such as the head orientation.

[0115] Optionally, spatial audio is generated based on the position information of the first target object, the posture information of the second target object, the mono audio, and the radial velocity of the first target object relative to the second target object. Generating spatial audio includes: inputting the position information of the first target object, the posture information of the second target object, the mono audio, and the radial velocity of the first target object relative to the second target object into a pre-trained machine learning model, so that the machine learning model outputs the spatial audio.

[0116] For example, server 22 stores the trained machine learning model as described above. Server 22 can input the position information of the first target object, the pose information of the second target object, the mono audio, and the radial velocity of the first target object relative to the second target object into the trained machine learning model, causing the machine learning model to output spatial audio. Since the trained machine learning model is an accurate model, the spatial audio output by the trained machine learning model is accurate spatial audio, that is, spatial audio that is infinitely close to the actual two-channel audio.

[0117] This embodiment acquires the mono audio corresponding to a moving first target object and calculates the radial velocity of the first target object relative to a second target object. Further, spatial audio is generated based on the position information of the first target object, the posture information of the second target object, the mono audio, and the radial velocity of the first target object relative to the second target object. This ensures that the generated spatial audio contains the motion information of the first target object, thereby matching the generated spatial audio with the actual stereo audio perceived by the second target object when the mono audio corresponding to the first target object reaches it. It also minimizes the difference between the phase of the audio signal in the spatial audio and the phase of the actual stereo audio, improving the accuracy of the phase calculation for the left and right ear channels in the spatial audio, and greatly enhancing the accuracy of the generated spatial audio. Therefore, when a user hears this spatial audio, they can accurately determine the position of the first target object, such as the sound source, based on the phase difference between the left and right ear audio signals in the spatial audio, and perceive the movement of the first target object, thus providing an immersive experience and greatly improving the user experience.

[0118] In some cases, the method described in this embodiment can also enable users to achieve consistency between auditory and visual perception of the first target object. Furthermore, the method described in this embodiment does not require introducing additional hyperparameters into the machine learning model, nor does it require modification of the loss function. Instead, it can be well extended to different types of machine learning models, such as... Figure 4 The twisted network shown Figure 5 The dual-channel gradient network shown is an example. Therefore, the method described in this embodiment can achieve a plug-and-play effect.

[0119] Optionally, the radial velocity of the first target object relative to the second target object includes the radial velocity of the first target object relative to the first pickup unit of the second target object and the radial velocity of the first target object relative to the second pickup unit of the second target object.

[0120] For example, the first pickup unit of the second target object can be the right ear of the second target object, or an audio acquisition device worn on the right ear. The second pickup unit of the second target object can be the left ear of the second target object, or an audio acquisition device worn on the left ear. The radial velocity of the first target object relative to the second target object includes the radial velocity of the first target object relative to the first pickup unit of the second target object and the radial velocity of the first target object relative to the second pickup unit of the second target object.

[0121] Optionally, calculating the radial velocity of the first target object relative to the second target object includes the following steps:

[0122] S701. In the first coordinate system with the second target object as the origin, calculate the position vector of the first pickup unit of the first target object relative to the second target object.

[0123] For example, after establishing a Cartesian coordinate system (the first coordinate system) with the center of the head of the second target object as the origin, the three-dimensional (3D) position of the first target object in this Cartesian coordinate system is denoted as... The first pickup unit of the second target object is, for example, the right ear of the second target object, and the three-dimensional position of the right ear in the Cartesian coordinate system is denoted as . The position vector of the first target object relative to the right ear is denoted as... It is calculated using the following formula (1):

[0124]

[0125] in, and P represents the three-dimensional coordinates respectively. x express x-axis coordinates and The position vector of the x-axis coordinate, P y express y-axis coordinates and The position vector of the y-axis coordinate, P z express z-axis coordinates and The position vector of the z-axis coordinate.

[0126] S702. Calculate the moving speed of the first target object based on the position vector.

[0127] For example, based on the position vector Calculate the movement speed of the first target object It is calculated using the following formula (2):

[0128]

[0129] in, P represents x Differentiate with respect to time, P represents y Differentiate with respect to time, P represents z Differentiate with respect to time.

[0130] S703. In the second coordinate system with the first pickup unit as the origin, the moving speed of the first target object is decomposed into the radial speed of the first target object relative to the first pickup unit.

[0131] For example, a spherical coordinate system is established with the right ear as the origin, and this spherical coordinate system is denoted as the second coordinate system. In this second coordinate system, the following formula (3) is used to... Decomposed into the radial velocity of the first target object relative to the right ear.

[0132]

[0133] in, This represents the radial unit vector in a spherical coordinate system. Specifically, it could be v as described above. r-right Similarly, v can be calculated as described above. r-left .

[0134] exist Figure 4 or Figure 5 In the illustrated embodiment, C0∈r 7 = (x, y, z, qx, qy, qz, qw), where (x, y, z) represents the position of the first target object in the Cartesian coordinate system. (qx, qy, qz, qw) represents the head direction of the second target object in the Cartesian coordinate system, that is, the head direction is represented by the quadruple (qx, qy, qz, qw). Since this embodiment introduces v... r-left and v r-right Therefore, the conditions used by machine learning models to generate spatial audio can be a new condition C.

[0135]

[0136] Since this embodiment can calculate the radial velocity of the first pickup unit of the first target object relative to the second target object and the radial velocity of the first target object relative to the second pickup unit of the second target object, and provide these two radial velocities as conditions to the machine learning model, the spatial audio generated by the machine learning model can contain the motion information of the first target object. That is, the machine learning model can perceive the motion information of the first target object when generating spatial audio. Experiments have verified that the method described in this embodiment can improve various accuracies, especially the phase accuracy (Phase L2) of the audio signal in the spatial audio generated by the machine learning model. Currently, Phase L2 can reach 0.780, which is a relatively good spatial audio phase loss in the industry. This loss can be understood as the difference between the phase of the audio signal in the spatial audio and the phase of the actual two-channel audio.

[0137] It is understood that the audio processing method provided in this embodiment can be applied to different application scenarios, such as virtual meetings, virtual assistant experiences, augmented reality, and virtual reality. The following describes different application scenarios with specific embodiments.

[0138] Figure 8 This is a flowchart illustrating an audio processing method according to another embodiment of the present disclosure. For example, this method can be applied to game scenes or metaverse scenes, and can be specifically executed by a cloud server. In this embodiment, the specific steps of the method are as follows:

[0139] S801. Calculate the radial velocity of the first target object relative to the second target object based on the position information of the first movable target object in the virtual space relative to the second target object in the virtual space.

[0140] For example, the first target object and the second target object are simultaneously located in the virtual space of the game scene or the virtual space of the metaverse scene, meaning the first target object and the second target object are in the same virtual space. In this virtual space, assume the first target object is a movable sound source, and the second target object is stationary. Since the positions of the first and second target objects in this virtual space are known, the cloud server can establish a Cartesian coordinate system with the second target object as the origin, and map the positions of the first and second target objects in this virtual space to the Cartesian coordinate system respectively, thus obtaining the positions of the first and second target objects in the Cartesian coordinate system. Further, based on the positions of the first and second target objects in the Cartesian coordinate system, the position vector of the first target object relative to the second target object is calculated, and the radial velocity of the first target object relative to the second target object is calculated based on this position vector; for example, the radial velocities of the first target object relative to the left and right ears of the second target object, respectively. The specific calculation process is as described above and will not be repeated here.

[0141] S802. Generate spatial audio based on the position information of the first target object, the posture information of the second target object, the mono audio pre-configured for the first target object, and the radial velocity of the first target object relative to the second target object.

[0142] For example, after the cloud server maps the positions of the first and second target objects in the virtual space to a Cartesian coordinate system, it can obtain the position information of the first target object in the Cartesian coordinate system and the head orientation of the second target object in the Cartesian coordinate system. Furthermore, this position information, head orientation, a pre-configured mono audio track for the first target object, and the radial velocity of the first target object relative to the second target object are input into a trained machine learning model, causing the machine learning model to output spatial audio.

[0143] S803, Play the spatial audio.

[0144] For example, a cloud server can send this spatial audio to a terminal that presents a game scene or metaverse scene. The terminal can then play the spatial audio, allowing the user to perceive the movement of the primary target object within the game scene or metaverse scene based on the spatial audio. This enhances the user's gaming experience or their experience with the metaverse scene.

[0145] Figure 9 This is a flowchart illustrating an audio processing method according to another embodiment of the present disclosure. For example, this method can be applied to virtual reality, and it can be specifically executed by a cloud server. In this embodiment, the specific steps of the method are as follows:

[0146] S901. Calculate the radial velocity of the movable target relative to the target user based on the position information of the movable target relative to the target user displayed in the virtual reality display device, wherein the target user is wearing the virtual reality display device.

[0147] For example Figure 10 As shown, target user 100 wears a virtual reality (VR) display device 101, such as VR glasses. The VR glasses include a display screen 102, which displays a virtual reality scene, for example, the scene includes a movable target 103. For example, a Cartesian coordinate system is established with target user 100 as the origin. Figure 10 The x-axis shown can be the x-axis of this Cartesian coordinate system, such as... Figure 10 The y-axis shown is the y-axis of this Cartesian coordinate system. In this Cartesian coordinate system, the position vector of the movable target 103 relative to the target user 100 can be calculated. Furthermore, based on this position vector, the radial velocity of the movable target 103 relative to the target user 100 can be calculated, for example, the radial velocities of the movable target 103 relative to the left and right ears of the target user 100, respectively.

[0148] S902. Generate spatial audio based on the position information of the movable target, the posture information of the target user, the mono audio pre-configured for the movable target, and the radial velocity of the movable target relative to the target user.

[0149] In addition, such as Figure 10 As shown, the position information of the movable target 103 in the Cartesian coordinate system and the head orientation of the target user 100 in the Cartesian coordinate system can also be calculated. The cloud server can input the position information, head orientation, pre-configured mono audio for the movable target 103, and the radial velocities of the movable target 103 relative to the left and right ears of the target user 100 into the trained machine learning model, so that the machine learning model outputs spatial audio.

[0150] S903. Play the spatial audio through the virtual reality display device.

[0151] For example, a cloud server can send the spatial audio to a virtual reality display device 101, which can then play the spatial audio, allowing the target user 100 to perceive the movement of the movable target 103 based on the spatial audio while viewing the virtual reality scene. This provides the target user with a more immersive experience.

[0152] Figure 11This is a flowchart illustrating an audio processing method according to another embodiment of the present disclosure. For example, this method can be applied to augmented reality scenarios, and it can be specifically executed by a cloud server. In this embodiment, the specific steps of the method are as follows:

[0153] S1101, Acquire mono audio emitted by a movable target.

[0154] For example Figure 12 As shown, the movable target 121 may be equipped with an audio acquisition device, which can acquire mono audio emitted by the movable target 121 during its movement. This audio acquisition device can then send the acquired mono audio to a cloud server.

[0155] S1102. Calculate the radial velocity of the movable target relative to the target user based on the position information of the movable target relative to the target user, wherein the target user is wearing an augmented reality device.

[0156] For example, target user 122 wears an augmented reality (AR) device 123, such as AR glasses. The AR glasses can measure the position information of the movable target 121 relative to target user 122, and then convert the position information into a position vector in a Cartesian coordinate system with target user 122 as the origin. Based on the position vector, the radial velocity of the movable target 121 relative to target user 122 is calculated, for example, the radial velocity of the movable target 121 relative to the left and right ears of target user 122, respectively.

[0157] S1103. Generate spatial audio based on the position information of the movable target, the posture information of the target user, the mono audio, and the radial velocity of the movable target relative to the target user.

[0158] like Figure 12 As shown, the AR glasses are equipped with sensors such as gyroscopes, enabling them to sense the head orientation of the target user 122 in the Cartesian coordinate system. Furthermore, the AR glasses send the position information of the movable target 121 in the Cartesian coordinate system, the head orientation of the target user 122 in the Cartesian coordinate system, and the radial velocities of the movable target 121 relative to the left and right ears of the target user 122, respectively, to the cloud server. This allows the cloud server to input the position information, head orientation, mono audio, and the radial velocities of the movable target 121 relative to the left and right ears of the target user 122 into a trained machine learning model, causing the machine learning model to output spatial audio.

[0159] S1104. Play the spatial audio through the augmented reality device.

[0160] For example, the cloud server can send the spatial audio to the AR glasses, which can then play the spatial audio, allowing the target user 122 to perceive the movement of the movable target 121 relative to the target user 122 while watching the game, thus enhancing the augmented reality experience.

[0161] Figure 13 This is a flowchart illustrating an audio processing method according to another embodiment of the present disclosure. For example, this method can be applied to a meeting scenario within the same space, and it can be executed by a cloud server. In this embodiment, the specific steps of the method are as follows:

[0162] S1301, Collect the mono audio emitted by the first user in motion.

[0163] For example Figure 14 As shown, the first user 141 and the second user 142 are located in the same space, which is denoted as the preset space 143. The first user 141 is in motion, while the second user 142 is stationary. The first user 141 may wear an audio acquisition device that can acquire the mono audio emitted by the first user and send the mono audio to the cloud server.

[0164] S1302. Calculate the radial velocity of the first user relative to the second user in the preset space based on the position information of the first user relative to the second user in the preset space.

[0165] For example, a camera, such as a recording device, is installed in the preset space 143. This recording device can capture images of the first user 141 and the second user 142 in real time. The recording device can send the captured images to a cloud server, allowing the cloud server to determine the position information of the first user 141 relative to the second user 142 based on the images. Furthermore, the cloud server can establish a Cartesian coordinate system with the second user 142 as the origin. The position information of the first user 141 relative to the second user 142 is then converted into a position vector in the Cartesian coordinate system. Based on this position vector, the radial velocity of the first user 141 relative to the second user 142 is calculated; for example, the radial velocities of the first user 141 relative to the left and right ears of the second user 142.

[0166] S1303. Generate spatial audio based on the position information of the first user, the posture information of the second user, the mono audio, and the radial velocity of the first user relative to the second user.

[0167] For example, the cloud server can input the position information of the first user 141 in the Cartesian coordinate system, the head direction of the second user 142 in the Cartesian coordinate system, the mono audio, and the radial velocities of the first user 141 relative to the left and right ears of the second user 142 into the trained machine learning model, so that the machine learning model outputs spatial audio.

[0168] S1304. The spatial audio is sent to the audio playback device worn by the second user.

[0169] For example, the cloud server can send the spatial audio to an audio playback device, such as a headset, worn by the second user 142. This audio playback device then plays the spatial audio, allowing the second user 142 to perceive the movements of the first user 141, thus improving the user experience in a meeting setting.

[0170] Figure 15 This is a flowchart illustrating an audio processing method according to another embodiment of the present disclosure. For example, this method can be applied to remote conferencing scenarios, and it can be specifically executed by a cloud server. In this embodiment, the specific steps of the method are as follows:

[0171] S1501, Collect the mono audio emitted by the first user in motion.

[0172] like Figure 16 As shown, the first user 161 is located in the first space 162, and the second user 163 is located in the second space 164, meaning that the first user 161 and the second user 163 are located in different spaces. The first user 161 is in motion in the first space 162. The second user 163 is stationary in the second space 164. The first user 161 may wear an audio acquisition device that can acquire the mono audio emitted by the first user and send the mono audio to a cloud server.

[0173] S1502. Calculate the radial velocity of the first user relative to the target position based on the position information of the first user in the preset space relative to the target position.

[0174] For example, a first space 162 is denoted as a preset space, within which a target position 165 exists. The coordinates of the target position 165 within this preset space can be pre-stored in a cloud server. Additionally, a camera, such as a video camera, is installed in the preset space. This camera can capture images of the first user 161 in real time and send the captured images to the cloud server. The cloud server can then determine the position information of the first user 161 relative to the target position 165 based on the images of the first user 161. Furthermore, the cloud server can establish a Cartesian coordinate system with the target position 165 as the origin. The position information of the first user 161 relative to the target position 165 is then converted into a position vector in the Cartesian coordinate system. The radial velocity of the first user 161 relative to the target position 165 is calculated based on this position vector.

[0175] S1503. Generate spatial audio based on the position information of the first user, the posture information of the second user, the mono audio, and the radial velocity of the first user relative to the target position.

[0176] For example in Figure 16 In the second space 164, a camera can be configured to capture images of the second user 163 and send the captured images to a cloud server, allowing the cloud server to determine the head orientation of the second user 163 based on the images. Alternatively, a gyroscope can be installed on the headset worn by the second user 163, enabling the headset to detect the head orientation of the second user 163 and send the head orientation to the cloud server. Furthermore, the cloud server can input the position information of the first user 161 in the Cartesian coordinate system, the head orientation of the second user 163, mono audio, and the radial velocity of the first user 161 relative to the target position 165 into a trained machine learning model, causing the machine learning model to output spatial audio.

[0177] S1504. The spatial audio is sent to the audio playback device worn by the second user, where the second user and the first user are located in different spaces.

[0178] For example, a cloud server can send this spatial audio to the audio playback device worn by the second user 163, such as a headset. This enables remote conferencing, allowing the remote user in the conferencing, such as the second user 163, to perceive the movement of the speaker, such as the first user, through the spatial audio, providing the second user with an immersive experience, that is, giving the second user the feeling that the second user and the first user are in the same space.

[0179] Figure 17This is a schematic diagram of the structure of an audio processing apparatus provided in an embodiment of this disclosure. The audio processing apparatus provided in this embodiment of the disclosure can execute the processing flow provided in the audio processing method embodiment, such as... Figure 17 As shown, the audio processing device 170 includes:

[0180] Acquisition module 171 is used to acquire the mono audio corresponding to the first target object in motion;

[0181] Calculation module 172 is used to calculate the radial velocity of the first target object relative to the second target object;

[0182] The generation module 173 is used to generate spatial audio based on the position information of the first target object, the posture information of the second target object, the mono audio, and the radial velocity of the first target object relative to the second target object.

[0183] Optionally, the generation module 173 generates spatial audio based on the position information of the first target object, the posture information of the second target object, the mono audio, and the radial velocity of the first target object relative to the second target object. Specifically, when generating spatial audio, it is used for:

[0184] The position information of the first target object, the posture information of the second target object, the mono audio, and the radial velocity of the first target object relative to the second target object are input into a pre-trained machine learning model, so that the machine learning model outputs the spatial audio.

[0185] Optionally, the position information of the first target object is the position information of the first target object in a first coordinate system, wherein the first coordinate system is a coordinate system with the second target object as the origin;

[0186] The pose information of the second target object is the orientation information of the second target object.

[0187] Optionally, the radial velocity of the first target object relative to the second target object includes the radial velocity of the first target object relative to the first pickup unit of the second target object and the radial velocity of the first target object relative to the second pickup unit of the second target object.

[0188] Optionally, when calculating the radial velocity of the first target object relative to the second target object, the calculation module 172 is specifically used for:

[0189] In a first coordinate system with the second target object as the origin, calculate the position vector of the first pickup unit of the first target object relative to the second target object;

[0190] Calculate the moving speed of the first target object based on the position vector;

[0191] In a second coordinate system with the first pickup unit as the origin, the moving speed of the first target object is decomposed into the radial speed of the first target object relative to the first pickup unit.

[0192] Figure 17 The audio processing apparatus of the illustrated embodiment can be used to execute the technical solutions of the above method embodiments. Its implementation principle and technical effect are similar, and will not be repeated here.

[0193] The above describes the internal functions and structure of an audio processing device, which can be implemented as an electronic device. Figure 18 A schematic diagram illustrating the structure of an electronic device embodiment provided in this disclosure. (See attached diagram.) Figure 18 As shown, the electronic device includes a memory 181 and a processor 182.

[0194] Memory 181 is used to store programs. In addition to the programs described above, memory 181 can also be configured to store various other data to support operation on the electronic device. Examples of this data include instructions for any application or method used to operate on the electronic device, contact data, phone book data, messages, pictures, videos, etc.

[0195] The memory 181 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic storage, flash memory, magnetic disk or optical disk.

[0196] The processor 182 is coupled to the memory 181 and executes the program stored in the memory 181 for:

[0197] Obtain the mono audio corresponding to the first target object in motion;

[0198] Calculate the radial velocity of the first target object relative to the second target object;

[0199] Spatial audio is generated based on the position information of the first target object, the posture information of the second target object, the mono audio, and the radial velocity of the first target object relative to the second target object.

[0200] Furthermore, such as Figure 18 As shown, the electronic device may also include other components such as a communication component 183, a power supply component 184, an audio component 185, and a display 186. Figure 18The diagram only shows some components and does not mean that the electronic device includes only these components. Figure 18 The components shown.

[0201] Communication component 183 is configured to facilitate wired or wireless communication between electronic devices and other devices. The electronic devices can access wireless networks based on communication standards, such as WiFi, 2G, or 3G, or combinations thereof. In one exemplary embodiment, communication component 183 receives broadcast signals or broadcast-related information from an external broadcast management system via a broadcast channel. In one exemplary embodiment, communication component 183 also includes a near-field communication (NFC) module to facilitate short-range communication. For example, the NFC module may be implemented based on radio frequency identification (RFID) technology, Infrared Data Association (IrDA) technology, ultra-wideband (UWB) technology, Bluetooth (BT) technology, and other technologies.

[0202] Power supply component 184 provides power to various components of an electronic device. Power supply component 184 may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power to the electronic device.

[0203] Audio component 185 is configured to output and / or input audio signals. For example, audio component 185 includes a microphone (MIC) configured to receive external audio signals when the electronic device is in an operating mode, such as call mode, recording mode, and voice recognition mode. The received audio signals may be further stored in memory 181 or transmitted via communication component 183. In some embodiments, audio component 185 also includes a speaker for outputting audio signals.

[0204] Display 186 includes a screen, which may include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen can be implemented as a touchscreen to receive input signals from a user. The touch panel includes one or more touch sensors to sense touches, swipes, and gestures on the touch panel. The touch sensors can sense not only the boundaries of the touch or swipe action but also the duration and pressure associated with the touch or swipe operation.

[0205] In addition, this disclosure also provides a computer-readable storage medium storing a computer program thereon, which is executed by a processor to implement the audio processing method and model training method described in the above embodiments.

[0206] It should be noted that, in this document, relational terms such as "first" and "second" are used merely to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0207] The above description is merely a specific embodiment of this disclosure, enabling those skilled in the art to understand or implement it. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of this disclosure. Therefore, this disclosure is not to be limited to the embodiments described herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. An audio processing method, wherein, The method includes: Obtain the mono audio corresponding to the first target object in motion; Calculate the radial velocity of the first target object relative to the second target object, wherein the radial velocity includes the radial velocity of the first target object relative to the first pickup unit of the second target object and the radial velocity of the first target object relative to the second pickup unit of the second target object; Spatial audio is generated based on the position information of the first target object, the posture information of the second target object, the mono audio, and the radial velocity of the first target object relative to the second target object, wherein the posture information of the second target object is the orientation information of the second target object; The calculation of the radial velocity of the first target object relative to the second target object includes: calculating the position vector of the first target object relative to the first pickup unit of the second target object in a first coordinate system with the second target object as the origin; calculating the moving velocity of the first target object based on the position vector; and decomposing the moving velocity of the first target object into the radial velocity of the first target object relative to the first pickup unit in a second coordinate system with the first pickup unit as the origin.

2. The method according to claim 1, wherein, The spatial audio is generated based on the position information of the first target object, the attitude information of the second target object, the mono audio, and the radial velocity of the first target object relative to the second target object, including: The position information of the first target object, the posture information of the second target object, the mono audio, and the radial velocity of the first target object relative to the second target object are input into a pre-trained machine learning model, so that the machine learning model outputs the spatial audio.

3. The method according to claim 1, wherein, The position information of the first target object is the position information of the first target object in the first coordinate system, and the first coordinate system is the coordinate system with the second target object as the origin.

4. An audio processing method, wherein, The method includes: Based on the position information of a movable first target object in the virtual space relative to a second target object in the virtual space, the radial velocity of the first target object relative to the second target object is calculated. The radial velocity includes the radial velocity of the first target object relative to a first pickup unit of the second target object and the radial velocity of the first target object relative to a second pickup unit of the second target object. Calculating the radial velocity of the first target object relative to the second target object includes: calculating the position vector of the first target object relative to the first pickup unit of the second target object in a first coordinate system with the second target object as the origin; calculating the movement velocity of the first target object based on the position vector; and decomposing the movement velocity of the first target object into the radial velocity of the first target object relative to the first pickup unit in a second coordinate system with the first pickup unit as the origin. Spatial audio is generated based on the position information of the first target object, the posture information of the second target object, the mono audio pre-configured for the first target object, and the radial velocity of the first target object relative to the second target object, wherein the posture information of the second target object is the orientation information of the second target object; Play the spatial audio.

5. An audio processing method, wherein, The method includes: Based on the position information of a movable target relative to a target user displayed in a virtual reality display device, the radial velocity of the movable target relative to the target user is calculated. The target user is wearing the virtual reality display device. The radial velocity includes the radial velocity of the movable target relative to a first microphone unit of the target user and the radial velocity of the movable target relative to a second microphone unit of the target user. Calculating the radial velocity of the movable target relative to the target user includes: calculating the position vector of the movable target relative to the first microphone unit of the target user in a first coordinate system with the target user as the origin; calculating the movement speed of the movable target based on the position vector; and decomposing the movement speed of the movable target into the radial velocity of the movable target relative to the first microphone unit in a second coordinate system with the first microphone unit as the origin. Spatial audio is generated based on the position information of the movable target, the posture information of the target user, the mono audio pre-configured for the movable target, and the radial velocity of the movable target relative to the target user, wherein the posture information of the target user is the orientation information of the target user; The spatial audio is played through the virtual reality display device.

6. An audio processing method, wherein, The method includes: Acquire mono audio emitted by a moving target; Based on the position information of the movable target relative to the target user, the radial velocity of the movable target relative to the target user is calculated. The target user is wearing an augmented reality device. The radial velocity includes the radial velocity of the movable target relative to a first microphone unit of the target user and the radial velocity of the movable target relative to a second microphone unit of the target user. Calculating the radial velocity of the movable target relative to the target user includes: calculating the position vector of the movable target relative to the first microphone unit of the target user in a first coordinate system with the target user as the origin; calculating the movement speed of the movable target based on the position vector; and decomposing the movement speed of the movable target into the radial velocity of the movable target relative to the first microphone unit in a second coordinate system with the first microphone unit as the origin. Spatial audio is generated based on the position information of the movable target, the posture information of the target user, the mono audio, and the radial velocity of the movable target relative to the target user, wherein the posture information of the target user is the orientation information of the target user; The spatial audio is played through the augmented reality device.

7. An audio processing method, wherein, The method includes: Collect mono audio emitted by the first user during movement; Based on the position information of the first user relative to the second user in the preset space, the radial velocity of the first user relative to the second user is calculated. The radial velocity includes the radial velocity of the first user relative to the first pickup unit of the second user and the radial velocity of the first user relative to the second user's second pickup unit. Calculating the radial velocity of the first user relative to the second user includes: calculating the position vector of the first user relative to the first pickup unit of the second user in a first coordinate system with the second user as the origin; calculating the movement speed of the first user based on the position vector; and decomposing the movement speed of the first user into the radial velocity of the first user relative to the first pickup unit in a second coordinate system with the first pickup unit as the origin. Spatial audio is generated based on the position information of the first user, the posture information of the second user, the mono audio, and the radial velocity of the first user relative to the second user, wherein the posture information of the second user is the orientation information of the second user; The spatial audio is sent to the audio playback device worn by the second user.

8. An audio processing method, wherein, The method includes: Collect mono audio emitted by the first user during movement; Based on the position information of the first user in the preset space relative to the target position in the preset space, the radial velocity of the first user relative to the target position is calculated. The radial velocity includes the radial velocity of the first user relative to a first pickup unit at the target position and the radial velocity of the first user relative to a second pickup unit at the target position. Calculating the radial velocity of the first user relative to the target position includes: calculating the position vector of the first user relative to the first pickup unit at the target position in a first coordinate system with the target position as the origin; calculating the movement speed of the first user based on the position vector; and decomposing the movement speed of the first user into the radial velocity of the first user relative to the first pickup unit in a second coordinate system with the first pickup unit as the origin. Spatial audio is generated based on the location information of the first user, the posture information of the second user, the mono audio, and the radial velocity of the first user relative to the target position, wherein the posture information of the second user is the orientation information of the second user; The spatial audio is sent to the audio playback device worn by the second user, who is located in a different space from the first user.

9. A model training method, wherein, The method includes: Acquire mono audio emitted by a movable sound source; Acquire the two-channel audio collected by the first and second pickup units of the target object; The position information of the movable sound source, the posture information of the target object, the mono audio, and the radial velocity of the movable sound source relative to the target object are input into the machine learning model to be trained, so that the machine learning model outputs spatial audio. The radial velocity includes the radial velocity of the movable sound source relative to the first pickup unit of the target object and the radial velocity of the movable sound source relative to the second pickup unit of the target object. The posture information of the target object is the orientation information of the target object. The calculation of the radial velocity of the movable sound source relative to the target object includes: calculating the position vector of the movable sound source relative to the first pickup unit of the target object in a first coordinate system with the target object as the origin; calculating the moving speed of the movable sound source based on the position vector; and decomposing the moving speed of the movable sound source into the radial velocity of the movable sound source relative to the first pickup unit in a second coordinate system with the first pickup unit as the origin. The machine learning model is trained based on the spatial audio and the dual-channel audio, and the trained machine learning model is used to perform the audio processing method as described in any one of claims 1-8.

10. An audio processing apparatus, wherein, include: The acquisition module is used to acquire the mono audio corresponding to the first target object in motion; A calculation module is used to calculate the radial velocity of the first target object relative to the second target object, wherein the radial velocity includes the radial velocity of the first target object relative to the first pickup unit of the second target object and the radial velocity of the first target object relative to the second pickup unit of the second target object. Calculating the radial velocity of the first target object relative to the second target object includes: calculating the position vector of the first target object relative to the first pickup unit of the second target object in a first coordinate system with the second target object as the origin; calculating the moving velocity of the first target object based on the position vector; and decomposing the moving velocity of the first target object into the radial velocity of the first target object relative to the first pickup unit in a second coordinate system with the first pickup unit as the origin. The generation module is used to generate spatial audio based on the position information of the first target object, the posture information of the second target object, the mono audio, and the radial velocity of the first target object relative to the second target object, wherein the posture information of the second target object is the orientation information of the second target object.

11. An electronic device, wherein, include: Memory; processor; as well as Computer programs; The computer program is stored in the memory and configured to be executed by the processor to implement the method as described in any one of claims 1-9.

12. A computer-readable storage medium having a computer program stored thereon, wherein, When the computer program is executed by a processor, it implements the method as described in any one of claims 1-9.

Citation Information

Patent Citations

  • Audio processing method and device

    CN114598985A

  • System and method for determining audio context in augmented-reality applications

    US20170208415A1