Multi-audio playing method and device, electronic equipment and storage medium

By determining the scene category, focus audio and non-focus audio in multi-audio stream playback, and allocating virtual space locations to play audio data, the problems of confusion and incomplete content in the existing multi-audio stream playback methods are solved, and the clear and complete playback of audio content is achieved, and the multi-audio playback effect is improved.

CN119946544APending Publication Date: 2025-05-06GUANGDONG OPPO MOBILE TELECOMMUNICATIONS CORP LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311452289.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-11-02
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

The existing multi-audio streaming method has problems such as confusing sound or incomplete content, which cannot take into account the clarity and integrity of the audio content, resulting in poor playback results.

Method used

By determining the scene categories corresponding to the audio data, dividing focus audio and non-focus audio, and assigning virtual spatial locations based on these audio data for playback, the clear and complete playback of the audio data is achieved.

Benefits of technology

This method can take into account the clarity and integrity of the audio content, improve the effect of multi-audio playback, and ensure that different audio data do not interfere with each other in the virtual space to play.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119946544A_ABST
    Figure CN119946544A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-audio playing method and device, electronic equipment and a storage medium. The method comprises the following steps: determining a first scene category corresponding to first audio data and a second scene category corresponding to second audio data; determining a first focus audio and a first non-focus audio in the first audio data and the second audio data according to the first scene category and the second scene category; determining a first virtual space position corresponding to the first audio data and a second virtual space position corresponding to the second audio data based on the first focus audio and the first non-focus audio; and playing the first audio data based on the first virtual space position, and playing the second audio data based on the second virtual space position.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of audio playback, and in particular to a multi-audio playback method and device, electronic equipment and storage medium. Background Art

[0002] At present, there are two common methods for playing multiple audio streams. One is to directly linearly superimpose different audio stream data and then play them out. The other is to selectively duck certain audio streams (i.e., lower the volume) to highlight the audio streams that have not been ducked.

[0003] However, the common method of playing multiple audio streams has the problem of confusing sound or incomplete content, which cannot take into account the clarity and integrity of the audio content and leads to poor playback effect. Summary of the invention

[0004] The embodiments of the present application provide a multi-audio playback method and device, an electronic device, and a storage medium, which can take into account both the clarity and integrity of the audio content and effectively improve the multi-audio playback effect.

[0005] The technical solution of the embodiment of the present application is implemented as follows:

[0006] In a first aspect, an embodiment of the present application provides a method for playing multiple audios, the method comprising:

[0007] Determine a first scene category corresponding to the first audio data and a second scene category corresponding to the second audio data;

[0008] determining first focus audio and first non-focus audio in the first audio data and the second audio data according to the first scene category and the second scene category;

[0009] Determining, based on the first focused audio and the first non-focused audio, a first virtual space position corresponding to the first audio data and a second virtual space position corresponding to the second audio data;

[0010] The first audio data is played based on the first virtual space position, and the second audio data is played based on the second virtual space position.

[0011] In a second aspect, an embodiment of the present application provides a multi-audio playback device, the multi-audio playback device comprising: a determination unit, a playback unit,

[0012] The determining unit is used to determine a first scene category corresponding to the first audio data and a second scene category corresponding to the second audio data; determine a first focused audio and a first non-focused audio in the first audio data and the second audio data according to the first scene category and the second scene category; and determine a first virtual space position corresponding to the first audio data and a second virtual space position corresponding to the second audio data based on the first focused audio and the first non-focused audio;

[0013] The playing unit is used to play the first audio data based on the first virtual space position, and play the second audio data based on the second virtual space position.

[0014] In a third aspect, an embodiment of the present application provides an electronic device, comprising a processor and a memory storing instructions executable by the processor, wherein when the instructions are executed by the processor, the method described in the first aspect above is implemented.

[0015] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium having a program stored thereon, and when the program is executed by a processor, the method described in the first aspect above is implemented.

[0016] The embodiment of the present application provides a multi-audio playback method and device, electronic device and storage medium, determining the first scene category corresponding to the first audio data and the second scene category corresponding to the second audio data; determining the first focus audio and the first non-focus audio in the first audio data and the second audio data according to the first scene category and the second scene category; determining the first virtual space position corresponding to the first audio data and the second virtual space position corresponding to the second audio data based on the first focus audio and the first non-focus audio; playing the first audio data based on the first virtual space position, and playing the second audio data based on the second virtual space position. It can be seen that in the embodiment of the present application, in the process of multi-audio parallel playback, the scene categories corresponding to different audio data can be used to determine the focus audio and non-focus audio in multiple audio data, so that each audio data can be assigned a corresponding virtual space position based on the focus audio and the non-focus audio, so that different audio data can be played at different virtual space positions. That is to say, the multi-audio playback method proposed in the embodiment of the present application can not only play each audio data completely, but also assign non-interfering virtual space positions to different audio data for rendering and playback, so as to take into account the clarity and integrity of the audio content and effectively improve the playback effect of multi-audio. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] Figure 1 Schematic diagram of the implementation process of the multi-audio playback method proposed in the embodiment of the present application Figure 1;

[0018] Figure 2 A schematic diagram of the correspondence between scenes and non-focus positions proposed in an embodiment of the present application;

[0019] Figure 3 Schematic diagram of the preset center position proposed in the embodiment of the present application Figure 1 ;

[0020] Figure 4 Schematic diagram of the preset center position proposed in the embodiment of the present application Figure 2 ;

[0021] Figure 5 Schematic diagram of the virtual space position proposed in the embodiment of the present application Figure 1 ;

[0022] Figure 6 Schematic diagram of the virtual space position proposed in the embodiment of the present application Figure 2 ;

[0023] Figure 7 Schematic diagram of the implementation of the gradual entry method proposed in the embodiment of the present application Figure 1 ;

[0024] Figure 8 Schematic diagram of the implementation of the gradual entry method proposed in the embodiment of the present application Figure 2 ;

[0025] Fig. 9 Schematic diagram of the implementation process of the multi-audio playback method proposed in the embodiment of the present application Figure 2 ;

[0026] Fig.10 Schematic diagram of the implementation of the gradual entry method proposed in the embodiment of the present application Figure 3 ;

[0027] Fig.11 Schematic diagram of the implementation of the gradual exit method proposed in the embodiment of the present application Figure 1 ;

[0028] Fig.12 Schematic diagram of the implementation of switching focus audio proposed in the embodiment of the present application Figure 1 ;

[0029] Fig.13 Schematic diagram of the implementation of the gradual exit method proposed in the embodiment of the present application Figure 2 ;

[0030] Fig.14 Schematic diagram of the implementation of switching focus audio proposed in the embodiment of the present application Figure 2 ;

[0031] Fig.15Schematic diagram of the implementation process of the multi-audio playback method proposed in the embodiment of the present application Figure 3 ;

[0032] Fig.16 Schematic diagram of the implementation process of the multi-audio playback method proposed in the embodiment of the present application Figure 4 ;

[0033] Fig.17 Schematic diagram of the implementation of switching focus audio proposed in the embodiment of the present application Figure 3 ;

[0034] Fig.18 Schematic diagram of the implementation process of the multi-audio playback method proposed in the embodiment of the present application Figure 5 ;

[0035] Fig.19 Schematic diagram of the implementation process of the multi-audio playback method proposed in the embodiment of the present application Figure 6 ;

[0036] Fig. 20 Schematic diagram of the implementation process of the multi-audio playback method proposed in the embodiment of the present application Figure 7 ;

[0037] Fig.21 Schematic diagram of the implementation of switching focus audio proposed in the embodiment of the present application Figure 4 ;

[0038] Fig. 22 A schematic diagram of the transformation of the motion state of audio data proposed in an embodiment of the present application;

[0039] Fig.23 Schematic diagram of the implementation of multi-audio playback proposed in the embodiment of the present application Figure 1 ;

[0040] Fig.24 Schematic diagram of the implementation of multi-audio playback proposed in the embodiment of the present application Figure 2 ;

[0041] Fig.25 Schematic diagram of the implementation of multi-audio playback proposed in the embodiment of the present application Figure 3 ;

[0042] Fig.26 Schematic diagram of the implementation of multi-audio playback proposed in the embodiment of the present application Figure 4 ;

[0043] Fig. 27 A schematic diagram of the structure of a multi-audio playback device proposed in an embodiment of the present application;

[0044] Fig.28 A schematic diagram of the structure of an electronic device according to an embodiment of the present application. DETAILED DESCRIPTION

[0045] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. It is understood that the specific embodiments described herein are only used to explain the related applications, rather than to limit the applications. It should also be noted that, for ease of description, only the parts related to the related applications are shown in the drawings.

[0046] With the continuous development of terminal devices, the functions of terminal devices are becoming more and more abundant, and they can integrate different functions required by users. For example, taking a mobile phone as an example, in one possible scenario, a user can open a game application to play a game, and at the same time, the user can also open a music application to listen to music. During this process, the phone may receive an incoming call and ring. Therefore, in the scenario where multiple applications play audio in parallel, the terminal device also faces the problem of how to play multiple audio streams.

[0047] At present, there are two common methods for playing multiple audio streams. One is to directly linearly superimpose different audio stream data and then play them out. The other is to selectively duck certain audio streams (i.e., lower the volume) to highlight the audio streams that have not been ducked.

[0048] However, directly adding different audio streams linearly and then playing them back usually results in serious problems such as confusion, inability to highlight the key points, and even unclear audio content. Using ducking processing can only ensure that the sound that is not suppressed can be heard relatively clearly, while the sound that is suppressed usually cannot be heard clearly.

[0049] That is to say, the common method of playing multiple audio streams has problems of confusing sound or incomplete content, and cannot take into account the clarity and integrity of the audio content, resulting in poor playback effect.

[0050] In order to solve the above problems, in an embodiment of the present application, the first scene category corresponding to the first audio data and the second scene category corresponding to the second audio data are determined; according to the first scene category and the second scene category, the first focus audio and the first non-focus audio are determined in the first audio data and the second audio data; based on the first focus audio and the first non-focus audio, the first virtual space position corresponding to the first audio data and the second virtual space position corresponding to the second audio data are determined; the first audio data is played based on the first virtual space position, and the second audio data is played based on the second virtual space position. It can be seen that in an embodiment of the present application, in the process of multi-audio parallel playback, the focus audio and non-focus audio can be determined in multiple audio data using the scene categories corresponding to different audio data, so that each audio data can be assigned a corresponding virtual space position based on the focus audio and the non-focus audio, so that different audio data can be played at different virtual space positions respectively. That is to say, the multi-audio playback method proposed in the embodiment of the present application can not only play each audio data completely, but also assign non-interfering virtual space positions to different audio data for rendering and playback, so as to take into account the clarity and integrity of the audio content and effectively improve the playback effect of multi-audio.

[0051] The technical solutions in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application.

[0052] An embodiment of the present application provides a multi-audio playback method, which can be applied to a multi-audio playback device or electronic device, or to an electronic device integrated with a multi-audio playback device. The multi-audio playback device or electronic device can simultaneously play different audio data created by different applications, which is not specifically limited in the present application.

[0053] It should be noted that, in the embodiment of the present application, the multi-audio playback method can be used to play at least two different audio data. Below, an exemplary description is given by taking the multi-audio playback device playing the first audio data and the second audio data simultaneously as an example.

[0054] Of course, the multi-audio playback method proposed in the embodiment of the present application can also be applied to any other device with a playback function, and can also complete the playback processing of any other number of different audio data. This application does not make specific limitations.

[0055] Furthermore, in the embodiments of the present application, Figure 1 Schematic diagram of the implementation process of the multi-audio playback method proposed in the embodiment of the present application Figure 1 ,like Figure 1As shown, the method for playing multiple audios by a multiple audio playback device may include the following steps:

[0056] Step 101: Determine a first scene category corresponding to first audio data and a second scene category corresponding to second audio data.

[0057] In an embodiment of the present application, the multi-audio playback device may first determine a first scene category corresponding to the first audio data and a second scene category corresponding to the second audio data, wherein the first audio data and the second audio data are two different audio data.

[0058] It should be noted that, in the embodiment of the present application, the first audio data and the second audio data may be data from different audio sources. The first audio data and the second audio data may be audio data created by the system or by an application.

[0059] Exemplarily, in some embodiments, the first audio data and the second audio data may be any type of audio data. For example, the first audio data and the second audio data may include any different audio data such as music, voice, ringtone, video, etc.

[0060] Exemplarily, in some embodiments, the first audio data may be music audio created by a music application, and the second audio data may be voice call audio created by a social application.

[0061] It should be noted that in an embodiment of the present application, after obtaining different audio data created by different applications, the multi-audio playback device can further determine the scene category corresponding to each audio data. For example, the first scene category corresponding to the first audio data is determined, and the second scene category corresponding to the second audio data is determined at the same time. The first scene category and the second scene category can be the same or different, and this application does not specifically limit this.

[0062] It can be understood that in the embodiments of the present application, when determining the scene category corresponding to the audio data, the scene category corresponding to each audio data can be divided based on any parameters such as the function, content, created application, etc. corresponding to the audio data, and the present application does not make specific limitations.

[0063] Exemplarily, in some embodiments, the scene category corresponding to the audio data may be determined based on the application or system that created the audio data. For example, for an audio data, the scene type of the application that created the audio data may be determined first, wherein the scene type of the application refers to the playback scene type or application type to which the application belongs, and then the scene type of the application is determined as the scene category corresponding to the audio data.

[0064] Exemplarily, in some embodiments, the scene category of the audio data determined based on the scene type of the application or system may include games, music, audiobooks, navigation, notifications, voice, incoming call ringtones, alarm ringtones or videos, etc.

[0065] Exemplarily, in some embodiments, the first audio data is created for a music application, and the corresponding first scene category may be music; the second audio data is created for a social application, and the corresponding second scene category may be an Internet call.

[0066] Step 102: Determine a first focus audio and a first non-focus audio in the first audio data and the second audio data according to the first scene category and the second scene category.

[0067] In an embodiment of the present application, after determining the first scene category corresponding to the first audio data and the second scene category corresponding to the second audio data, the multi-audio playback device can further determine the first focus audio and the first non-focus audio in the first audio data and the second audio data based on the first scene category and the second scene category.

[0068] It should be noted that, in the embodiments of the present application, the first focus audio can be understood as a holographic focus, which is similar to the concept of audio focus in Android. When multiple applications play sounds at the same time, there is only one holographic focus.

[0069] That is to say, in the embodiment of the present application, corresponding to the multi-audio playback scenario, there is one focus audio among the multiple audio data, that is, there is one holographic focus.

[0070] Accordingly, in the embodiment of the present application, the first non-focus audio can be understood as other audio outside the holographic focus. Among them, when multiple applications play sounds at the same time, there can be a holographic focus, and other audio outside the holographic focus can be understood as non-focus audio.

[0071] Further, in an embodiment of the present application, when determining the first focus audio and the first non-focus audio in the first audio data and the second audio data according to the first scene category and the second scene category, the first priority can be first determined according to the first scene category, and the second priority can be determined according to the second scene category; and then the first focus audio and the first non-focus audio can be determined in the first audio data and the second audio data according to the first priority and the second priority.

[0072] It can be understood that in the embodiments of the present application, for the first audio data and the second audio data, a focus audio, such as the first focus audio, and a non-focus audio, such as the first non-focus audio, can be determined respectively.

[0073] Accordingly, in an embodiment of the present application, for more than three audio data, since there is only one holographic focus, one focus audio and multiple non-focus audios can be determined.

[0074] It should be noted that in an embodiment of the present application, the priority corresponding to the scene category of the audio data can be determined first, and then the priorities corresponding to the audio data can be compared, so as to further determine the importance of the audio data based on the comparison results, and then determine the focus audio and non-focus audio therein.

[0075] Furthermore, in an embodiment of the present application, when determining the first priority according to the first scene category and determining the second priority according to the second scene category, the first priority corresponding to the first scene category and the second priority corresponding to the second priority can be determined based on the correspondence between the preset scenes and the priorities.

[0076] It is understandable that in the embodiment of the present application, corresponding priorities can be preset for different scene categories, that is, the correspondence between the preset scenes and the priorities can be determined. The correspondence between the preset scenes and the priorities can be continuously adjusted and updated.

[0077] For example, in some embodiments, Table 1 is a correspondence between preset scenes and priorities, wherein the smaller the value of the priority, the higher the priority. For different scene categories, the corresponding priorities may be the same or different. For example, the priorities corresponding to ringtones and alarms are the same, while the priorities corresponding to ringtones and videos are different, and the priority of videos is higher than that of ringtones.

[0078] Table 1

[0079] Scene Category Priority Internet calling 1 Game Live Broadcast 1 video 2 Audiobook 3 music 4 Navigation broadcast 5 ring 6 Alarm clock 6 Notification Tone 7

[0080] Exemplarily, in some embodiments, based on the correspondence between preset scenes and priorities shown in Table 1, assuming that the first scene category corresponding to the first audio data is music, then the corresponding first priority is 4; assuming that the second scene category corresponding to the second audio data is navigation broadcast, then the corresponding second priority is 5.

[0081] Further, in an embodiment of the present application, when determining the first focus audio and the first non-focus audio in the first audio data and the second audio data according to the first priority and the second priority, when the first priority is higher than the second priority, the first audio data is determined as the first focus audio, and the second audio data is determined as the first non-focus audio; when the first priority is lower than the second priority, the second audio data is determined as the first focus audio, and the first audio data is determined as the first non-focus audio; when the first priority is equal to the second priority, the first focus audio and the first non-focus audio are determined in the first audio data and the second audio data according to the first creation time corresponding to the first audio data and the second creation time corresponding to the second audio data.

[0082] It should be noted that in an embodiment of the present application, when using the priority corresponding to the audio data to determine the focus audio and non-focus audio, the priorities can be compared and the audio data with the highest priority can be determined as the focus audio, while the other audio data can be used as non-focus audio.

[0083] Exemplarily, in some embodiments, if the first priority is higher, the first audio data may be used as the focus audio, and if the second priority is higher, the second audio data may be used as the focus audio.

[0084] Accordingly, in an embodiment of the present application, when using the priority corresponding to the audio data to determine the focus audio and non-focus audio, if there are at least two audio data with the highest priority, then the creation time corresponding to the audio data with the highest priority can be further determined, and then the creation time can be used to determine the focus audio in the audio data with the highest priority.

[0085] Exemplarily, in some embodiments, if the first priority and the second priority are the same, then the first focus audio and the first non-focus audio need to be further determined in combination with the first creation time corresponding to the first audio data and the second creation time corresponding to the second audio data.

[0086] Further, in an embodiment of the present application, when determining the first focus audio and the first non-focus audio in the first audio data and the second audio data based on the first creation time corresponding to the first audio data and the second creation time corresponding to the second audio data, when the first creation time is greater than or equal to the second creation time, the first audio data is determined as the first focus audio, and the second audio data is determined as the first non-focus audio; when the first creation time is less than the second creation time, the second audio data is determined as the first focus audio, and the first audio data is determined as the first non-focus audio.

[0087] It can be understood that in an embodiment of the present application, when determining the focus audio among the audio data with the highest priority according to the creation time corresponding to the audio data, the audio data created later can be selected as the focus audio, that is, the audio data with the largest creation time can be selected as the focus audio.

[0088] It is understandable that in the embodiments of the present application, the larger the creation time, the later the corresponding audio data was created, and the audio data can be considered to be more important, so the audio data with the largest creation time can be used as the focus audio. For example, if the first creation time is greater than the second creation time, it can be considered that the first audio data is more important, so the first audio data can be determined as the focus audio.

[0089] It should be noted that, in the embodiment of the present application, for audio data with the same creation time and the highest priority, one of the audio data can be randomly selected as the focus audio.

[0090] Step 103: Determine a first virtual space position corresponding to the first audio data and a second virtual space position corresponding to the second audio data based on the first focused audio and the first non-focused audio.

[0091] In an embodiment of the present application, after determining the first focus audio and the first non-focus audio in the first audio data and the second audio data according to the first scene category and the second scene category, the multi-audio playback device can further determine the first virtual space position corresponding to the first audio data and the second virtual space position corresponding to the second audio data based on the first focus audio and the first non-focus audio.

[0092] It should be noted that in an embodiment of the present application, after the focus audio and the non-focus audio are determined in multiple audio data, the virtual space position corresponding to each audio data can be determined based on the focus audio and the non-focus audio. Among them, for the focus audio, the corresponding virtual space position is fixed, that is, no matter which audio data is used as the focus audio, it corresponds to the same position; and for the non-focus audio, the corresponding virtual space position is adaptively set based on different audio data, that is, different audio data may have different corresponding positions when used as non-focus audio.

[0093] It can be understood that, in the embodiments of the present application, the virtual space position can be understood as the sound image position in the virtual sound field space.

[0094] That is to say, in an embodiment of the present application, after determining the focus audio and the non-focus audio, the virtual space position corresponding to the audio data as the focus audio can be determined, and the virtual space position corresponding to the audio data as the non-focus audio can also be determined.

[0095] It should be noted that in the embodiments of the present application, when the spatial audio rendering algorithm renders multi-channel audio data, it is often assumed that multiple virtual speakers are arranged in the virtual listening space, and the data of each channel is fed to each virtual speaker. After being processed by the rendering algorithm, the sound is like it is emitted from the position of the virtual speaker. Among them, the basic sound unit processed by the spatial audio rendering algorithm generally includes audio data and sound position information, and the sound position information can be the virtual space position corresponding to the audio data.

[0096] Accordingly, in an embodiment of the present application, based on the division of focus audio and non-focus audio, the virtual space positions corresponding to different audio data may be divided into two categories: Center position and Home position.

[0097] It can be understood that in the embodiment of the present application, the Center position can be the position of the sound and image of the application holding the holographic focus, that is, it can be the virtual space position corresponding to the audio data serving as the focus audio.

[0098] It is understandable that in the embodiment of the present application, the Home position can be the position of the sound and image of the application that does not have a holographic focus, that is, it can be the virtual space position corresponding to the audio data as non-focus audio. Among them, audio data of different scene categories correspond to different Home positions.

[0099] Furthermore, in an embodiment of the present application, the correspondence between scenes and non-focus positions can be determined. Specifically, for each scene category, a non-focus position corresponding to the scene category can be set, and when the audio data corresponding to the scene category is used as non-focus audio, the non-focus position can be used to determine the virtual space position of the audio data.

[0100] It should be noted that in the embodiments of the present application, the correspondence between the scene and the non-focus position can be set based on the preset distribution principle. Among them, the core of the preset distribution principle is that the Home position (non-focus position) of different audio scenes is different, that is, different scene categories correspond to different non-focus positions. Among them, the overall Home position layout must ensure that the sound position of the entire sound field distribution can be distributed as balanced as possible in the front, back, left and right.

[0101] For example, in some embodiments, the distribution of the Home position can be multiple. Based on the preset distribution principle, for audio scenes (scene categories) with rich audio content, such as music and video, the Home position can be set in a central symmetrical position, which will have a better stereo effect. For audio scenes with relatively simple sound content, such as Internet calls, listening to books, prompt sounds, etc., the Home position can be set in a position biased in a certain direction.

[0102] Exemplarily, in some embodiments, Figure 2 This is a schematic diagram of the correspondence between the scene and the non-focus position proposed in the embodiment of the present application, such as Figure 2 As shown, you can set the non-focal position (Home position) corresponding to different scene categories in the virtual sound field space, where the position of the headphones is the center of the virtual sound field space. Based on the preset distribution principle, for scene categories such as music and video, you can choose to set the Home position in a central symmetrical position, which will have a better stereo effect. For scene categories such as online calls, listening to books, notification sounds, etc., you can choose to set the Home position in a position biased in a certain direction.

[0103] It should be noted that, in the embodiments of the present application, the correspondence between the scene and the non-focus position may be determined in a variety of ways, which are not specifically limited in the present application.

[0104] Exemplarily, in some embodiments, when determining the correspondence between a scene and a non-focus position, the preset correspondence between the scene and the non-focus position may be directly acquired.

[0105] Exemplarily, in some embodiments, when determining the correspondence between scenes and non-focus positions, the setting instructions corresponding to the first application can be obtained first; then the non-focus position corresponding to the audio category created by the first application is determined based on the setting instructions, thereby determining the correspondence between scenes and non-focus positions.

[0106] Exemplarily, in some embodiments, when determining the correspondence between scenes and non-focus positions, custom information of the first application can be obtained first; the non-focus position corresponding to the audio category created by the first application is determined based on the custom information, so that the correspondence between the scene and the non-focus position can be determined.

[0107] Further, in an embodiment of the present application, when determining a first virtual space position corresponding to the first audio data and a second virtual space position corresponding to the second audio data based on the first focus audio and the first non-focus audio, when the first audio data is the first focus audio and the second audio data is the first non-focus audio, the non-focus position corresponding to the second audio data can be determined based on the correspondence between the scene and the non-focus position and the second scene category; then the preset center position is determined as the first virtual space position, and the non-focus position corresponding to the second audio data is determined as the second virtual space position.

[0108] Further, in an embodiment of the present application, when determining a first virtual space position corresponding to the first audio data and a second virtual space position corresponding to the second audio data based on the first focus audio and the first non-focus audio, when the second audio data is the first focus audio and the first audio data is the first non-focus audio, the non-focus position corresponding to the first audio data can be determined based on the correspondence between the scene and the non-focus position and the first scene category; then the preset center position is determined as the second virtual space position, and the non-focus position corresponding to the first audio data is determined as the first virtual space position.

[0109] It can be understood that in the embodiments of the present application, the preset center position can be the spatial position corresponding to the focus audio, wherein the preset center position can be fixed or adaptively changed, and the present application does not make specific limitations.

[0110] It is understandable that in the embodiments of the present application, the preset center position can be determined based on the playback configuration and playback parameters of the multi-audio playback device or equipment. For example, for a dual-channel playback device, the corresponding preset center position can be a spatial position concentrated at one point, or can include two symmetrical spatial positions; for a quad-channel playback device, the corresponding preset center position can include four mutually symmetrical spatial positions, which is not specifically limited in the present application.

[0111] Exemplarily, in some embodiments, taking dual channels as an example, assuming that the preset center position is fixed, it can be two symmetrical spatial positions, that is, for any audio data as the focus audio, the corresponding virtual spatial position is the preset center position.

[0112] Exemplarily, in some embodiments, Figure 3 Schematic diagram of the preset center position proposed in the embodiment of the present application Figure 1 ,like Figure 3 As shown, taking dual channels as an example, assuming that the preset center position is fixed, symmetrical spatial positions A and B, then when the first audio data is the first focus audio, the first virtual spatial position is the preset center position; when the second audio data is the first focus audio, the second virtual spatial position is also the preset center position.

[0113] That is to say, if the preset center position is fixed, then for the first audio data or the second audio data as the first focus audio, the corresponding virtual space position is the preset center position. That is, if the first audio data is the first focus audio, the first virtual space position is the preset center position, and if the second audio data is the second focus audio, the second virtual space position is the preset center position.

[0114] Exemplarily, in some embodiments, taking dual channels as an example, assuming that the preset center position is adaptively changeable, it can be adaptively determined according to the scene category corresponding to the audio data, that is, for different audio data serving as focus audio, the corresponding preset center position may be different or the same, and this application does not make specific limitations.

[0115] Exemplarily, in some embodiments, Figure 4 Schematic diagram of the preset center position proposed in the embodiment of the present application Figure 2 ,like Figure 4 As shown, taking dual channels as an example, assuming that the preset center position is adaptively changed, when the first audio data is the first focus audio, the first virtual spatial position, that is, the preset center position corresponding to the first audio data is symmetrical spatial positions A and B; when the second audio data is the first focus audio, the second virtual spatial position, that is, the preset center position corresponding to the second audio data is a spatial position C concentrated on one point.

[0116] That is to say, if the preset center position is adaptively changed, then for the first audio data or the second audio data as the first focus audio, the corresponding preset center position may be different. That is, if the first audio data is the first focus audio, the first virtual space position may be the preset center position corresponding to the first audio data, and if the second audio data is the second focus audio, the second virtual space position may be the preset center position corresponding to the second audio data.

[0117] It can be understood that in an embodiment of the present application, the spatial position corresponding to the non-focus audio can be adaptively changed. For example, based on the correspondence between the scene and the non-focus position, the virtual spatial position corresponding to the first non-focus audio can be determined.

[0118] For example, in some embodiments, if the second audio data is the first non-focus audio, then the non-focus position corresponding to the second audio data can be determined according to the correspondence between the scene and the non-focus position and the second scene category; and the non-focus position is determined as the second virtual space position. Figure 5 Schematic diagram of the virtual space position proposed in the embodiment of the present application Figure 1 ,like Figure 5 As shown, taking dual channels as an example, when the first audio data is the first focus audio and the second audio data is the first non-focus audio, the first virtual spatial position is a preset center position including two symmetrical spatial positions, and the second virtual spatial position is the non-focus position corresponding to the second audio data.

[0119] For example, in some embodiments, if the first audio data is the first non-focus audio, then the non-focus position corresponding to the first audio data can be determined according to the correspondence between the scene and the non-focus position and the first scene category; and the non-focus position is determined as the second virtual space position. Figure 6 Schematic diagram of the virtual space position proposed in the embodiment of the present application Figure 2 ,like Figure 6 As shown, taking dual channels as an example, when the second audio data is the first focus audio and the first audio data is the first non-focus audio, the second virtual spatial position is a preset center position including two symmetrical spatial positions, and the first virtual spatial position is the non-focus position corresponding to the first audio data.

[0120] Step 104: Play the first audio data based on the first virtual space position, and play the second audio data based on the second virtual space position.

[0121] In an embodiment of the present application, after determining the first virtual space position corresponding to the first audio data and the second virtual space position corresponding to the second audio data based on the first focus audio and the first non-focus audio, the multi-audio playback device can further play the first audio data based on the first virtual space position, and simultaneously play the second audio data based on the second virtual space position.

[0122] Furthermore, in an embodiment of the present application, when the first audio data is played based on the first virtual space position and the second audio data is played based on the second virtual space position, when the first audio data is the first focus audio and the second audio data is the first non-focus audio, the first audio data can be played in a fade-in manner based on the first virtual space position; and the second audio data can be played directly based on the second virtual space position.

[0123] It should be noted that, in an embodiment of the present application, if the first audio data is the first focus audio, then when playing the first audio data based on the first virtual space position, in addition to playing the first audio data in a fade-in manner based on the first virtual space position, the first audio data can also be played directly at the first virtual space position.

[0124] Furthermore, in an embodiment of the present application, when the first audio data is played based on the first virtual space position and the second audio data is played based on the second virtual space position, when the second audio data is the first focus audio and the first audio data is the first non-focus audio, the second audio data can be played in a fade-in manner based on the second virtual space position; and the first audio data can be played directly based on the first virtual space position.

[0125] It should be noted that, in an embodiment of the present application, if the second audio data is the first focus audio, then when playing the second audio data based on the second virtual space position, in addition to playing the second audio data in a fade-in manner based on the second virtual space position, the second audio data can also be played directly at the second virtual space position.

[0126] That is to say, in the embodiments of the present application, for the focus audio, the playback mode adopted can be a gradual-in playback mode or a direct playback mode, and the present application does not make specific limitations.

[0127] Exemplarily, in some embodiments, when the first audio data is the first focus audio, the first audio data is played in a fade-in manner based on the first virtual space position; or, the first audio data is played directly at the first virtual space position; when the first audio data is the first non-focus audio, the first audio data is played directly at the first virtual space position.

[0128] Exemplarily, in some embodiments, when the second audio data is the first focus audio, the second audio data is played in a fade-in manner based on the second virtual space position; or, the second audio data is played directly at the second virtual space position; when the second audio data is the first non-focus audio, the second audio data is played directly at the second virtual space position.

[0129] Exemplarily, in some embodiments, for the focus audio among the multiple audio data to be played, you can choose to use a gradual-in method to play the focus audio; for the non-focus audio among the multiple audio data to be played, you can choose to directly play the non-focus audio.

[0130] It should be noted that in an embodiment of the present application, for the focus audio among the multiple audio data to be played, when the focus audio is played in a fade-in manner, the focus audio can be played in a fade-in manner according to a preset fade-in path, from a preset starting spatial position to a preset center position.

[0131] It can be understood that, in the embodiments of the present application, the preset gradual entry path may include a straight path or a curved path, and the present application does not make any specific limitation thereto.

[0132] Exemplarily, in some embodiments, Figure 7 Schematic diagram of the implementation of the gradual entry method proposed in the embodiment of the present application Figure 1 ,like Figure 7 As shown, a preset starting spatial position is predetermined in the virtual space, the first audio data is the first focus audio, and the first virtual spatial position is a preset center position including two symmetrical spatial positions. Then, the first audio data can be played in a fade-in manner by pointing from the preset starting spatial position to the preset center position.

[0133] Exemplarily, in some embodiments, Figure 8 Schematic diagram of the implementation of the gradual entry method proposed in the embodiment of the present application Figure 2 ,like Figure 8 As shown, a preset starting spatial position is predetermined in the virtual space, the second audio data is the first focus audio, and the second virtual spatial position is a preset center position concentrated in one spatial position. Then, the second audio data can be played in a gradual manner by pointing from the preset starting spatial position to the preset center position.

[0134] Furthermore, in the embodiments of the present application, Fig. 9 Schematic diagram of the implementation process of the multi-audio playback method proposed in the embodiment of the present application Figure 2 ,like Fig. 9 As shown, after playing the first audio data based on the first virtual space position and playing the second audio data based on the second virtual space position, that is, after step 104, the method for playing multiple audios by the multiple audio playback device further includes the following steps:

[0135] Step 105: Receive a play instruction corresponding to the third audio data.

[0136] In an embodiment of the present application, during the playback of the first audio data and the second audio data, a playback instruction of the third audio data may be continuously received, wherein the playback instruction is used to instruct the third audio data to perform multi-audio parallel playback processing.

[0137] It is understandable that, in the embodiment of the present application, the third audio data may be audio data created by the system or audio data created by the application. The third audio data may be audio data different from the first audio data and the second audio data.

[0138] Exemplarily, in some embodiments, the third audio data may be any type of audio data. For example, the third audio data may include any different audio data such as music, voice, ringtone, video, etc.

[0139] It should be noted that in an embodiment of the present application, when multiple audio data are being played in parallel, if new audio data, such as third audio data, is added, then the focus data and non-focus data in the multiple audio data may change accordingly. Therefore, it is necessary to further update and determine the focus data and non-focus data in the multiple audio data.

[0140] Step 106: Determine a third scene category corresponding to the third audio data.

[0141] In an embodiment of the present application, after receiving a play instruction corresponding to the third audio data, the multi-audio play device may further determine a third scene category corresponding to the third audio data.

[0142] Exemplarily, in some embodiments, the scene category corresponding to the audio data may be determined according to the application or system that creates the audio data. For example, for the third audio data, the scene type of the application that creates the third audio data may be first determined, and then the scene type of the application may be determined as the third scene category corresponding to the third audio data.

[0143] It should be noted that, in the embodiments of the present application, the third scene category may be the same as or different from the first scene category or the second scene category, and the present application does not make any specific limitation thereto.

[0144] Exemplarily, in some embodiments, the first audio data is created for a music application, and the corresponding first scene category may be music; the second audio data is created for a social application, and the corresponding second scene category may be an Internet call; the third audio data is created for a clock in the system, and the corresponding third scene category may be an alarm.

[0145] Step 107: Determine a second focus audio from the first audio data, the second audio data, and the third audio data according to the third scene category, and determine a second non-focus audio.

[0146] In an embodiment of the present application, after determining the third scene category corresponding to the third audio data, the multi-audio playback device can further determine the second focus audio from the first audio data, the second audio data, and the third audio data according to the third scene category, and simultaneously determine the second non-focus audio. The second focus audio can be one of the first audio data, the second audio data, and the third audio data, and the second non-focus audio can be the other two audio data other than the second focus audio.

[0147] Further, in an embodiment of the present application, when determining a second focus audio in the first audio data, the second audio data, and the third audio data according to the third scene category, and determining the second non-focus audio at the same time, the third priority corresponding to the third scene category can be determined based on the correspondence between the preset scene and the priority; when the first audio data is the first focus audio and the first priority is higher than the third priority, the first audio data is determined as the second focus audio, and the second audio data and the third audio data are determined as the second non-focus audio; when the first audio data is the first focus audio and the first priority is lower than or equal to the third priority, the third audio data is determined as the second focus audio, and the first audio data and the second audio data are determined as the second non-focus audio.

[0148] Further, in an embodiment of the present application, when determining a second focus audio in the first audio data, the second audio data, and the third audio data according to the third scene category, and determining the second non-focus audio at the same time, the third priority corresponding to the third scene category can be determined based on the correspondence between the preset scene and the priority; when the second audio data is the first focus audio and the second priority is higher than the third priority, the second audio data is determined as the second focus audio, and the first audio data and the third audio data are determined as the second non-focus audio; when the second audio data is the first focus audio and the second priority is lower than or equal to the third priority, the third audio data is determined as the second focus audio, and the first audio data and the second audio data are determined as the second non-focus audio.

[0149] It should be noted that in an embodiment of the present application, when new audio data needs to be played in parallel, the priority corresponding to the new audio data can be determined first, and then the new focus audio and non-focus audio can be further determined in combination with the priority corresponding to the audio data being played, thereby completing the update of the focus audio.

[0150] It should be noted that in the embodiments of the present application, the second focus audio can be understood as the holographic focus. When multiple applications play sounds at the same time, there is only one holographic focus. The second non-focus audio can be understood as other audio other than the holographic focus. Among them, when multiple applications play sounds at the same time, there can be one holographic focus, and other audio other than the holographic focus can be understood as non-focus audio.

[0151] It can be understood that in the embodiments of the present application, for the first audio data, the second audio data and the third audio data, one focus audio, such as the second focus audio, and two non-focus audios, such as the second non-focus audio, can be determined respectively.

[0152] It should be noted that in an embodiment of the present application, based on the correspondence between preset scenes and priorities, the priority corresponding to the scene category of the audio data can be determined first, and then the priorities corresponding to the audio data can be compared, so as to further determine the importance of the audio data based on the comparison results, and then determine the focus audio and non-focus audio therein.

[0153] Exemplarily, in some embodiments, based on the correspondence between preset scenes and priorities shown in Table 1, assuming that the first scene category corresponding to the first audio data is music, then the corresponding first priority is 4, assuming that the second scene category corresponding to the second audio data is navigation broadcast, then the corresponding second priority is 5, assuming that the third scene category corresponding to the third audio data is video, then the corresponding third priority is 2. At this time, the priorities of the audio data from high to low are the third audio data, the first audio data, and the second audio data.

[0154] It is understandable that in the embodiment of the present application, for the first audio data and the second audio data being played in parallel, if the first audio data is the original focus audio, i.e., the first focus audio, and at the same time, it is determined that the first priority is higher than the third priority, then it can be considered that the first audio data is still the audio data with the highest priority. At this time, it can be determined that the updated focus audio is still the first audio data, i.e., the first audio data is determined as the second focus audio. At the same time, the second audio data and the third audio data can be used as the second non-focus audio.

[0155] Accordingly, in an embodiment of the present application, if the first audio data is the original focus audio, and it is determined that the first priority is lower than or equal to the third priority, then the third audio data can be considered as the audio data with the highest priority. At this time, the updated focus audio can be determined to be the third audio data, that is, the third audio data is determined as the second focus audio. Accordingly, the second audio data and the first audio data can be used as the second non-focus audio.

[0156] It should be noted that, in an embodiment of the present application, when the third priority corresponding to the newly added third audio data is the same as the first priority corresponding to the first audio data which is the original focus audio, since the third creation time corresponding to the third audio data is greater than the first creation time corresponding to the first audio data by default, it is possible to choose not to compare the creation times of the two, but to directly determine the newly added third audio data as the new focus audio, i.e., the second focus audio.

[0157] It is understandable that in the embodiment of the present application, for the first audio data and the second audio data being played in parallel, if the second audio data is the original focus audio, that is, the first focus audio, and at the same time, it is determined that the second priority is higher than the third priority, then it can be considered that the second audio data is still the audio data with the highest priority. At this time, it can be determined that the updated focus audio is still the second audio data, that is, the second audio data is determined as the second focus audio. At the same time, the first audio data and the third audio data can be used as the second non-focus audio.

[0158] Accordingly, in an embodiment of the present application, if the second audio data is the original focus audio, and it is determined that the second priority is lower than or equal to the third priority, then the third audio data can be considered as the audio data with the highest priority. At this time, the updated focus audio can be determined to be the third audio data, that is, the third audio data is determined as the second focus audio. Accordingly, the second audio data and the first audio data can be used as the second non-focus audio.

[0159] It should be noted that, in an embodiment of the present application, when the third priority corresponding to the newly added third audio data is the same as the second priority corresponding to the second audio data which is the original focus audio, since the third creation time corresponding to the third audio data is greater than the second creation time corresponding to the second audio data by default, it is possible to choose not to compare the creation times of the two, but to directly determine the newly added third audio data as the new focus audio, i.e., the second focus audio.

[0160] Step 108: Based on the second focused audio and the second non-focused audio, determine the third virtual space position corresponding to the third audio data, the fourth virtual space position corresponding to the first audio data, and the fifth virtual space position corresponding to the second audio data.

[0161] In an embodiment of the present application, after determining a second focus audio in the first audio data, the second audio data and the third audio data according to the third scene category and determining the second non-focus audio, the third virtual space position corresponding to the third audio data, the fourth virtual space position corresponding to the first audio data, and the fifth virtual space position corresponding to the second audio data can be further determined based on the second focus audio and the second non-focus audio.

[0162] It should be noted that in an embodiment of the present application, after the focus audio and the non-focus audio are determined in multiple audio data, the virtual space position corresponding to each audio data can be determined based on the focus audio and the non-focus audio. Among them, for the focus audio, the corresponding virtual space position is fixed, and for the non-focus audio, different audio data may have different corresponding positions when used as non-focus audio.

[0163] Further, in an embodiment of the present application, when determining the third virtual space position corresponding to the third audio data, the fourth virtual space position corresponding to the first audio data, and the fifth virtual space position corresponding to the second audio data based on the second focus audio and the second non-focus audio, when the third audio data is the second focus audio, the preset center position can be directly determined as the third virtual space position; when the third audio data is the second non-focus audio, the non-focus position corresponding to the third audio data can be first determined based on the correspondence between the scene and the non-focus position and the third scene category, and then the non-focus position corresponding to the third audio data can be determined as the third virtual space position.

[0164] That is to say, in an embodiment of the present application, for the third audio data, if the third audio data becomes the new focus audio, that is, the third audio data is the second focus audio, then the corresponding virtual space position can be the preset center position, that is, the preset center position can be set to the third virtual space position corresponding to the third audio data.

[0165] Correspondingly, in an embodiment of the present application, for the third audio data, if the third audio data becomes a new non-focus audio, that is, the third audio data is the second non-focus audio, then the non-focus position corresponding to the third audio data, that is, the third virtual space position, can be determined based on the correspondence between the scene and the non-focus position and the third scene category.

[0166] Further, in an embodiment of the present application, when the third virtual space position corresponding to the third audio data, the fourth virtual space position corresponding to the first audio data, and the fifth virtual space position corresponding to the second audio data are determined based on the second focus audio and the second non-focus audio, when the first audio data is the second focus audio, the first virtual space position is determined as the fourth virtual space position; when the first audio data is the second non-focus audio, the non-focus position corresponding to the first audio data is determined based on the correspondence between the scene and the non-focus position and the first scene category, and the non-focus position corresponding to the first audio data is determined as the fourth virtual space position.

[0167] It should be noted that, in the embodiment of the present application, the fourth virtual space position may be an updated virtual space position corresponding to the first audio data, wherein the first virtual space position and the fourth virtual space position may be the same or different.

[0168] It is understandable that in the embodiment of the present application, if the first audio data remains as the focus audio or the non-focus audio after the third audio data is added, then it is considered that the first virtual space position corresponding to the first audio data can remain unchanged. If the first audio data changes from the first focus audio to the second non-focus audio, then it is considered that the first virtual space position corresponding to the first audio data will change.

[0169] Exemplarily, in some embodiments, as the first audio data of the first focus audio, if after the third audio data is added, the first audio data is updated to the second focus audio, then it can be considered that the first audio data can still occupy the preset center position, and therefore, the corresponding spatial position can be kept unchanged, that is, the first virtual spatial position is determined to be the fourth virtual spatial position.

[0170] Correspondingly, in an embodiment of the present application, as the first audio data of the first focus audio, if after the third audio data is added, the first audio data is updated to the second non-focus audio, then it can be considered that the first audio data no longer occupies the preset center position, and the corresponding spatial position needs to be updated. At this time, the non-focus position corresponding to the first audio data can be determined based on the correspondence between the scene and the non-focus position and the first scene category, and the non-focus position can be determined as the fourth virtual spatial position.

[0171] Further, in an embodiment of the present application, when the third virtual space position corresponding to the third audio data, the fourth virtual space position corresponding to the first audio data, and the fifth virtual space position corresponding to the second audio data are determined based on the second focus audio and the second non-focus audio, when the second audio data is the second focus audio, the second virtual space position is determined as the fifth virtual space position; when the second audio data is the second non-focus audio, the non-focus position corresponding to the second audio data is determined based on the correspondence between the scene and the non-focus position and the second scene category, and the non-focus position corresponding to the second audio data is determined as the fifth virtual space position.

[0172] It should be noted that, in the embodiment of the present application, the fifth virtual space position may be an updated virtual space position corresponding to the second audio data, wherein the second virtual space position and the fifth virtual space position may be the same or different.

[0173] It is understandable that in the embodiment of the present application, if the second audio data remains as the focus audio or the non-focus audio after the third audio data is added, then it is considered that the second virtual space position corresponding to the second audio data can remain unchanged. If the second audio data changes from the first focus audio to the second non-focus audio, then it is considered that the second virtual space position corresponding to the second audio data will change.

[0174] Exemplarily, in some embodiments, as the second audio data of the first focus audio, if the second audio data is updated to the second focus audio after the third audio data is added, then it can be considered that the second audio data can still occupy the preset center position, and therefore, the corresponding spatial position can be kept unchanged, that is, the second virtual spatial position is determined to be the fifth virtual spatial position.

[0175] Correspondingly, in an embodiment of the present application, as the second audio data of the first focus audio, if after the third audio data is added, the second audio data is updated to the second non-focus audio, then it can be considered that the second audio data no longer occupies the preset center position, and the corresponding spatial position needs to be updated. At this time, the non-focus position corresponding to the second audio data can be determined based on the correspondence between the scene and the non-focus position and the second scene category, and the non-focus position can be determined as the fifth virtual spatial position.

[0176] Step 109: Play the first audio data based on the fourth virtual space position, play the second audio data based on the fifth virtual space position; and play the third audio data based on the third virtual space position.

[0177] In an embodiment of the present application, after determining the third virtual space position corresponding to the third audio data, the fourth virtual space position corresponding to the first audio data, and the fifth virtual space position corresponding to the second audio data based on the second focus audio and the second non-focus audio, the multi-audio playback device can further play the first audio data based on the fourth virtual space position, play the second audio data based on the fifth virtual space position; and play the third audio data based on the third virtual space position.

[0178] Furthermore, in an embodiment of the present application, when the third audio data is played based on the third virtual space position, when the third audio data is the second focus audio, the third audio data is played in a fade-in manner based on the third virtual space position; when the third audio data is the second non-focus audio, the third audio data is played directly based on the third virtual space position.

[0179] It should be noted that, in an embodiment of the present application, if the third audio data is the second focus audio, then when playing the third audio data based on the third virtual space position, in addition to playing the third audio data in a fade-in manner based on the third virtual space position, the third audio data can also be played directly at the third virtual space position.

[0180] That is to say, in the embodiments of the present application, for the focus audio, the playback mode adopted can be a gradual-in playback mode or a direct playback mode, and the present application does not make specific limitations.

[0181] Exemplarily, in some embodiments, when the third audio data is the second focus audio, the third audio data is played in a fade-in manner based on the third virtual space position; or, the third audio data is played directly at the third virtual space position; when the third audio data is the second non-focus audio, the third audio data is played directly at the third virtual space position.

[0182] It can be understood that in an embodiment of the present application, for the third audio data to be played, if it is determined to be new focus audio, then the third audio data can be played in a gradual manner; if it is determined to be fat focus audio data, then the third audio data can be played directly.

[0183] It should be noted that, in the embodiment of the present application, when the third audio data is played in a fade-in manner, the third audio data can be played in a fade-in manner according to a preset fade-in path, from a preset starting spatial position to a preset center position. The preset fade-in path may include a straight path or a curved path, which is not specifically limited in the present application.

[0184] Exemplarily, in some embodiments, Fig.10 Schematic diagram of the implementation of the gradual entry method proposed in the embodiment of the present application Figure 3 ,like Fig.10 As shown, a preset starting spatial position is predetermined in the virtual space, the third audio data is the second focus audio, and the third virtual spatial position includes a preset center position of two symmetrical spatial positions. Then, the third audio data can be played in a fade-in manner by pointing from the preset starting spatial position to the preset center position.

[0185] Further, in an embodiment of the present application, if the first audio data is the first focus audio, then when the first audio data is played based on the fourth virtual space position, when the first audio data is the second focus audio, the first audio data is directly played based on the fourth virtual space position; when the first audio data is the second non-focus audio, the first audio data is played in a fade-out manner based on the fourth virtual space position.

[0186] It should be noted that in an embodiment of the present application, if the first audio data remains as the focus audio and does not change after the third audio data is added, then it is considered that the virtual space position corresponding to the first audio data can be maintained as the preset center position unchanged. Accordingly, when playing the first audio data, the first audio data can be played directly based on the fourth virtual space position, that is, the first audio data is played directly at the fourth virtual space position.

[0187] It should be noted that in the embodiment of the present application, if the first audio data changes from the first focus audio to the second non-focus audio after the third audio data is added, then it is considered that the first virtual space position corresponding to the first audio data will change, that is, the first virtual space position is changed to the fourth virtual space position. Therefore, when playing the first audio data, you can choose to play the first audio data in a fade-out manner.

[0188] It can be understood that in the embodiment of the present application, when the first audio data is played in a fade-out manner, the first audio data can be played in a fade-out manner from the first virtual space position to the fourth virtual space position according to a preset fade-out path.

[0189] It can be understood that, in the embodiments of the present application, the preset fade-out path includes a straight path or a curved path, which is not specifically limited in the present application.

[0190] That is to say, in an embodiment of the present application, for the audio data being played, after adding new audio data, if the original audio data is updated from focus audio to non-focus audio, and the corresponding spatial position is updated from the Center position to the Home position, then a gradual-out method can be selected to play the audio data.

[0191] Exemplarily, in some embodiments, Fig.11 Schematic diagram of the implementation of the gradual exit method proposed in the embodiment of the present application Figure 1 ,like Fig.11 As shown, assuming that the first audio data is the first focus audio, the first virtual space position is a preset center position corresponding to the first audio data and concentrated in a spatial position, after adding the third audio data and updating the focus audio, the first audio data is updated to the second non-focus audio, and the corresponding fourth virtual space position is the non-focus position corresponding to the first audio data determined based on the correspondence between the scene and the non-focus position. Then, the first virtual space position can be pointed to the fourth virtual space position, and the first audio data can be played in a gradual out manner.

[0192] Exemplarily, in some embodiments, Fig.12 Schematic diagram of the implementation of switching focus audio proposed in the embodiment of the present application Figure 1 ,like Fig.12As shown, assuming that the first audio data is the first focus audio, the first virtual space position is the preset center position corresponding to the first audio data and concentrated in one space position. After adding the third audio data and updating the focus audio, the third audio serves as the second focus audio, and the third virtual space position includes the preset center positions of two symmetrical space positions corresponding to the third audio data, while the first audio data is updated to the second non-focus audio, and the corresponding fourth virtual space position is the non-focus position corresponding to the first audio data determined based on the correspondence between the scene and the non-focus position. Then, in the process of playing the audio data, the third audio data can be played in a fade-in manner by pointing to the third virtual space position from the preset starting space position, and the first audio data can be played in a fade-out manner by pointing to the fourth virtual space position from the first virtual space position.

[0193] Further, in an embodiment of the present application, if the second audio data is the first focus audio, then when the second audio data is played based on the fifth virtual space position, when the second audio data is the second focus audio, the second audio data is played directly based on the fifth virtual space position; when the second audio data is the second non-focus audio, based on the fifth virtual space position, the second audio data is played in a fade-out manner, that is, the second audio data is played directly at the fifth virtual space position.

[0194] It should be noted that in an embodiment of the present application, if the second audio data remains as the focus audio without change after the third audio data is added, then it is considered that the virtual space position corresponding to the second audio data can be maintained at the preset center position unchanged. Accordingly, when playing the second audio data, the second audio data can be played directly based on the fifth virtual space position.

[0195] It should be noted that in the embodiment of the present application, if the second audio data changes from the first focus audio to the second non-focus audio after the third audio data is added, then it is considered that the second virtual space position corresponding to the second audio data will change, that is, the second virtual space position is changed to the fifth virtual space position. Therefore, when playing the second audio data, you can choose to play the second audio data in a fade-out manner.

[0196] It can be understood that in the embodiment of the present application, when the second audio data is played in a fade-out manner, the second audio data can be played in a fade-out manner from the second virtual space position to the fifth virtual space position according to a preset fade-out path.

[0197] It can be understood that, in the embodiments of the present application, the preset fade-out path includes a straight path or a curved path, which is not specifically limited in the present application.

[0198] That is to say, in an embodiment of the present application, for the audio data being played, after adding new audio data, if the original audio data is updated from focus audio to non-focus audio, and the corresponding spatial position is updated from the Center position to the Home position, then a gradual-out method can be selected to play the audio data.

[0199] Exemplarily, in some embodiments, Fig.13 Schematic diagram of the implementation of the gradual exit method proposed in the embodiment of the present application Figure 2 ,like Fig.13 As shown, assuming that the second audio data is the first focus audio, the second virtual space position is fixedly set to a preset center position including two symmetrical space positions, after adding the third audio data and updating the focus audio, the second audio data is updated to the second non-focus audio, and the corresponding fifth virtual space position is the non-focus position corresponding to the second audio data determined based on the correspondence between the scene and the non-focus position. Then, the second virtual space position can be pointed to the fifth virtual space position, and the second audio data can be played in a gradual out manner.

[0200] Exemplarily, in some embodiments, Fig.14 Schematic diagram of the implementation of switching focus audio proposed in the embodiment of the present application Figure 2 ,like Fig.14 As shown, it is assumed that the second audio data is the first focus audio, and the second virtual space position is a fixed preset center position including two symmetrical space positions. After adding the third audio data and updating the focus audio, the third audio serves as the second focus audio, and the third virtual space position is also a fixed preset center position including two symmetrical space positions, while the second audio data is updated to the second non-focus audio, and the corresponding fifth virtual space position is the non-focus position corresponding to the second audio data determined based on the correspondence between the scene and the non-focus position. Then, in the process of playing the audio data, the third audio data can be played in a fade-in manner by pointing to the third virtual space position from the preset starting space position, and the second audio data can be played in a fade-out manner by pointing to the fifth virtual space position from the second virtual space position.

[0201] Furthermore, in the embodiments of the present application, Fig.15 Schematic diagram of the implementation process of the multi-audio playback method proposed in the embodiment of the present application Figure 3 ,like Fig.15 As shown, after playing the first audio data based on the first virtual space position and playing the second audio data based on the second virtual space position, that is, after step 104, the method for playing multiple audios by the multiple audio playback device further includes the following steps:

[0202] Step 110: After receiving a stop instruction corresponding to the first audio data, stop playing the first audio data.

[0203] In an embodiment of the present application, during the process of playing the first audio data and the second audio data in parallel, a stop instruction of the first audio data may also be received, and then, in response to the stop instruction, the playing of the first audio data may be stopped.

[0204] It is understandable that, in the embodiment of the present application, when multiple audio data are played in parallel, it is also possible to choose to stop playing any one or more of the audio data based on the received instruction.

[0205] Step 111: When the first audio data is the first focus audio, the second audio data is updated to the first focus audio, and the second virtual space position is updated according to the preset center position.

[0206] In an embodiment of the present application, when a stop instruction corresponding to the first audio data is received and the playing of the first audio data is stopped, when the first audio data is the first focus audio, the second audio data can be selected to be updated to the first focus audio, and the second virtual space position can be updated according to the preset center position.

[0207] It can be understood that, in an embodiment of the present application, if the audio data that is paused is the focus audio, then the remaining other audio data can be determined as the new focus audio, thereby improving the audio playback effect.

[0208] Exemplarily, in some embodiments, if the first audio data serving as the first focus audio stops playing, at this time, continuing to play the second audio data at a non-focus position will reduce the audio playback effect to a certain extent. Therefore, the second audio data can be used as the new focus audio again, that is, the second audio data can be updated to the first focus audio. At the same time, the second virtual space position corresponding to the second audio data can also be determined as the preset center position.

[0209] Step 112: Play the second audio data based on the updated second virtual space position.

[0210] In an embodiment of the present application, when the first audio data is the first focus audio, the second audio data is updated to the first focus audio, and after the second virtual space position is updated according to the preset center position, the second audio data can be further played based on the updated second virtual space position.

[0211] Furthermore, in the embodiments of the present application, Fig.16 Schematic diagram of the implementation process of the multi-audio playback method proposed in the embodiment of the present application Figure 4 ,like Fig.16As shown, after receiving the stop instruction corresponding to the first audio data, after stopping playing the first audio data, that is, after step 110, the method for playing multiple audios by the multiple audio playing device further includes the following steps:

[0212] Step 113: When the first audio data is the first non-focus audio, continue to play the second audio data based on the second virtual space position.

[0213] In an embodiment of the present application, when a stop instruction corresponding to the first audio data is received and the playing of the first audio data is stopped, if the first audio data is the first non-focus audio, you can choose to continue playing the second audio data based on the second virtual space position.

[0214] It can be understood that, in an embodiment of the present application, if the audio data whose playback is paused is non-focus audio, the playback status of the remaining other audio data can be kept unchanged.

[0215] Exemplarily, in some embodiments, if the first audio data serving as the first non-focus audio stops playing, then continuing to play the second audio data at the focus position does not reduce the audio playback effect, and therefore, the playback state of the second audio data at the second virtual space position can continue to be maintained.

[0216] That is to say, in the embodiment of the present application, during the parallel playback of multiple audio data, there may be a situation where audio data is reduced. Among them, as the audio data is reduced, the original focus audio and non-focus audio can be selected to be updated in real time, thereby ensuring the playback effect of the parallel playback of multiple audios.

[0217] Exemplarily, in some embodiments, Fig.17 Schematic diagram of the implementation of switching focus audio proposed in the embodiment of the present application Figure 3 ,like Fig.17 As shown, assuming that the second audio data is the first focus audio, the second virtual space position is a preset center position that is fixedly set and includes two symmetrical space positions. The first audio data is the first non-focus audio, and the corresponding first virtual space position is the non-focus position corresponding to the first audio data determined based on the correspondence between the scene and the non-focus position. After receiving the stop instruction corresponding to the second audio data, the first audio data can be updated to the first focus audio, and at the same time, the first virtual space position can be updated to a preset center position that is fixedly set and includes two symmetrical space positions. Then, during the playback of the audio data, while stopping the playback of the second audio data, the first virtual space position before the update, that is, the non-focus position corresponding to the first audio data, can be pointed to the updated first virtual space position, that is, the preset center position, and the first audio data can be played in a gradual manner.

[0218] Furthermore, in the embodiments of the present application, Fig.18 Schematic diagram of the implementation process of the multi-audio playback method proposed in the embodiment of the present application Figure 5 ,like Fig.18 As shown, after playing the first audio data based on the fourth virtual space position, playing the second audio data based on the fifth virtual space position, and playing the third audio data based on the third virtual space position, that is, after step 109, the method for playing multiple audios by the multiple audio playback device further includes the following steps:

[0219] Step 114: After receiving a stop instruction corresponding to the first audio data, stop playing the first audio data.

[0220] In an embodiment of the present application, during the process of playing the first audio data, the second audio data and the third audio data in parallel, a stop instruction for the first audio data may also be received, and then, the playing of the first audio data may be stopped in response to the stop instruction.

[0221] It is understandable that, in the embodiment of the present application, when multiple audio data are played in parallel, it is also possible to choose to stop playing any one or more of the audio data based on the received instruction.

[0222] Step 115: When the first audio data is the second focus audio, if the second priority is higher than the third priority, the second audio data is updated to the second focus audio, and the fifth virtual space position is updated according to the preset center position.

[0223] In an embodiment of the present application, when a stop instruction corresponding to the first audio data is received and the playing of the first audio data is stopped, when the first audio data is the second focus audio, if the second priority corresponding to the second audio data is higher than the third priority corresponding to the third audio data, then the second audio data can be selected to be updated to the second focus audio, and the fifth virtual space position can be updated according to the preset center position.

[0224] It can be understood that, in an embodiment of the present application, if the audio data that is paused is the focus audio, then one audio data can be determined from the remaining other audio data as the new focus audio, thereby improving the audio playback effect.

[0225] Exemplarily, in some embodiments, if the first audio data as the second focus audio is stopped from playing, at this time, continuing to play the second audio data and the third audio data at the non-focus position will reduce the audio playback effect to a certain extent. Therefore, the priorities of the remaining second audio data and the third audio data can be compared to further re-determine a new focus audio. For example, when the second priority is higher than the third priority, the second audio data can be re-used as the new focus audio, that is, the second audio data can be updated to the second focus audio. At the same time, the fifth virtual space position corresponding to the second audio data can also be determined as the preset center position.

[0226] Step 116: Play the second audio data based on the updated fifth virtual space position; and continue to play the third audio data based on the third virtual space position.

[0227] In an embodiment of the present application, when the first audio data is the second focus audio, if the second priority is higher than the third priority, the second audio data is updated to the second focus audio. After the fifth virtual space position is updated according to the preset center position, the second audio data can be further played based on the updated fifth virtual space position; and the third audio data continues to be played based on the third virtual space position.

[0228] It should be noted that, in an embodiment of the present application, after stopping the playback of the original focus audio (such as the first audio data), in order to ensure the parallel playback effect of multiple audios, a new focus audio can be re-determined based on the priority corresponding to the remaining audio data, and the new focus audio (such as the second audio data) can be played at the updated focus position (such as the updated fifth virtual space position), while the remaining non-focus audio (third audio data) continues to play at the original non-focus position (such as the third virtual space position).

[0229] Furthermore, in the embodiments of the present application, Fig.19 Schematic diagram of the implementation process of the multi-audio playback method proposed in the embodiment of the present application Figure 6 ,like Fig.19 As shown, after receiving the stop instruction corresponding to the first audio data, after stopping playing the first audio data, that is, after step 114, the method for playing multiple audios by the multiple audio playing device further includes the following steps:

[0230] Step 117: When the first audio data is the second focus audio, if the second priority is lower than or equal to the third priority, update the third audio data to the second focus audio, and update the third virtual space position according to the preset center position.

[0231] In an embodiment of the present application, when a stop instruction corresponding to the first audio data is received and the playing of the first audio data is stopped, when the first audio data is the second focus audio, if the second priority corresponding to the second audio data is lower than or equal to the third priority corresponding to the third audio data, then the third audio data can be selected to be updated to the second focus audio, and the third virtual space position can be updated according to the preset center position.

[0232] It can be understood that, in an embodiment of the present application, if the audio data that is paused is the focus audio, then one audio data can be determined from the remaining other audio data as the new focus audio, thereby improving the audio playback effect.

[0233] Exemplarily, in some embodiments, if the first audio data as the second focus audio is stopped from playing, at this time, continuing to play the second audio data and the third audio data at the non-focus position will reduce the audio playback effect to a certain extent. Therefore, the priorities of the remaining second audio data and the third audio data can be compared to further re-determine a new focus audio. For example, when the second priority is lower than or equal to the third priority, the third audio data can be re-used as the new focus audio, that is, the third audio data can be updated to the second focus audio, and at the same time, the third virtual space position corresponding to the third audio data can also be determined as the preset center position.

[0234] Step 118: Continue to play the second audio data based on the fifth virtual space position; and play the third audio data based on the updated third virtual space position.

[0235] In an embodiment of the present application, when the first audio data is the second focus audio, if the second priority is lower than or equal to the third priority, the third audio data is updated to the second focus audio. After the third virtual space position is updated according to the preset center position, the second audio data can be further played based on the fifth virtual space position; and the third audio data is played based on the updated third virtual space position.

[0236] It should be noted that, in an embodiment of the present application, after stopping the playback of the original focus audio (such as the first audio data), in order to ensure the parallel playback effect of multiple audios, a new focus audio can be re-determined based on the priority corresponding to the remaining audio data, and played at the updated focus position (such as the updated third virtual space position), while the remaining non-focus audio (second audio data) continues to play at the original non-focus position (such as the fifth virtual space position).

[0237] Furthermore, in the embodiments of the present application, Fig. 20Schematic diagram of the implementation process of the multi-audio playback method proposed in the embodiment of the present application Figure 7 ,like Fig. 20 As shown, after receiving the stop instruction corresponding to the first audio data, after stopping playing the first audio data, that is, after step 114, the method for playing multiple audios by the multiple audio playing device further includes the following steps:

[0238] Step 119: When the first audio data is the second non-focus audio, continue to play the second audio data based on the fifth virtual space position; and continue to play the third audio data based on the third virtual space position.

[0239] In an embodiment of the present application, when a stop instruction corresponding to the first audio data is received and the playing of the first audio data is stopped, when the first audio data is the second non-focus audio, you can choose to continue playing the second audio data based on the fifth virtual space position; and continue playing the third audio data based on the third virtual space position.

[0240] It can be understood that, in an embodiment of the present application, if the audio data whose playback is paused is non-focus audio, the playback status of the remaining other audio data can be kept unchanged.

[0241] Exemplarily, in some embodiments, if the first audio data serving as the second non-focus audio stops playing, at this time, continuing to play the second audio data or the third audio data at the focus position will not reduce the audio playback effect. Therefore, the playback state of the second audio data at the fifth virtual space position and the third audio data at the third virtual space position can continue to be maintained.

[0242] That is to say, in the embodiment of the present application, during the parallel playback of multiple audio data, there may be a situation where audio data is reduced. Among them, as the audio data is reduced, the original focus audio and non-focus audio can be selected to be updated in real time, thereby ensuring the playback effect of the parallel playback of multiple audios.

[0243] Exemplarily, in some embodiments, Fig.21 Schematic diagram of the implementation of switching focus audio proposed in the embodiment of the present application Figure 4 ,like Fig.21As shown, it is assumed that the second audio data is the second focus audio, and the fifth virtual space position is a preset center position that is fixedly set and includes two symmetrical space positions. The first audio data and the third audio data are the second non-focus audio, and the corresponding fourth virtual space position is the non-focus position corresponding to the first audio data determined based on the corresponding relationship between the scene and the non-focus position, and the corresponding third virtual space position is the non-focus position corresponding to the third audio data determined based on the corresponding relationship between the scene and the non-focus position. After receiving the stop instruction corresponding to the second audio data, if the first priority is higher than the third priority, the first audio data can be updated to the second focus audio, and at the same time, the fourth virtual space position can be updated to the preset center position that is fixedly set and includes two symmetrical space positions. Then, during the playback of the audio data, while stopping the playback of the second audio data, the fourth virtual space position before the update, that is, the non-focus position corresponding to the first audio data, can be pointed to the updated fourth virtual space position, that is, the preset center position, and the first audio data can be played in a gradual manner; while the third audio data continues to be played at the corresponding third virtual space position, that is, the non-focus position corresponding to the third audio data.

[0244] Furthermore, in the embodiments of the present application, in the process of playing multiple audio data in parallel, there are both cases where new audio data is added and cases where audio data is reduced. Accordingly, with the addition or reduction of audio data, the focus audio and non-focus audio therein will also change and update. At the same time, the virtual space position corresponding to each audio data may also change and update. Ultimately, the playback process of each audio data may have different playback methods such as playing in a gradual-in manner, playing in a gradual-out manner, and playing directly.

[0245] It should be noted that in the embodiments of the present application, in the process of changing and updating the virtual space position and playback mode of each audio data, the motion state of each audio data can be defined, and then the change and update of the virtual space position and playback mode can be completed based on the motion state of each audio data.

[0246] Exemplarily, in some embodiments, the audio Track (audio data) in the holographic audio may include five motion states: an initial state Start, a fade-in state FadeIn, a fade-out state FadeOut, a focus state Center, and a non-focus state Home.

[0247] Exemplarily, in some embodiments, Table 2 is a description of the change in the motion state of audio data, wherein the audio Track is in the Start state when it is created, and changes to the FadeIn state after the audio Track seizes the holographic focus. At this time, the sound and image of the audio Track will move from the Home position to the Center position. When the sound and image of the audio Track moves to the Center position, the Track state changes to Center. If the holographic focus of the Track is seized, the Track state changes to FadeOut, and its sound and image moves from the Center position to the Home position to which the Track belongs. When the Track moves to the Home position, the Track state changes to Home.

[0248] Table 2

[0249]

[0250] Exemplarily, in some embodiments, Fig. 22 This is a schematic diagram of the transformation of the motion state of the audio data proposed in the embodiment of the present application, such as Fig. 22 As shown, the audioTrack structure can be modified to add necessary information for audio and video control, such as scene category, priority, creation time, etc. Before playing, the motion state corresponding to the audio data is Start. Then, based on parameters such as scene category, priority, and creation time, the focused audio with Focus and the non-focus audio without Focus can be determined from multiple audio data.

[0251] For example, in some embodiments, the focus audio to be played can be rendered and played in a fade-in manner, that is, for the focus audio, the corresponding motion state can first change from Start to FadeIn, and then from FadeIn to Center. When the spatial position corresponding to the focus audio gradually changes and enters the Center range, it can be considered that the corresponding motion state has changed from FadeIn to Center.

[0252] For example, in some embodiments, for the focused audio being played, when the motion state is Center, if the Focus is seized, that is, the focused audio changes to non-focused audio, then it can be rendered and played in a fade-out manner. At this time, the corresponding motion state can first change from Center to FadeOut, and then from FadeOut to Home. Among them, when the spatial position corresponding to the focused audio gradually changes and enters the Home range, it can be considered that the corresponding motion state has changed from FadeOut to Home.

[0253] For example, in some embodiments, for the focused audio being played, when the motion state is FadeIn, if the Focus is seized, that is, the focused audio changes to non-focused audio, then it can be rendered and played in a fade-out manner. At this time, the corresponding motion state can first change from FadeIn to FadeOut, and then from FadeOut to Home. Among them, when the spatial position corresponding to the focused audio gradually changes to enter the Home range, it can be considered that the corresponding motion state has changed from FadeOut to Home.

[0254] Exemplarily, in some embodiments, non-focus audio to be played can be directly rendered and played, that is, for non-focus audio, the corresponding motion state can be directly changed from Start to Home.

[0255] For example, in some embodiments, for the non-focus audio being played, when the motion state is Home, if the Focus is seized, that is, the non-focus audio is changed to the focus audio, then it can be rendered and played in a gradual manner. At this time, the corresponding motion state can first change from Home to FadeIn, and then from FadeIn to Center. Among them, when the spatial position corresponding to the focus audio gradually changes and enters the Center range, it can be considered that the corresponding motion state has changed from FadeIn to Center.

[0256] For example, in some embodiments, for the non-focus audio being played, when the motion state is FadeOut, if the Focus is seized, that is, the non-focus audio is changed to the focus audio, then it can be rendered and played in a gradual manner. At this time, the corresponding motion state can first change from FadeOut to FadeIn, and then from FadeIn to Center. Among them, when the spatial position corresponding to the focus audio gradually changes and enters the Center range, it can be considered that the corresponding motion state has changed from FadeIn to Center.

[0257] It should be noted that, in the embodiment of the present application, the Center range and the Home range may be pre-set, wherein the Center range and the Home range may be different virtual space ranges. The Center range may be a range closer to the center of the virtual space than the Home range.

[0258] Exemplarily, in some embodiments, based on the above Fig. 22 , Fig.23 Schematic diagram of the implementation of multi-audio playback proposed in the embodiment of the present application Figure 1 ,like Fig.23As shown, during the parallel playback of multiple audios, for the input audio data played in parallel, the focus audio FocusTrack in the current state can be obtained according to the priority and creation time of each Track (audio data) (step 201), and then the motion state of each Track itself is updated according to the current FocusTrack and the spatial position and motion state of each Track itself (step 202), and after the Track updates its own motion state, it updates its own spatial position (step 203), and then the position coordinates of each Track are set in the sound renderer (step 204), and finally the data of the Track is given to the sound renderer for sound rendering, and the dual-channel data is output (step 205). The dual-channel data output by the final sound renderer includes the position effect of each Track in the virtual space.

[0259] That is to say, in the embodiment of the present application, after rendering, what is finally played is the two-channel data after 3D mixing.

[0260] Exemplarily, in some embodiments, Fig.24 Schematic diagram of the implementation of multi-audio playback proposed in the embodiment of the present application Figure 2 ,like Fig.24 As shown, when new audio data is added during the process of playing multiple audios, the update of the focus audio can be completed first. Among them, the priority of the newly added audio data and the priority of the original focus audio can be compared first (step 301). If the priority of the newly added audio data is higher than the priority of the original focus audio, then the newly added audio data can be updated to the new focus audio (step 302); if the priority of the newly added audio data is lower than the priority of the original focus audio, then the original focus audio can be maintained, and the newly added audio data can be used as non-focus audio (step 303); if the priority of the newly added audio data is equal to the priority of the original focus audio, then the creation time of the newly added audio data and the creation time of the original focus audio can be further compared (step 304), and when the creation time of the newly added audio data is greater than or equal to the creation time of the original focus audio, the newly added audio data is updated to the new focus audio (step 302), otherwise the original focus audio is maintained (step 303).

[0261] To sum up, the multi-audio playback method proposed through the above steps can determine the focused audio with Focus and the non-focused audio without Focus among multiple audio data through the scene category corresponding to the audio data, and then combine the focused audio and the non-focused audio to determine the corresponding virtual space position of each audio data in the virtual sound field space, so that different audio data can be allocated to different positions with complementary interference for playback.

[0262] That is to say, the multi-audio playback method proposed in the embodiment of the present application can distribute different audio data in different virtual space positions, realize the separation of sound hearing through the different virtual space positions, and then achieve the clarity and integrity of the audio content when different audio data are played simultaneously.

[0263] It should be noted that in the time limit process of the multi-audio playback method proposed in the embodiment of the present application, the audio data can be first divided into different scene categories, and different scene categories have different priorities; then, according to the priority (and creation time), it can be determined which audio data is the focus audio and which audio data is the non-focus audio. Finally, by combining the focus audio and the non-focus audio, different audio data can be distributed in different virtual space locations for rendering and playback.

[0264] It can be understood that the multi-audio playback method proposed in the embodiment of the present application is not limited to the headphone playback scenario, but can also be applied to the speaker playback scenario, and the present application does not make specific limitations.

[0265] The embodiment of the present application provides a multi-audio playback method, determining a first scene category corresponding to the first audio data and a second scene category corresponding to the second audio data; determining a first focus audio and a first non-focus audio in the first audio data and the second audio data according to the first scene category and the second scene category; determining a first virtual space position corresponding to the first audio data and a second virtual space position corresponding to the second audio data based on the first focus audio and the first non-focus audio; playing the first audio data based on the first virtual space position, and playing the second audio data based on the second virtual space position. It can be seen that in the embodiment of the present application, in the process of multi-audio parallel playback, the scene categories corresponding to different audio data can be used to determine the focus audio and non-focus audio in multiple audio data, so that each audio data can be assigned a corresponding virtual space position based on the focus audio and the non-focus audio, so that different audio data can be played at different virtual space positions. That is to say, the multi-audio playback method proposed in the embodiment of the present application can not only play each audio data completely, but also assign non-interfering virtual space positions to different audio data for rendering and playback, so as to take into account the clarity and integrity of the audio content and effectively improve the playback effect of multi-audio.

[0266] Based on the above embodiments, the embodiments of the present application provide a multi-audio playback method, which can solve the problem of multiple sounds (audio data) interfering with each other when played simultaneously, or the sound being incomplete, making the sound experience more intelligent and immersive. Among them, the multi-audio playback method proposed in the embodiments of the present application can distribute different sounds played simultaneously to different virtual spaces according to the priority and the start time through an innovative audio-visual control algorithm, and perform sound rendering in the virtual space with the help of a sound rendering algorithm, and finally render different sounds to different positions without interfering with each other.

[0267] It should be noted that in the embodiments of the present application, when the spatial audio rendering algorithm renders multi-channel audio data, it is often assumed that multiple virtual speakers are arranged in the virtual listening space, and the data of each channel is fed to each virtual speaker. After being processed by the rendering algorithm, the sound seems to be emitted from the position of the virtual speaker.

[0268] It should be noted that, in the embodiments of the present application, the basic sound unit processed by the spatial audio rendering algorithm generally includes audio data and sound position information (ie, the virtual spatial position corresponding to the audio data).

[0269] Furthermore, the core content of the image control algorithm proposed in the embodiment of the present application may include: dividing the audio stream (audio data) into different scene categories, and different scene categories correspond to different priorities. You can first determine which audio stream enters the Center position as the focus of the whole system and which audio streams enter the Home position based on the priority. Among them, when the priorities are the same, you can choose to judge according to the order of playback time (creation time), for example, the audio data played later enters the Center position as the focus of the whole system.

[0270] It should be noted that in the embodiments of the present application, the holographic focus is the focus audio, which is similar to the concept of audio focus in Android. When multiple applications play sounds at the same time, there is only one holographic focus.

[0271] It should be noted that, in the embodiment of the present application, the Center position is the position where the sound and image of the application holding the holographic focus are located.

[0272] It should be noted that, in the embodiment of the present application, when the holographic focus of the application is seized, the sound and image of the application will gradually fade out from the Center position to the Home position, and each audio scene has its own Home position.

[0273] Furthermore, in an embodiment of the present application, the audio streams (audio data) created by the system and third-party applications can be classified into a certain scene, such as music, listening, video, game voice, ringing, alarm, etc. The audio streams of the same audio scene (scene category) have the same effect performance.

[0274] For example, in some embodiments, different applications may be classified first, and then the scene category of the created audio data may be determined according to the scene type for which the application is used, wherein different scene categories may have the same or different priorities.

[0275] That is to say, in the embodiment of the present application, for different scene categories, corresponding priorities can be preset, that is, the corresponding relationship between the preset scenes and the priorities can be determined. Among them, the corresponding relationship between the preset scenes and the priorities can be continuously adjusted and updated.

[0276] For example, in some embodiments, based on the correspondence between the preset scenes and the priorities shown in Table 1 above, the smaller the value of the priority, the higher the priority. For different scene categories, the corresponding priorities may be the same or different. For example, the priorities corresponding to the ringtone and the alarm are the same, while the priorities corresponding to the ringtone and the video are different, and the priority of the video is higher than that of the ringtone.

[0277] Further, in the embodiments of the present application, different scene categories may correspond to different audio data, and the corresponding relationship between the scene and the non-focus position may be set for different scene categories. Specifically, for each scene category, a non-focus position (i.e., Home position) corresponding to the scene category may be set, and when the audio data corresponding to the scene category is used as non-focus audio, the non-focus position may be used to determine the virtual space position of the audio data.

[0278] It should be noted that in the embodiments of the present application, the correspondence between the scene and the non-focus position can be set based on the preset distribution principle. Among them, the core of the preset distribution principle is that the Home position (non-focus position) of different audio scenes is different, that is, different scene categories correspond to different non-focus positions. Among them, the overall Home position layout must ensure that the sound position of the entire sound field distribution can be distributed as balanced as possible in the front, back, left and right.

[0279] For example, in some embodiments, the distribution of the Home position can be multiple. Based on the preset distribution principle, for audio scenes (scene categories) with rich audio content, such as music and video, the Home position can be set in a central symmetrical position, which will have a better stereo effect. For audio scenes with relatively simple sound content, such as Internet calls, listening to books, prompt sounds, etc., the Home position can be set in a position biased in a certain direction.

[0280] For example, in some embodiments, Figure 2 As shown, you can set the non-focal position (Home position) corresponding to different scene categories in the virtual sound field space, where the position of the headphones is the center of the virtual sound field space. Based on the preset distribution principle, for scene categories such as music and video, you can choose to set the Home position in a central symmetrical position, which will have a better stereo effect. For scene categories such as online calls, listening to books, notification sounds, etc., you can choose to set the Home position in a position biased in a certain direction.

[0281] Furthermore, in an embodiment of the present application, in the process of changing and updating the virtual space position and playback mode of each audio data, the motion state of each audio data can be defined, and then the change and update of the virtual space position and playback mode can be completed based on the motion state of each audio data.

[0282] Exemplarily, in some embodiments, the audio Track (audio data) in the holographic audio may include five motion states: an initial state Start, a fade-in state FadeIn, a fade-out state FadeOut, a focus state Center, and a non-focus state Home.

[0283] Exemplarily, in some embodiments, Table 2 is a description of the change in the motion state of audio data, wherein the audio Track is in the Start state when it is created, and changes to the FadeIn state after the audio Track seizes the holographic focus. At this time, the sound and image of the audio Track will move from the Home position to the Center position. When the sound and image of the audio Track moves to the Center position, the Track state changes to Center. If the holographic focus of the Track is seized, the Track state changes to FadeOut, and its sound and image moves from the Center position to the Home position to which the Track belongs. When the Track moves to the Home position, the Track state changes to Home.

[0284] Exemplarily, in some embodiments, the audioTrack structure can be modified to add necessary information for audio and video control, such as scene category, priority, creation time, etc. Before playing, the motion state corresponding to the audio data is Start, and then, based on parameters such as scene category, and / or priority and / or creation time, the focused audio with Focus and the non-focus audio without Focus can be determined from multiple audio data.

[0285] For example, in some embodiments, the focus audio to be played can be rendered and played in a fade-in manner, that is, for the focus audio, the corresponding motion state can first change from Start to FadeIn, and then from FadeIn to Center. When the spatial position corresponding to the focus audio gradually changes and enters the Center range, it can be considered that the corresponding motion state has changed from FadeIn to Center.

[0286] For example, in some embodiments, for the focused audio being played, when the motion state is Center, if the Focus is seized, that is, the focused audio changes to non-focused audio, then it can be rendered and played in a fade-out manner. At this time, the corresponding motion state can first change from Center to FadeOut, and then from FadeOut to Home. Among them, when the spatial position corresponding to the focused audio gradually changes and enters the Home range, it can be considered that the corresponding motion state has changed from FadeOut to Home.

[0287] For example, in some embodiments, for the focused audio being played, when the motion state is FadeIn, if the Focus is seized, that is, the focused audio changes to non-focused audio, then it can be rendered and played in a fade-out manner. At this time, the corresponding motion state can first change from FadeIn to FadeOut, and then from FadeOut to Home. Among them, when the spatial position corresponding to the focused audio gradually changes to enter the Home range, it can be considered that the corresponding motion state has changed from FadeOut to Home.

[0288] Exemplarily, in some embodiments, non-focus audio to be played can be directly rendered and played, that is, for non-focus audio, the corresponding motion state can be directly changed from Start to Home.

[0289] For example, in some embodiments, for the non-focus audio being played, when the motion state is Home, if the Focus is seized, that is, the non-focus audio is changed to the focus audio, then it can be rendered and played in a gradual manner. At this time, the corresponding motion state can first change from Home to FadeIn, and then from FadeIn to Center. Among them, when the spatial position corresponding to the focus audio gradually changes and enters the Center range, it can be considered that the corresponding motion state has changed from FadeIn to Center.

[0290] For example, in some embodiments, for the non-focus audio being played, when the motion state is FadeOut, if the Focus is seized, that is, the non-focus audio is changed to the focus audio, then it can be rendered and played in a gradual manner. At this time, the corresponding motion state can first change from FadeOut to FadeIn, and then from FadeIn to Center. Among them, when the spatial position corresponding to the focus audio gradually changes and enters the Center range, it can be considered that the corresponding motion state has changed from FadeIn to Center.

[0291] It should be noted that, in the embodiment of the present application, the Center range and the Home range may be pre-set, wherein the Center range and the Home range may be different virtual space ranges. The Center range may be a range closer to the center of the virtual space than the Home range.

[0292] Furthermore, in the embodiments of the present application, Fig.25 Schematic diagram of the implementation of multi-audio playback proposed in the embodiment of the present application Figure 3 ,like Fig.25As shown, taking the multi-audio data including the first audio data and the second audio data as an example, after obtaining the multi-audio parallel playback instruction of the first audio data and the second audio data, the scene category corresponding to the first audio data and the second audio data can be determined first, that is, the first scene category and the second scene category (step 401); then, the priority corresponding to the first audio data and the second audio data is further determined based on the correspondence between the preset scene and the priority, that is, the first priority and the second priority (step 402); then, according to the priority (and creation time) of the first audio data and the second audio data, the focus audio FocusTrack and the non-focus audio in the current state, that is, the first focus audio and the first non-focus audio (step 403); then, according to the spatial position and motion state of the first focus audio and the first non-focus audio, the state and the corresponding virtual space position of the first audio data and the second audio data can be updated, thereby obtaining the first virtual space position corresponding to the first audio data and the second virtual space position corresponding to the second audio data (step 404). Finally, the position coordinates of each Track can be set in the sound renderer, and finally the data of the Track is given to the sound renderer for sound rendering, and the dual-channel data is output (step 405).

[0293] Exemplarily, in an embodiment of the present application, in the process of determining and updating the virtual space position and motion state corresponding to the audio data, assuming that the first audio data is determined to be the first focus audio, the motion state corresponding to the first audio data can first be changed from Start to FadeIn, and then from FadeIn to Center. At the same time, the virtual space position corresponding to the first audio data can be changed from a preset starting space position to a preset center position. Among them, when the spatial position corresponding to the first audio data gradually changes and enters the Center range, it can be considered that the corresponding motion state has been changed from FadeIn to Center. The preset center position is in the Center range.

[0294] Exemplarily, in an embodiment of the present application, in the process of determining and updating the virtual space position and motion state corresponding to the audio data, assuming that the second audio data is determined to be the first non-focus audio, then the motion state corresponding to the second audio data can be directly changed from Start to Home, and at the same time, the second virtual space position corresponding to the second audio data can be directly determined by the correspondence between the scene and the non-focus position. Among them, when the spatial position corresponding to the second audio data gradually changes and enters the Home range, it can be considered that the corresponding motion state has changed from Start to Home. The non-focus position corresponding to the second audio data, that is, the second virtual space position is in the Home range.

[0295] Furthermore, in the embodiments of the present application, Fig.26 Schematic diagram of the implementation of multi-audio playback proposed in the embodiment of the present application Figure 4 ,like Fig.26 As shown, taking the third audio data as the newly added audio data as an example, in the process of playing the first audio data and the second audio data in parallel, after obtaining the playback instruction of the third audio data, the third scene category corresponding to the third audio data can be determined first (step 501); then, based on the correspondence between the preset scene and the priority, the third priority corresponding to the third audio data can be further determined (step 502); then, according to the priority update of the first audio data, the second audio data and the third audio data, the focus audio FocusTrack and non-focus audio in the current state, that is, the second focus audio and the second non-focus audio, are obtained (step 503).

[0296] Exemplarily, in some embodiments, assuming that the first audio data is the original first focus audio, if the first priority is higher than the third priority, then it can be considered that the first audio data is still the audio data with the highest priority. At this time, it can be determined that the updated focus audio is still the first audio data, that is, the first audio data is determined as the second focus audio. At the same time, the second audio data and the third audio data can be used as the second non-focus audio.

[0297] Exemplarily, in some embodiments, assuming that the first audio data is the original first focus audio, if the first priority is not higher than the third priority, then the updated focus audio can be determined to be the third audio data, that is, the third audio data is determined to be the second focus audio. Correspondingly, the second audio data and the first audio data can be used as the second non-focus audio.

[0298] It is understandable that in an embodiment of the present application, the first audio data, the second audio data, and the third audio state and the corresponding virtual space position can be updated according to the spatial position and motion state of the second focus audio and the second non-focus audio, thereby obtaining the third virtual space position corresponding to the third audio data, the fourth virtual space position corresponding to the first audio data, and the fifth virtual space position corresponding to the second audio data (step 504). Finally, you can choose to set the position coordinates of each Track in the sound renderer, and finally give the Track data to the sound renderer for sound rendering, and output dual-channel data (step 505).

[0299] Exemplarily, in an embodiment of the present application, in the process of determining and updating the virtual space position and motion state corresponding to the audio data, assuming that the newly added third audio data is updated to the second focus audio, the motion state corresponding to the third audio data can first be changed from Start to FadeIn, and then from FadeIn to Center. At the same time, the virtual space position corresponding to the third audio data can be changed from the preset starting space position to the preset center position. Among them, when the spatial position corresponding to the third audio data gradually changes and enters the Center range, it can be considered that the corresponding motion state has been changed from FadeIn to Center. The preset center position is in the Center range.

[0300] Exemplarily, in an embodiment of the present application, in the process of determining and updating the virtual space position and motion state corresponding to the audio data, assuming that the first audio data as the first focus audio is updated to the second non-focus audio, then the motion state corresponding to the first audio data can first be changed from Center to FadeOut, and then from FadeOut to Home. At the same time, the fourth virtual space position corresponding to the first audio data can be directly determined by the correspondence between the scene and the non-focus position. Among them, when the spatial position corresponding to the first audio data gradually changes to enter the Home range, it can be considered that the corresponding motion state has changed from FadeOut to Home. The non-focus position corresponding to the first audio data, that is, the fourth virtual space position is in the Home range.

[0301] Exemplarily, in an embodiment of the present application, in the process of determining and updating the virtual space position and motion state corresponding to the audio data, assuming that the second audio data as the first non-focus audio is updated to the second non-focus audio, then the motion state corresponding to the second audio data can continue to be maintained as Home, and at the same time, the fifth virtual space position corresponding to the second audio data can continue to be maintained as the corresponding second virtual space position unchanged.

[0302] It can be understood that, in the embodiments of the present application, for the FadeIn state, the preset fade-in path may include a straight path or a curved path, which is not specifically limited in the present application.

[0303] It can be understood that, in the embodiment of the present application, for the FadeOut state, the preset fade-out path may include a straight path or a curved path, which is not specifically limited in the present application.

[0304] To sum up, the multi-audio playback method proposed in the embodiment of the present application can be a set of sound and image control algorithms that can automatically assign different sounds to different virtual space positions. It can distribute different sounds in different virtual space positions with the help of the sound object rendering capability of the spatial audio rendering algorithm, and achieve the separation of sound hearing through the different virtual space positions, thereby achieving the effect of different sounds being able to be heard clearly when played at the same time.

[0305] The embodiment of the present application provides a multi-audio playback method, determining a first scene category corresponding to the first audio data and a second scene category corresponding to the second audio data; determining a first focus audio and a first non-focus audio in the first audio data and the second audio data according to the first scene category and the second scene category; determining a first virtual space position corresponding to the first audio data and a second virtual space position corresponding to the second audio data based on the first focus audio and the first non-focus audio; playing the first audio data based on the first virtual space position, and playing the second audio data based on the second virtual space position. It can be seen that in the embodiment of the present application, in the process of multi-audio parallel playback, the scene categories corresponding to different audio data can be used to determine the focus audio and non-focus audio in multiple audio data, so that each audio data can be assigned a corresponding virtual space position based on the focus audio and the non-focus audio, so that different audio data can be played at different virtual space positions. That is to say, the multi-audio playback method proposed in the embodiment of the present application can not only play each audio data completely, but also assign non-interfering virtual space positions to different audio data for rendering and playback, so as to take into account the clarity and integrity of the audio content and effectively improve the playback effect of multi-audio.

[0306] Based on the above embodiment, in another embodiment of the present application, Fig. 27 This is a schematic diagram of the structure of the multi-audio playback device proposed in the embodiment of the present application, such as Fig. 27 As shown, the multi-audio playback device 10 proposed in the embodiment of the present application may include a determination unit 11, a playback unit 12,

[0307] The determining unit 11 is used to determine a first scene category corresponding to the first audio data and a second scene category corresponding to the second audio data; determine a first focus audio and a first non-focus audio in the first audio data and the second audio data according to the first scene category and the second scene category; and determine a first virtual space position corresponding to the first audio data and a second virtual space position corresponding to the second audio data based on the first focus audio and the first non-focus audio;

[0308] The playing unit 12 is used to play the first audio data based on the first virtual space position, and play the second audio data based on the second virtual space position.

[0309] In the embodiments of the present application, further, Fig.28 This is a schematic diagram of the structure of the electronic device proposed in the embodiment of the present application, such as Fig.28 As shown, the electronic device 20 proposed in the embodiment of the present application may also include a processor 21 and a memory 22 storing executable instructions of the processor 21. Furthermore, the multi-audio playback device 10 may also include a communication interface 23 and a bus 24 for connecting the processor 21, the memory 22 and the communication interface 23.

[0310] In the embodiment of the present application, the processor 21 can be at least one of an Application Specific Integrated Circuit (ASIC), a Digital Signal Processor (DSP), a Digital Signal Processing Device (DSPD), a Programmable Logic Device (PLD), a Field Programmable Gate Array (FPGA), a Central Processing Unit (CPU), a controller, a microcontroller, and a microprocessor. It can be understood that for different devices, the electronic device used to implement the function of the processor can also be other, and the embodiment of the present application is not specifically limited. The multi-audio playback device 10 can also include a memory 22, which can be connected to the processor 21, wherein the memory 22 is used to store executable program code, the program code includes computer operation instructions, and the memory 22 may include a high-speed RAM memory, and may also include a non-volatile memory, for example, at least two disk memories.

[0311] In the embodiment of the present application, the bus 24 is used to connect the communication interface 23, the processor 21 and the memory 22, and the mutual communication between these devices.

[0312] In the embodiment of the present application, the memory 22 is used to store instructions and data.

[0313] Further, in an embodiment of the present application, the above-mentioned processor 21 is used to determine a first scene category corresponding to the first audio data and a second scene category corresponding to the second audio data; determine a first focus audio and a first non-focus audio in the first audio data and the second audio data according to the first scene category and the second scene category; determine a first virtual space position corresponding to the first audio data and a second virtual space position corresponding to the second audio data based on the first focus audio and the first non-focus audio; play the first audio data based on the first virtual space position, and play the second audio data based on the second virtual space position at the same time.

[0314] In practical applications, the memory 22 may be a volatile memory, such as a random access memory (RAM); or a non-volatile memory, such as a read-only memory (ROM), a flash memory, a hard disk (HDD) or a solid-state drive (SSD); or a combination of the above types of memory, and provide instructions and data to the processor 21.

[0315] In addition, each functional module in this embodiment can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The above integrated unit can be implemented in the form of hardware or software functional modules.

[0316] If the integrated unit is implemented in the form of a software function module and is not sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this embodiment is essentially or the part that contributes to the prior art or the whole or part of the technical solution can be embodied in the form of a software product. The computer software product is stored in a storage medium, including several instructions for a computer device (which can be a personal computer, server, or network device, etc.) or a processor to perform all or part of the steps of the method of this embodiment. The aforementioned storage medium includes: U disk, mobile hard disk, read only memory (ROM), random access memory (RAM), disk or optical disk, etc., which can store program code.

[0317] The embodiment of the present application provides a multi-audio playback device and an electronic device, which determines a first scene category corresponding to the first audio data and a second scene category corresponding to the second audio data; according to the first scene category and the second scene category, determines the first focus audio and the first non-focus audio in the first audio data and the second audio data; based on the first focus audio and the first non-focus audio, determines the first virtual space position corresponding to the first audio data and the second virtual space position corresponding to the second audio data; plays the first audio data based on the first virtual space position, and plays the second audio data based on the second virtual space position. It can be seen that in the embodiment of the present application, in the process of multi-audio parallel playback, the scene categories corresponding to different audio data can be used to determine the focus audio and non-focus audio in multiple audio data, so that each audio data can be assigned a corresponding virtual space position based on the focus audio and the non-focus audio, so that different audio data can be played at different virtual space positions. That is to say, the multi-audio playback method proposed in the embodiment of the present application can not only play each audio data completely, but also assign virtual space positions that do not interfere with each other to different audio data for rendering and playback, so as to take into account the clarity and integrity of the audio content and effectively improve the playback effect of multi-audio.

[0318] An embodiment of the present application provides a computer-readable storage medium on which a program is stored. When the program is executed by a processor, the multi-audio playback method as described above is implemented.

[0319] Specifically, the program instructions corresponding to the multi-audio playback method in this embodiment can be stored in a storage medium such as a CD, a hard disk, a USB flash drive, etc. When the program instructions corresponding to the multi-audio playback method in the storage medium are read or executed by an electronic device, the following steps are included:

[0320] Determine a first scene category corresponding to the first audio data and a second scene category corresponding to the second audio data;

[0321] determining first focus audio and first non-focus audio in the first audio data and the second audio data according to the first scene category and the second scene category;

[0322] Determining, based on the first focused audio and the first non-focused audio, a first virtual space position corresponding to the first audio data and a second virtual space position corresponding to the second audio data;

[0323] The first audio data is played based on the first virtual space position, and the second audio data is played based on the second virtual space position.

[0324] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, systems, or computer program products. Therefore, the present application may adopt the form of hardware embodiments, software embodiments, or embodiments in combination with software and hardware. Moreover, the present application may adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage and optical storage, etc.) that contain computer-usable program code.

[0325] The present application is described with reference to implementation flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the process in the flowchart. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0326] These computer program instructions may also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce a manufactured product including an instruction device, which is implemented in the implementation flow diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.

[0327] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operation steps are executed on the computer or other programmable device to produce a computer-implemented process, so that the instructions executed on the computer or other programmable device provide for implementing the process described in the flowchart. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.

[0328] The above description is only a preferred embodiment of the present application and is not intended to limit the protection scope of the present application.

Claims

1. A multi-audio playback method, characterized in that: The method comprises: Determine a first scene category corresponding to the first audio data and a second scene category corresponding to the second audio data; determining first focus audio and first non-focus audio in the first audio data and the second audio data according to the first scene category and the second scene category; Determining, based on the first focused audio and the first non-focused audio, a first virtual space position corresponding to the first audio data and a second virtual space position corresponding to the second audio data; The first audio data is played based on the first virtual space position, and the second audio data is played based on the second virtual space position.

2. The method according to claim 1, characterized in that The determining, according to the first scene category and the second scene category, a first focus audio and a first non-focus audio in the first audio data and the second audio data comprises: Determine a first priority according to the first scene category, and determine a second priority according to the second scene category; The first focus audio and the first non-focus audio are determined in the first audio data and the second audio data according to the first priority and the second priority.

3. The method according to claim 2, characterized in that The determining of the first priority according to the first scene category and the determining of the second priority according to the second scene category includes: Based on a correspondence between preset scenarios and priorities, the first priority corresponding to the first scenario category and the second priority corresponding to the second priority are determined.

4. The method according to claim 3, characterized in that The determining, according to the first priority and the second priority, the first focus audio and the first non-focus audio in the first audio data and the second audio data comprises: When the first priority is higher than the second priority, determining the first audio data as the first focus audio, and determining the second audio data as the first non-focus audio; When the first priority is lower than the second priority, determining the second audio data as the first focus audio, and determining the first audio data as the first non-focus audio; When the first priority is equal to the second priority, the first focus audio and the first non-focus audio are determined in the first audio data and the second audio data according to a first creation time corresponding to the first audio data and a second creation time corresponding to the second audio data.

5. The method according to claim 4, characterized in that The determining the first focus audio and the first non-focus audio in the first audio data and the second audio data according to the first creation time corresponding to the first audio data and the second creation time corresponding to the second audio data includes: When the first creation time is greater than or equal to the second creation time, determining the first audio data as the first focus audio, and determining the second audio data as the first non-focus audio; When the first creation time is less than the second creation time, the second audio data is determined as the first focus audio, and the first audio data is determined as the first non-focus audio.

6. The method according to claim 4 or 5, characterized in that: The method further comprises: Determine the correspondence between the scene and the non-focus position.

7. The method according to claim 6, characterized in that The determining, based on the first focused audio and the first non-focused audio, a first virtual space position corresponding to the first audio data and a second virtual space position corresponding to the second audio data comprises: When the first audio data is the first focused audio and the second audio data is the first non-focused audio, determining the non-focused position corresponding to the second audio data according to the correspondence between the scene and the non-focused position and the second scene category; A preset center position is determined as the first virtual space position, and a non-focus position corresponding to the second audio data is determined as the second virtual space position.

8. The method according to claim 6, characterized in that The determining, based on the first focused audio and the first non-focused audio, a first virtual space position corresponding to the first audio data and a second virtual space position corresponding to the second audio data comprises: In a case where the second audio data is the first focused audio and the first audio data is the first non-focused audio, determining the non-focused position corresponding to the first audio data according to the correspondence between the scene and the non-focused position and the first scene category; A preset center position is determined as the second virtual space position, and a non-focus position corresponding to the first audio data is determined as the first virtual space position.

9. The method according to claim 7 or 8, characterized in that: The determining the correspondence between the scene and the non-focus position includes: A preset correspondence between the scene and the non-focus position is obtained.

10. The method according to claim 7 or 8, characterized in that: The determining the correspondence between the scene and the non-focus position includes: Obtaining a setting instruction corresponding to the first application; The non-focus position corresponding to the audio category created by the first application is determined according to the setting instruction to determine the corresponding relationship between the scene and the non-focus position.

11. The method according to claim 7 or 8, characterized in that: The determining the correspondence between the scene and the non-focus position includes: Obtaining custom information of the first application; The non-focus position corresponding to the audio category created by the first application is determined according to the customized information to determine the corresponding relationship between the scene and the non-focus position.

12. The method according to any one of claims 1 to 3, characterized in that: The playing the first audio data based on the first virtual space position includes: In the case where the first audio data is the first focus audio, the first audio data is played in a fade-in manner based on the first virtual space position; or the first audio data is played directly at the first virtual space position; In a case where the first audio data is the first non-focus audio, the first audio data is played directly at the first virtual space position.

13. The method according to any one of claims 1 to 3, characterized in that: The playing the second audio data based on the second virtual space position includes: In the case where the second audio data is the first focus audio, the second audio data is played in a fade-in manner based on the second virtual space position; or the second audio data is played directly at the second virtual space position; In the case where the second audio data is the first non-focus audio, the second audio data is played directly at the second virtual space position.

14. The method according to claim 6, characterized in that After playing the first audio data based on the first virtual space position and playing the second audio data based on the second virtual space position, the method further includes: receiving a play instruction corresponding to the third audio data; Determining a third scene category corresponding to the third audio data; Determine a second focus audio from the first audio data, the second audio data, and the third audio data according to the third scene category, and simultaneously determine a second non-focus audio; Based on the second focused audio and the second non-focused audio, determining a third virtual space position corresponding to the third audio data, a fourth virtual space position corresponding to the first audio data, and a fifth virtual space position corresponding to the second audio data; The first audio data is played based on the fourth virtual space position, the second audio data is played based on the fifth virtual space position; and the third audio data is played based on the third virtual space position.

15. The method according to claim 14, characterized in that The step of determining a second focus audio from the first audio data, the second audio data, and the third audio data according to the third scene category, and determining a second non-focus audio at the same time, comprises: Determining a third priority corresponding to the third scene category based on the correspondence between the preset scenes and the priorities; When the first audio data is the first focus audio and the first priority is higher than the third priority, determining the first audio data as the second focus audio, and determining the second audio data and the third audio data as the second non-focus audio; When the first audio data is the first focus audio and the first priority is lower than or equal to the third priority, the third audio data is determined as the second focus audio, and the first audio data and the second audio data are determined as the second non-focus audio.

16. The method according to claim 14, characterized in that The step of determining a second focus audio from the first audio data, the second audio data, and the third audio data according to the third scene category, and determining a second non-focus audio at the same time, comprises: Determining a third priority corresponding to the third scene category based on the correspondence between the preset scenes and the priorities; When the second audio data is the first focus audio and the second priority is higher than the third priority, the second audio data is determined as the second focus audio, and the first audio data and the third audio data are determined as the second non-focus audio; When the second audio data is the first focus audio and the second priority is lower than or equal to the third priority, the third audio data is determined as the second focus audio, and the first audio data and the second audio data are determined as the second non-focus audio.

17. The method according to claim 15 or 16, characterized in that The determining, based on the second focused audio and the second non-focused audio, a third virtual space position corresponding to the third audio data, a fourth virtual space position corresponding to the first audio data, and a fifth virtual space position corresponding to the second audio data comprises: In a case where the third audio data is the second focus audio, determining a preset center position as the third virtual space position; In the case where the third audio data is the second non-focus audio, the non-focus position corresponding to the third audio data is determined according to the correspondence between the scene and the non-focus position and the third scene category, and the non-focus position corresponding to the third audio data is determined as the third virtual space position.

18. The method according to claim 15, characterized in that The determining, based on the second focused audio and the second non-focused audio, a third virtual space position corresponding to the third audio data, a fourth virtual space position corresponding to the first audio data, and a fifth virtual space position corresponding to the second audio data comprises: In a case where the first audio data is the second focus audio, determining the first virtual space position as the fourth virtual space position; In the case where the first audio data is the second non-focus audio, the non-focus position corresponding to the first audio data is determined according to the correspondence between the scene and the non-focus position and the first scene category, and the non-focus position corresponding to the first audio data is determined as the fourth virtual space position.

19. The method according to claim 16, characterized in that The determining, based on the second focused audio and the second non-focused audio, a third virtual space position corresponding to the third audio data, a fourth virtual space position corresponding to the first audio data, and a fifth virtual space position corresponding to the second audio data comprises: In a case where the second audio data is the second focus audio, determining the second virtual space position as the fifth virtual space position; In the case where the second audio data is the second non-focus audio, the non-focus position corresponding to the second audio data is determined according to the correspondence between the scene and the non-focus position and the second scene category, and the non-focus position corresponding to the second audio data is determined as the fifth virtual space position.

20. The method according to claim 17, characterized in that The playing the third audio data based on the third virtual space position includes: In the case where the third audio data is the second focus audio, the third audio data is played in a fade-in manner based on the third virtual space position; or the third audio data is played directly at the third virtual space position; In a case where the third audio data is the second non-focus audio, the third audio data is directly played at the third virtual space position.

21. The method according to claim 18, characterized in that The playing the first audio data based on the fourth virtual space position includes: When the first audio data is the second focus audio, directly playing the first audio data at the fourth virtual space position; In the case where the first audio data is the second non-focus audio, the first audio data is played in a fade-out manner based on the fourth virtual space position.

22. The method according to claim 19, characterized in that The playing the second audio data based on the fifth virtual space position includes: When the second audio data is the second focus audio, directly playing the second audio data at the fifth virtual space position; In the case where the second audio data is the second non-focus audio, the second audio data is played in a fade-out manner based on the fifth virtual space position.

23. The method according to any one of claims 12, 13, and 20, characterized in that: The first audio data, the second audio data, or the third audio data is used as the target audio data, and the first virtual space position, the second virtual space position, or the third virtual space position is used as the end space position, the method further includes: According to a preset fade-in path, the target audio data is played in a fade-in manner from a preset starting spatial position to the end spatial position; wherein the preset fade-in path includes a straight path or a curved path.

24. The method according to claim 21, characterized in that The step of playing the first audio data in a fade-out manner based on the fourth virtual space position includes: According to a preset fade-out path, the first audio data is played in a fade-out manner from the first virtual space position to the fourth virtual space position; wherein the preset fade-out path includes a straight path or a curved path.

25. The method according to claim 22, characterized in that The step of playing the second audio data in a fade-out manner based on the fifth virtual space position includes: According to a preset fade-out path, the second virtual space position points to the fifth virtual space position, and the second audio data is played in a fade-out manner; wherein the preset fade-out path includes a straight path or a curved path.

26. The method according to claim 7 or 8, characterized in that After playing the first audio data based on the first virtual space position and playing the second audio data based on the second virtual space position, the method further includes: After receiving a stop instruction corresponding to the first audio data, stop playing the first audio data; In a case where the first audio data is the first focus audio, updating the second audio data to the first focus audio, and updating the second virtual space position according to the preset center position; The second audio data is played based on the updated second virtual space position.

27. The method according to claim 26, characterized in that The method further comprises: When the first audio data is the first non-focus audio, continue to play the second audio data based on the second virtual space position.

28. The method according to claim 15 or 16, characterized in that The first audio data is played based on the fourth virtual space position, and the second audio data is played based on the fifth virtual space position; After playing the third audio data based on the third virtual space position, the method further includes: After receiving a stop instruction corresponding to the first audio data, stop playing the first audio data; In the case where the first audio data is the second focus audio, if the second priority is higher than the third priority, updating the second audio data to the second focus audio, and updating the fifth virtual space position according to the preset center position; The second audio data is played based on the updated fifth virtual space position; and the third audio data is continued to be played based on the third virtual space position.

29. The method according to claim 28, characterized in that The method further comprises: In the case where the first audio data is the second focus audio, if the second priority is lower than or equal to the third priority, updating the third audio data to the second focus audio, and updating the third virtual space position according to the preset center position; Continue to play the second audio data based on the fifth virtual space position; and play the third audio data based on the updated third virtual space position.

30. The method according to claim 29, characterized in that The method further comprises: In a case where the first audio data is the second non-focus audio, the second audio data continues to be played based on the fifth virtual space position; and the third audio data continues to be played based on the third virtual space position.

31. A multi-audio playback device, characterized in that: The multi-audio playback device comprises: a determination unit, a playback unit, The determining unit is used to determine a first scene category corresponding to the first audio data and a second scene category corresponding to the second audio data; determine a first focused audio and a first non-focused audio in the first audio data and the second audio data according to the first scene category and the second scene category; and determine a first virtual space position corresponding to the first audio data and a second virtual space position corresponding to the second audio data based on the first focused audio and the first non-focused audio; The playing unit is used to play the first audio data based on the first virtual space position, and play the second audio data based on the second virtual space position.

32. An electronic device, characterized in that: The electronic device comprises a processor and a memory storing instructions executable by the processor, and when the instructions are executed by the processor, the method according to any one of claims 1 to 30 is implemented.

33. A computer-readable storage medium having a program stored thereon, characterized in that: When the program is executed by a processor, the method according to any one of claims 1 to 30 is implemented.