Audio processing method and electronic device

CN119999221APending Publication Date: 2025-05-13HISENSE VISUAL TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202380070272.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2022-12-19
Filing Date
2023-09-19
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

Smart devices only support a single audio track output when playing videos, which cannot meet the needs of multiple users to play multiple audio tracks according to their own needs, resulting in the inability to achieve synchronous playback of multiple audio tracks.

Method used

By parsing and processing multiple audio tracks in the multimedia file to be played, the target audio track is determined and the corresponding relationship with the playing headphones is established. The preset decoder is used for decoding, the pulse modulation code is generated, and the audio track is played according to the relationship between the audio track and the headphones. .

Benefits of technology

It realizes the synchronous playback of multiple audio tracks, meets the audio playback needs of different users, and improves the user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119999221A_ABST
    Figure CN119999221A_ABST
Patent Text Reader

Abstract

The invention discloses an audio processing method and an electronic device, and the method comprises the steps: carrying out the analysis of a to-be-played multimedia file under the condition that the to-be-played multimedia file comprises a plurality of audio tracks, and obtaining the plurality of audio tracks of the to-be-played multimedia file; determining at least two paths of target audio tracks in the multiple paths of audio tracks, and establishing a corresponding relationship between each path of target audio track and each playing earphone; based on a preset decoder corresponding to each path of target audio track, decoding each path of target audio track to obtain a pulse modulation code corresponding to each path of target audio track; and based on the pulse modulation code corresponding to each path of target audio track and the corresponding relationship between each path of target audio track and each playing earphone, playing the target audio track through the playing earphone corresponding to each path of target audio track, so that playing of multiple paths of audio tracks can be realized, the audio playing requirements of different users on the to-be-played multimedia file are met, and the user experience is improved. Therefore, each user can hear the required audio, and the user experience is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Audio processing method and electronic device

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS

[0002] This application claims priority to the Chinese patent application filed on December 14, 2022, with application number 202211611352.1; and filed on December 19, 2022, with application number 202211632453.7, the entire contents of which are incorporated by reference into this application. Technical Field

[0003] The present application relates to the field of audio processing technology, and in particular to an audio processing method and electronic device. Background Art

[0004] Currently, when users play videos through smart devices, since the smart devices only support single-channel audio track output, the multiple audio tracks contained in the played video are only played according to the default audio track of the smart device system, or the single audio track that a user determines to be played among the multiple audio tracks, while other audio tracks are prohibited from being played. As a result, in scenarios with multiple users, it is impossible to play the audio track corresponding to each user's needs. That is, the related art cannot meet the needs of each user and play multiple audio tracks.

[0005] Summary of the Invention

[0006] An embodiment of the present application provides an audio processing method, comprising: upon determining that a multimedia file to be played contains multiple audio tracks, parsing the multimedia file to be played to obtain the multiple audio tracks of the multimedia file to be played; determining at least two target audio tracks from the multiple audio tracks, and establishing a correspondence between each target audio track and each playback headphone; decoding each target audio track based on a preset decoder corresponding to each target audio track to obtain a pulse modulation code corresponding to each target audio track; and playing the target audio track through the playback headphone corresponding to each target audio track based on the pulse modulation code corresponding to each target audio track and the correspondence between each target audio track and each playback headphone.

[0007] An embodiment of the present application provides an electronic device, comprising: a memory configured to store computer instructions;

[0008] A processor, connected to the memory, is configured to execute when executing the computer instructions:

[0009] If it is determined that the multimedia file to be played contains multiple audio tracks, parsing the multimedia file to be played to obtain the multiple audio tracks of the multimedia file to be played;

[0010] Determining at least two target audio tracks from the multiple audio tracks, and establishing a correspondence between each target audio track and each playback headphone;

[0011] Decoding each target audio track based on a preset decoder corresponding to each target audio track to obtain a pulse modulation code corresponding to each target audio track;

[0012] Based on the pulse modulation code corresponding to each target audio track and the corresponding relationship between each target audio track and each playback earphone, the target audio track is played through the playback earphone corresponding to each target audio track. BRIEF DESCRIPTION OF THE DRAWINGS

[0013] FIG1A is a schematic diagram of a process of playing through a single audio track according to some embodiments;

[0014] FIG1B is a schematic diagram of another process of playing through a single audio track according to some embodiments;

[0015] FIG2 is a schematic diagram of software configuration of an electronic device according to some embodiments;

[0016] FIG3A is a schematic flow chart of an audio processing method according to some embodiments;

[0017] FIG3B is a schematic diagram illustrating determining a target audio track among multiple audio tracks according to some embodiments;

[0018] FIG3C is a schematic diagram of a process of playing multiple audio tracks according to some embodiments;

[0019] FIG4A is a schematic flow chart of another audio processing method according to some embodiments;

[0020] FIG4B is a flowchart of yet another audio processing method according to some embodiments;

[0021] FIG5A is a schematic flow chart of another audio processing method according to some embodiments;

[0022] FIG5B is a schematic diagram of another process of playing multiple audio tracks according to some embodiments;

[0023] FIG6A is a schematic flow chart of another audio processing method according to some embodiments;

[0024] FIG6B is a schematic diagram of another process of playing multiple audio tracks according to some embodiments;

[0025] FIG7 is a schematic flow chart of yet another audio processing method according to some embodiments;

[0026] FIG8 is a schematic flow chart of yet another audio processing method according to some embodiments;

[0027] FIG9 is a flowchart of another audio processing method according to some embodiments;

[0028] FIG10 is a scene architecture diagram of a subtitle display method according to some embodiments;

[0029] FIG11 is a block diagram of a hardware configuration of a control device according to some embodiments;

[0030] FIG12 is a block diagram of a hardware configuration of a display device according to some embodiments;

[0031] FIG13 is a diagram illustrating software configuration in a display device according to some embodiments;

[0032] FIG14 is a flowchart of steps of a subtitle display method according to some embodiments;

[0033] FIG15 is a schematic diagram of a scenario of a subtitle display method according to some embodiments;

[0034] FIG16 is a flowchart of steps of a subtitle display method according to some embodiments;

[0035] FIG17 is a schematic diagram of a subtitle display method according to some embodiments;

[0036] FIG18 is a flowchart of another method for displaying subtitles according to some embodiments;

[0037] FIG19 is a schematic diagram of another subtitle display method according to some embodiments;

[0038] FIG20 is a flowchart of another method for displaying subtitles according to some embodiments;

[0039] FIG21 is a schematic diagram of another subtitle display method according to some embodiments;

[0040] FIG22 is a flowchart illustrating the steps of another subtitle display method according to some embodiments;

[0041] FIG23 is a schematic diagram of yet another subtitle display method according to some embodiments;

[0042] FIG24 is a flowchart illustrating the steps of another subtitle display method according to some embodiments;

[0043] FIG25 is a schematic diagram of yet another subtitle display method according to some embodiments;

[0044] FIG26 is a flowchart illustrating the steps of another subtitle display method according to some embodiments;

[0045] FIG27 is a flowchart illustrating the steps of another subtitle display method according to some embodiments;

[0046] FIG28 is a schematic structural diagram of an electronic device according to some embodiments;

[0047] FIG29 is a schematic structural diagram of a smart device according to some embodiments. DETAILED DESCRIPTION

[0048] The following embodiments are described in detail, with examples shown in the accompanying drawings. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. It should be noted that the brief descriptions of terms in this application are only for the convenience of understanding the embodiments described below and are not intended to limit the embodiments of this application. Unless otherwise indicated, these terms should be understood according to their ordinary and usual meanings.

[0049] FIG1A is a schematic diagram of a process for playing a video using a single audio track according to some embodiments. As shown in FIG1A , when a user needs to use a smart device such as a smart TV to play a video, they input a play command. The smart TV first downloads the playback file to be played from a local device or a server. After downloading the playback file, if the playback file is determined to be a streaming media playback file, the protocol decapsulation module parses the streaming media playback file to obtain the address corresponding to the media segment contained in the streaming media playback file. The media segment is downloaded based on the address. After obtaining the media segment contained in the media playback file, the format decapsulation module parses the file to extract the different audio elementary streams, video elementary streams, and subtitle elementary streams contained in the playback file, and caches them to ensure smooth video playback. It should be noted that if the playback file contains multiple audio tracks, the multiple audio tracks will be parsed. The user selects the target audio track to be played from the multiple audio tracks through a selector. Simultaneously, the subtitle selector determines the corresponding subtitles to be played based on the target audio track selected by the user, decodes the target audio track, video, and subtitles, and plays them synchronously.

[0050] Figure 1B is a schematic diagram of another process of playing through a single audio track according to some embodiments. As shown in Figure 1B, when it is determined that the playback file is not a streaming media playback file, there is no need to parse the streaming media playback file through the protocol decapsulation module to obtain the address corresponding to the media segment contained in the streaming media playback file, and download the media segment based on the address.

[0051] However, using the above method, when a user plays a video through a smart device, the smart device only supports single-channel audio track output. For multiple audio tracks contained in the played video, a single audio track is played according to the default audio track of the smart device system, or a single audio track determined by a user to be played among the multiple audio tracks. Multiple audio tracks cannot be played.

[0052] To solve the above problems, the present invention provides an audio processing method. When it is determined that a multimedia file to be played contains multiple audio tracks, the multimedia file to be played is parsed to obtain the multiple audio tracks of the multimedia file to be played; at least two target audio tracks are determined from the multiple audio tracks, and a corresponding relationship between each target audio track and each playback headphone is established; each target audio track is decoded based on a preset decoder corresponding to each target audio track to obtain a pulse modulation code corresponding to each target audio track; and based on the pulse modulation code corresponding to each target audio track and the corresponding relationship between each target audio track and each playback headphone, the target audio track is played through the playback headphone corresponding to each target audio track. In the above process, the multiple audio tracks contained in the multimedia file to be played can be determined according to different user needs, and a corresponding relationship between each target audio track and each playback headphone is established, so that different playback headphones can play the multiple target audio tracks contained in the same multimedia file to be played at the same time, thereby achieving the playback of multiple audio tracks, meeting the audio playback needs of different users for the multimedia file to be played, allowing each user to listen to the audio they need, and improving the user experience.

[0053] The audio processing model training method and audio processing method provided by the embodiments of the present disclosure can be implemented based on an electronic device, or a functional module or functional entity in an electronic device.

[0054] The electronic device may be a smart TV, a personal computer (PC), a server, a mobile phone, a tablet computer, a laptop computer, a mainframe computer, etc., and the embodiments of the present disclosure do not specifically limit this.

[0055] Figure 2 is a schematic diagram of the software configuration of an electronic device according to some embodiments. As shown in Figure 2, the system is divided into four layers, from top to bottom, namely the application layer (referred to as the "application layer"), the application framework layer (referred to as the "framework layer"), the Android runtime (Android runtime) and system library layer (referred to as the "system runtime library layer"), and the kernel layer.

[0056] In some embodiments, at least one application runs in the application layer. These applications can be window programs, system settings programs, clock programs, etc. that come with the operating system, or applications developed by third-party developers. In specific implementations, the application packages in the application layer are not limited to the above examples.

[0057] The framework layer provides applications with an application programming interface (API) and programming framework. The application framework layer includes predefined functions. The application framework layer acts as a processing center, determining the actions taken by applications in the application layer. Applications use the API to access system resources and services during execution.

[0058] In some embodiments, the system runtime layer provides support for the upper layer, namely the framework layer. When the framework layer is used, the Android operating system will run the C / C++ library contained in the system runtime layer to implement the functions to be implemented by the framework layer.

[0059] In some embodiments, the kernel layer is a layer between hardware and software. The kernel layer includes at least one of the following drivers: audio driver, display driver, Bluetooth driver, camera driver, Wi-Fi driver, USB driver, HDMI driver, sensor driver (such as fingerprint sensor, temperature sensor, pressure sensor, etc.), and power driver.

[0060] The audio processing method provided in the embodiment of the present application can be implemented based on the above-mentioned electronic device.

[0061] In order to illustrate this solution in more detail, the following will be explained in an exemplary manner in conjunction with Figure 3A. It can be understood that the steps involved in Figure 3A may include more steps or fewer steps in actual implementation, and the order of these steps may also be different, so as to be able to implement the audio processing method provided in the embodiment of the present application.

[0062] FIG3A is a flow chart of an audio processing method according to some embodiments. The method of this embodiment is performed by an audio processing device (electronic device) applied to a smart device, which can be implemented in hardware and / or software. As shown in FIG3A , the audio processing method specifically includes the following steps:

[0063] S31 : When it is determined that the multimedia file to be played contains multiple audio tracks, the multimedia file to be played is parsed to obtain the multiple audio tracks of the multimedia file to be played.

[0064] Among them, the audio track refers to the attribute information corresponding to the audio contained in the multimedia file to be played. Each audio to be played contained in the multimedia file to be played corresponds to an audio track, such as language, timbre, timbre library, number of channels, input / output port, volume, but is not limited to this. This disclosure does not specifically limit it, and those skilled in the art can set it according to actual circumstances.

[0065] Specifically, when it is determined that the multimedia file to be played contains multiple audio tracks, the multimedia file to be played is parsed to obtain the multiple audio tracks contained in the multimedia file to be played.

[0066] For example, when a user uses a smart device such as a smart TV to play a movie XXX, it is determined based on the file information of the movie XXX that the movie XXX contains five audio tracks. That is, it can be understood that the movie file XXX currently contains five languages ​​for the user to select. After determining that the movie XXX contains five audio tracks, the movie XXX is parsed to obtain the corresponding five audio tracks, but the present disclosure does not specifically limit this. Those skilled in the art can make settings according to actual conditions.

[0067] The above-mentioned parsing and processing of the multimedia file to be played refers to related technologies and will not be repeated here.

[0068] S32: Determine at least two target audio tracks from the multiple audio tracks, and establish a correspondence between each target audio track and each playback headphone.

[0069] Among them, the playback earphones are used to play audio, that is, to play the target audio corresponding to the target audio track. One playback earphone corresponds to one target audio track, so as to avoid interference between multiple users using different audio tracks when watching the same multimedia file to be played, such as watching XXX movies in different languages. The playback earphones can be, for example, Bluetooth earphones, but are not limited to this. This disclosure does not specifically limit it, and those skilled in the art can set it according to actual conditions.

[0070] Specifically, after parsing the multimedia file to be played and obtaining multiple audio tracks contained in the multimedia file to be played, two or more target audio tracks to be played are determined from the multiple audio tracks. Since each target audio track requires a playback headphone to play, it is necessary to establish a correspondence between each target audio track and each playback headphone.

[0071] Optionally, based on the above embodiments, in some embodiments of the present disclosure, determining at least two target audio tracks from multiple audio tracks may be performed based on a selection instruction input by a user.

[0072] For example, referring to FIG3B , FIG3B is a schematic diagram of determining a target audio track among multiple audio tracks according to some embodiments. According to the needs of different users, each user inputs a selection instruction on a display interface 301 of a smart device such as a smart TV to determine Track 2 and Track 3 among the multiple audio tracks as the target audio tracks that different users need to play.

[0073] Optionally, based on the above embodiments, in some embodiments of the present disclosure, establishing the correspondence between each target audio track and each playback headphone can be done by determining the order in which the playback headphones are connected to the smart device and the order in which multiple target audio tracks are determined, thereby determining the correspondence between each target audio track and each playback headphone, or it can be configured through user customization.

[0074] S33 , decoding each target audio track based on a preset decoder corresponding to each target audio track to obtain a pulse modulation code corresponding to each target audio track.

[0075] Among them, pulse modulation coding refers to the digital sampling processing of analog signals, that is, the coding method of converting audio signals into digital signals, which mainly goes through sampling, quantization and encoding. Specifically, the sampling process converts the continuous-time audio signal into a discrete-time, continuous-amplitude sampling signal, the quantization process converts the sampling signal into a discrete-time, discrete-amplitude digital signal, and the encoding process encodes the quantized digital signal into a binary code group output. After obtaining the pulse modulation code corresponding to each target audio track, the pulse modulation code can be used for rendering and playback.

[0076] Specifically, for each target audio track, decoding processing is performed on each target audio track according to a preset decoder corresponding to each target audio track, so as to obtain the pulse modulation code corresponding to each target audio track.

[0077] It should be noted that the above is a decoding process for the basic code stream corresponding to each target audio track. The specific decoding process refers to the relevant technology and will not be repeated here.

[0078] S34, based on the pulse modulation code corresponding to each target audio track and the correspondence between each target audio track and each playback earphone, play the target audio track through the playback earphone corresponding to each target audio track.

[0079] Specifically, after obtaining the pulse modulation code corresponding to each target audio track, according to the correspondence between each target audio track and each playback earphone, each target audio track is played through the playback earphone corresponding to each target audio track and using the pulse modulation code corresponding to each target audio track.

[0080] Optionally, FIG3C is a schematic diagram of a process of playing multiple audio tracks according to some embodiments. The specific implementation process refers to the above steps S31-S34 and will not be described in detail here.

[0081] In this way, the audio processing method provided in the embodiment of the present disclosure is as follows: when it is determined that the multimedia file to be played contains multiple audio tracks, the multimedia file to be played is parsed to obtain the multiple audio tracks of the multimedia file to be played; at least two target audio tracks are determined in the multiple audio tracks, and a corresponding relationship between each target audio track and each playback headphone is established; after decoding each target audio track based on a preset decoder corresponding to each target audio track, a pulse modulation code corresponding to each target audio track is obtained; based on the pulse modulation code corresponding to each target audio track and the corresponding relationship between each target audio track and each playback headphone, the target audio track is played through the playback headphone corresponding to each target audio track. In the above process, the multiple audio tracks contained in the multimedia file to be played can be determined according to different user needs, and a corresponding relationship between each target audio track and each playback headphone is established, so that different playback headphones can play the multiple audio tracks contained in the same multimedia file to be played at the same time, thereby achieving the playback of multiple audio tracks, meeting the audio playback needs of different users for the multimedia file to be played, allowing each user to listen to the audio they need, and improving the user experience.

[0082] FIG4A is a flow chart of another audio processing method according to some embodiments. This embodiment is a further expansion and optimization based on the above embodiment. Optionally, referring to FIG4A , before executing S33, the method further includes:

[0083] S41, obtaining parameter information corresponding to each target audio track.

[0084] The parameter information includes, but is not limited to, the audio sampling rate, the number of channels, and the bit rate. This disclosure does not specifically limit this, and those skilled in the art may set it according to actual conditions.

[0085] S42: Based on the parameter information corresponding to each target audio track, a preset decoder corresponding to each target audio track is established.

[0086] Among them, the preset decoders include hard decoders and soft decoders. The hard decoder is a decoder established based on an independent hardware chip. The hard decoder can improve the decoding efficiency of the target audio track. The soft decoder is a decoder established according to the encoding.

[0087] Specifically, for each target audio track, parameter information corresponding to each target audio track is obtained, such as audio sampling rate, number of channels, and bit rate. According to the parameter information corresponding to each target audio track, a corresponding decoder is established for each target audio track.

[0088] Optionally, based on the above embodiment, FIG4B is a flowchart of another audio processing method according to some embodiments. This embodiment is a further expansion and optimization based on the above embodiment. Optionally, referring to FIG4B , one implementation of S42 may be:

[0089] S421 , performing a product operation on the audio sampling rate, the number of channels, and the bit rate of each target audio track to obtain a product operation result of each target audio track.

[0090] Specifically, after obtaining the audio sampling rate, number of channels and bit rate of each target audio track, a product operation is performed on the audio sampling rate, number of channels and bit rate of each target audio track to calculate the product operation result of each target audio track.

[0091] It should be noted that by calculating the product operation results of each target audio track, the target audio track corresponding to the maximum product operation result can be determined, which is the optimal target audio track among the multiple target audio tracks, that is, the target audio track with the best playback quality.

[0092] S422: Establish a hardware decoder for the target audio track corresponding to the maximum product operation result, and establish a soft decoder for other target audio tracks.

[0093] Specifically, after calculating the product operation result of each target audio track, for the target audio track with the maximum product operation result, since this target audio track is the optimal target audio track among multiple target audio tracks, that is, the target audio track with the best playback quality, and since more resources are required when decoding it, a hard decoder is established for the optimal target audio track, and resources in the decoding process are provided by the hardware chip, so as to improve the efficiency of decoding the target audio track corresponding to the maximum product operation result. For other target audio tracks, a soft decoder is established.

[0094] Exemplarily, following the above embodiment, for the five audio tracks contained in the multimedia file to be played: Track 1, Track 2, Track 3, Track 4 and Track 5, it is determined that Track 2 and Track 3 are two target audio tracks, namely, target Track 1 and target Track 2, and the audio sampling rate, number of channels and bit rate corresponding to the target Track 1 and the target Track 2 are obtained, and a multiplication operation is performed to obtain the product operation result 1 and the product operation result 2 corresponding to the target Track 1 and the target Track 2, respectively. It is determined that the product operation result 1 is greater than the product operation result 2, so a hard decoder is established for the target Track 1 and a soft decoder is established for the target Track 2, but the present disclosure is not specifically limited thereto, and those skilled in the art can set it according to actual conditions.

[0095] In this way, the audio processing method provided in the embodiment of the present invention, in the above process, performs a product operation according to the audio sampling rate, number of channels and bit rate corresponding to each target audio track, and establishes a hard decoder for the target audio track corresponding to the maximum product operation result based on the multiplication result, and establishes a soft decoder for the other target audio tracks. In this way, independent hardware chips can be used to provide resources during the decoding process of the target audio track with the best sound quality, thereby improving the efficiency of decoding the best target audio track, and saving resources of the smart device, and also ensuring the efficiency of decoding other target audio tracks to a certain extent.

[0096] Optionally, FIG5A is a flow chart of another audio processing method according to some embodiments. FIG5B is a flow chart of another process of playing multiple audio tracks according to some embodiments. This embodiment is a further expansion and optimization of the above embodiment. Referring to FIG5A , when executing S34, it also includes:

[0097] S51 , when playing the target audio track through the playback earphone corresponding to each target audio track, synchronously playing the video and subtitles included in the multimedia file to be played based on the target synchronization clock.

[0098] Among them, the target synchronization clock is used to ensure that multiple target audio tracks, subtitles and videos can be played synchronously.

[0099] Specifically, when the target audio track is played through the playback earphone corresponding to each target audio track, the video and subtitles included in the to-be-played multimedia file are played accordingly according to the target synchronization clock.

[0100] Thus, the audio processing method provided in the embodiment of the present disclosure utilizes the target synchronization clock in the above process to ensure that the video, subtitles, and multiple target audio tracks contained in the multimedia file to be played can be played synchronously.

[0101] Optionally, FIG6A is a flow chart of another audio processing method according to some embodiments. FIG6B is a flow chart of another process of playing multiple audio tracks according to some embodiments. This embodiment is a further expansion and optimization of the above embodiment. Referring to FIG6A , one implementation of S51 may be:

[0102] S61 , decoding the elementary code streams corresponding to the video and subtitles contained in the multimedia file to be played, and obtaining initial data corresponding to the video and subtitles.

[0103] The initial data refers to the uncompressed original data corresponding to the video and subtitles respectively.

[0104] Specifically, for the video and subtitles contained in the multimedia file to be played, the basic code stream of the video is decoded according to the decoder corresponding to the video, and the basic code stream of the subtitles is decoded according to the decoder corresponding to the subtitles, so as to obtain the original data before compression corresponding to the video and subtitles respectively. The specific process of basic code stream decoding processing can be referred to related technologies and will not be repeated here.

[0105] S62 : Synchronously play the video and subtitles included in the multimedia file to be played based on the target synchronization clock and the initial data corresponding to the video and subtitles respectively.

[0106] Specifically, after obtaining the initial data corresponding to the video and subtitles, the initial data corresponding to the video and subtitles are used for rendering according to the target synchronization clock, so as to play the video and subtitles contained in the multimedia file to be played synchronously with multiple target audio tracks.

[0107] Optionally, FIG7 is a flow chart of another audio processing method according to some embodiments. This embodiment is a further expansion and optimization based on the above embodiment. Referring to FIG7 , before executing S61, it further includes:

[0108] S71: Determine the audio clock corresponding to each target audio track.

[0109] S72: Determine a target synchronization clock among multiple audio clocks.

[0110] The target synchronization clock is used to synchronously play the video, subtitles and at least two target audio tracks contained in the multimedia file to be played.

[0111] Specifically, for multiple target audio tracks determined in the multiple audio tracks, an audio clock corresponding to each target audio track is determined, and an audio clock is selected from the multiple audio clocks as the target audio track for synchronously playing the video, subtitles and at least two target audio tracks contained in the multimedia file to be played.

[0112] Optionally, based on the above embodiment, in some embodiments of the present disclosure, the implementation of S72 includes but is not limited to the following two methods. Optionally, FIG8 is a flowchart of another audio processing method according to some embodiments. This embodiment further expands and optimizes the above embodiment. Referring to FIG8, one implementation of S72 may be:

[0113] S81: Based on the product operation results of each target audio track, determine the first target audio track corresponding to the maximum product operation result.

[0114] S82: Use the audio clock of the first target audio track as the target synchronization clock.

[0115] Specifically, since the product operation is performed on the parameter information corresponding to each target audio track, namely the audio sampling rate, the number of channels and the bit rate, the product operation result corresponding to each target audio track is obtained, it can be determined that the first target audio track corresponding to the maximum product operation result is the playback audio track with the best sound quality among the multiple target audio tracks. Therefore, according to the product operation result of each target audio track, after determining the first target audio track corresponding to the maximum product operation result, the audio clock of the first target audio track is used as the target synchronization clock, so that other target audio tracks and the videos and subtitles included in the multimedia file to be played can be played synchronously.

[0116] In this way, the audio processing method provided in the embodiment of the present disclosure, in the above process, determines the playback audio track with the best sound quality among multiple target audio tracks based on the parameter information corresponding to each target audio track, and uses the audio clock corresponding to the target audio track with the best sound quality as the target synchronization clock. In this way, during the playback of the multimedia file to be played, the smoothness of the rendering and playback of the video, subtitles and multiple target audio tracks is guaranteed, making the playback smoother.

[0117] Optionally, based on the above embodiment, FIG9 is a flowchart of another audio processing method according to some embodiments. This embodiment is a further expansion and optimization based on the above embodiment. Referring to FIG9 , another implementation of S72 may be:

[0118] S91 , when establishing a corresponding preset decoder for each target audio track, determining the last second target audio track for which the preset decoder is established.

[0119] S92: Use the audio clock of the second target audio track as the target synchronization clock.

[0120] Specifically, for each target audio track, a corresponding preset decoder needs to be established. When establishing corresponding preset decoders for multiple target audio tracks, the second target audio track for which the preset decoder is established is determined, and the audio clock corresponding to the second target audio track is used as the target synchronization clock, so that other target audio tracks, videos, and subtitles included in the multimedia file to be played can be played synchronously.

[0121] In this way, the audio processing method provided in the embodiment of the present disclosure uses the audio clock corresponding to the second target audio track of the preset decoder that is finally established as the target synchronization clock in the above process, thereby ensuring the smoothness of rendering and playback of videos, subtitles and multiple target audio tracks during the playback of the multimedia file to be played, making the playback smoother.

[0122] Thus, the audio processing method provided in the embodiments of the present disclosure, in the above process, optionally, based on the above embodiment, in some embodiments of the present disclosure, further includes:

[0123] When a switching instruction input by the user is received, the target audio track currently being played is switched.

[0124] Specifically, for a multimedia file to be played, when different users play the target audio tracks corresponding to their respective needs, when there is a user who needs to switch the target audio track to be played, the smart device receives the switching instruction input by the user, and in response to the switching instruction input by the user, switches the target audio track being played by the user, and then uses the target audio track that the user currently needs to listen to to play it.

[0125] It should be noted that during the target audio track switching process, when the switched target audio track is any target audio track or multiple target audio tracks corresponding to the soft decoder, or the target audio track corresponding to the hard decoder, when the switching is completed, the audio clock corresponding to the hard decoder is still used as the target synchronization clock to play the video, subtitles and multiple target audios contained in the multimedia file to be played.

[0126] Optionally, when all the currently playing target audio tracks are switched, the target synchronization clock is further re-determined in the process of switching the target audio tracks. The specific implementation method of the target synchronization clock is determined with reference to the above embodiments S81-S82, or S91-S92. This disclosure does not specifically limit it, and those skilled in the art can set it according to actual conditions.

[0127] In this way, the audio processing method provided in the embodiment of the present disclosure can, during the above process, switch the target audio track in real time according to the user's demand for the playback audio track during the playback of the multimedia file to be played, thereby improving the user experience.

[0128] Among them, when playing the target audio track through the playback headphones corresponding to each target audio track, the process of playing the subtitles included in the multimedia file to be played can be performed in the manner provided in the following embodiments.

[0129] FIG10 is a schematic diagram of a scenario architecture of a method for controlling a display device according to some embodiments. As shown in FIG10 , the scenario architecture provided by the embodiment of the present application includes: a control device 100 , a display device 200 , and a server 300 .

[0130] The display device provided in the embodiments of the present application can have various implementation forms. For example, the display device can be a television, a smart speaker refrigerator with a display function, curtains with a display function, a personal computer (PC), a laser projection device, a monitor, an electronic whiteboard, a wearable device, a vehicle-mounted device, an electronic table, etc.

[0131] In some embodiments, the control device 100 may be a remote control. Communication between the remote control and the display device may include infrared protocol communication, Bluetooth protocol communication, or other short-range communication methods, allowing the display device to be controlled wirelessly or wired. A user may control the display device by inputting user commands through buttons on the remote control, voice input, or control panel input.

[0132] In some embodiments, the control device 100 may be a terminal device, for example, a mobile terminal such as a mobile phone, a tablet computer, a computer, or a laptop computer.

[0133] In some embodiments, the display device may also be controlled in a manner other than the control device. For example, the user's selection operation may be directly received through a user selection receiving module configured within the display device.

[0134] In some embodiments, the display device may also communicate data with the server 300 to obtain relevant media resources from the server 300. The display device may be allowed to communicate with the server via a local area network (LAN) or a wireless local area network (WLAN). The server 300 may provide media resource services and various content and interactions to the display device. The server 300 may be a single cluster or multiple clusters, and may include one or more types of servers.

[0135] Figure 11 is a hardware configuration block diagram of a control device according to some embodiments. Figure 11 exemplifies a configuration block diagram of the control device 100 in the embodiment shown in Figure 10. As shown in Figure 11, the control device 100 includes a processor 110, a communication interface 130, a user interface 140, a memory, and a power supply. The control device 100 can receive user input of operational instructions, convert the operational instructions into instructions that the display device can recognize and respond to, and forward the operational instructions or instructions converted from voice instructions to the display device, acting as an intermediary for interaction between the user and the display device.

[0136] In some embodiments, the user interface 140 of the control device 100 is configured to perform the following steps: receiving a user's selection operation.

[0137] In some embodiments, the user interface 140 of the control device 100 is further configured to perform the following steps: receiving a deletion operation from the user, wherein the deletion operation is used to stop the synchronous display of the first subtitle in the subtitles to be output.

[0138] In some embodiments, the user interface 140 of the control device 100 is further configured to perform the following steps: receiving an adding operation from the user, wherein the adding operation is used to add a second subtitle other than the subtitles to be output for synchronous display.

[0139] Figure 12 is a hardware configuration block diagram of a display device according to some embodiments. As shown in Figure 12, the display device includes at least one of a tuner and demodulator 210, a communicator 220, a detector 230, an external device interface 240, a processor 250, a display 260, an audio output interface 270, a memory, a power supply, and a user interface 280.

[0140] In some embodiments, the processor 250 includes one or more processors, for example, a video processor, an audio processor, a graphics processor, RAM, ROM, and first to nth interfaces for input / output.

[0141] The display 260 includes a display screen component for presenting images, a driving component for driving image display, a component for receiving image signals output from a processor, and a component for displaying video content, image content, and a menu control interface and a user control UI interface.

[0142] The display 260 may be a liquid crystal display, an OLED display, or a projection display, and may also be a projection device and a projection screen.

[0143] Communicator 220 is a component used to communicate with external devices or servers using various communication protocols. For example, the communicator may include at least one of a Wi-Fi module, a Bluetooth module, a wired Ethernet module, or other network communication protocol chip, a near-field communication protocol chip, and an infrared receiver. The display device can use communicator 220 to send and receive control signals and data signals with the external control device 100 or server 300.

[0144] The user interface may be used to receive control signals input by the user through the control device 100 (eg, an infrared remote controller, etc.) or through touch or gestures.

[0145] Detector 230 is used to collect signals from the external environment or external interactions. For example, detector 230 includes a light receiver, a sensor for collecting ambient light intensity; or detector 230 includes an image collector, such as a camera, for collecting external environmental scenes, user attributes, or user interaction gestures; or detector 230 includes a sound collector, such as a microphone, for receiving external sounds.

[0146] The external device interface 240 may include, but is not limited to, any one or more of the following: a high-definition multimedia interface (HDMI), an analog or digital high-definition component input interface (component), a composite video input interface (CVBS), a USB input interface (USB), an RGB port, etc. It may also be a composite input / output interface formed by multiple of the above interfaces.

[0147] The tuner-demodulator 210 receives broadcast television signals via a wired or wireless reception method, and demodulates audio and video signals, such as EPG data signals, from a plurality of wireless or wired broadcast television signals.

[0148] In some embodiments, the processor 250 and the tuner / demodulator 210 may be located in different separate devices, that is, the tuner / demodulator 210 may also be located in an external device of the main device where the processor 250 is located, such as an external set-top box.

[0149] Processor 250 controls the operation of the display device and responds to user operations through various software control programs stored in memory. Processor 250 controls the overall operation of the display device. For example, in response to receiving a user command to select a UI object for display on display 260, processor 250 may perform operations related to the object selected by the user command.

[0150] In some embodiments, the processor includes at least one of a central processing unit (CPU), a video processor, an audio processor, a graphics processing unit (GPU), RAM Random Access Memory (RAM), ROM (Read-Only Memory, ROM), a first interface to an nth interface for input / output, a communication bus (Bus), etc.

[0151] The user may input a user command through a graphical user interface (GUI) displayed on the display 260, and the user interface receives the user input command through the graphical user interface (GUI). Alternatively, the user may input a user command through a specific voice or gesture, and the user interface may recognize the voice or gesture through a sensor to receive the user input command.

[0152] A user interface is the medium for interaction and information exchange between an application or operating system and the user. It converts information between its internal form and a user-friendly format. A common user interface is the graphical user interface (GUI), which refers to a graphical user interface related to computer operations. It can be an icon, window, control, or other interface element displayed on an electronic device's display. Controls can include visual interface elements such as icons, buttons, menus, tabs, text boxes, dialog boxes, status bars, navigation bars, and widgets.

[0153] In some embodiments, the user interface 280 in the display device is configured to perform the following steps: receiving a user selection operation; the processor 250 is configured to, in response to the selection operation, determine one or more subtitles to be output from the subtitles of the multimedia file to be played; and obtain subtitle data for each of the subtitles to be output;

[0154] Obtaining a global clock of a playback pipeline for the multimedia file to be played and synchronous rendering logic corresponding to each channel of subtitles to be output; and synchronously rendering subtitle data of each channel of subtitles to be output based on the global clock and the synchronous rendering logic corresponding to each channel of subtitles to be output, so as to synchronously display each channel of subtitles to be output when playing the multimedia file.

[0155] In some embodiments, the processor 250 may obtain the subtitle data of each channel of subtitles to be output by: obtaining an encapsulation file corresponding to the multimedia file to be played; decapsulating the encapsulation file to obtain the basic stream of the multimedia file to be played, where the basic stream of the multimedia file to be played includes the subtitle basic stream of at least one channel of subtitles; and decoding the subtitle basic stream of the subtitles to be output in the subtitle basic stream of the at least one channel of subtitles to obtain the subtitle data of the subtitles to be output.

[0156] In some embodiments, the user interface 140 of the control device 100 is further configured to perform the following steps: receiving a deletion operation from the user, wherein the deletion operation is used to stop the synchronous display of the first subtitle in the subtitles to be output.

[0157] In some embodiments, the user interface 140 of the control device 100 is further configured to perform the following steps: receiving an adding operation from the user, wherein the adding operation is used to add a second subtitle other than the subtitles to be output for synchronous display.

[0158] In some embodiments, in some embodiments, the processor 250 obtains the subtitle data of each channel of the subtitles to be output by: obtaining at least one external subtitle file of the multimedia file to be played based on the native layer of the display device; decapsulating the at least one external subtitle file respectively to obtain the subtitle basic stream corresponding to the at least one external subtitle file; decoding the subtitle basic stream of the subtitles to be output in the subtitle basic stream corresponding to the at least one external subtitle file to obtain the subtitle data of the subtitles to be output.

[0159] In some embodiments, the processor 250 may obtain the subtitle data of each channel of subtitles to be output by: obtaining at least one external subtitle file of the multimedia file to be played based on the application layer of the display device; and parsing the at least one external subtitle file respectively to obtain the subtitle data corresponding to each external subtitle file.

[0160] In some embodiments, the processor 250 synchronously renders the subtitle data of each channel of the subtitles to be output based on the global clock and the synchronous rendering logic corresponding to each channel of the subtitles to be output, so as to synchronously display each channel of the subtitles to be output when playing the multimedia file. The method can be: based on the global clock and the preset synchronous rendering logic, synchronously render the subtitle data of each channel of the subtitles to be output, so as to synchronously display each channel of the subtitles to be output when playing the multimedia file.

[0161] In some embodiments, the processor 250 is further configured to add identification information of the target subtitle in the rendering event corresponding to the target subtitle; wherein the target subtitle is any one of the one or more subtitles to be output, and the identification information is used to uniquely identify the target subtitle.

[0162] In some embodiments, the processor 250 may obtain the global clock of the playback pipeline of the multimedia file to be played by: obtaining the audio clock of the multimedia file to be played; and determining the audio clock as the global clock of the playback pipeline of the multimedia file to be played.

[0163] In some embodiments, the user interface 280 is further configured to: receive a user's deletion operation, wherein the deletion operation is used to stop the synchronous display of the first subtitle in the subtitles to be output; the processor 250 is further configured to: in response to the user's deletion operation, stop synchronous rendering of the first subtitle, so as to stop synchronously displaying the first subtitle when playing the multimedia file.

[0164] In some embodiments, the user interface 280 is further configured to: receive an addition operation from the user, wherein the addition operation is used to add a synchronous display of a second subtitle other than the subtitles to be output; the processor 250 is further configured to: obtain the subtitle data of the second subtitle and the synchronous rendering logic corresponding to the second subtitle, and synchronously render the subtitle data of the second subtitle according to the global clock and the synchronous rendering logic corresponding to the second subtitle to add the synchronous display of the second subtitle.

[0165] Figure 13 is a software configuration diagram in a display device according to some embodiments. Referring to Figures 2 and 13, in some embodiments, the operating system of the display device is divided into two parts, the application layer and the native layer. The application layer is the application layer (referred to as the "application layer"), and the native layer is, from top to bottom, the native application framework (Application Framework) layer (referred to as the "framework layer"), the Android runtime (Android runtime) and the system library layer (referred to as the "system runtime library layer"), and the kernel layer.

[0166] As shown in Figure 13, the application framework layer in the embodiment of the present application includes managers, content providers, etc., wherein the manager includes at least one of the following modules: an activity manager (Activity Manager) is used to interact with all activities running in the system; a location manager (Location Manager) is used to provide system services or applications with access to system location services; a package manager (Package Manager) is used to retrieve various information related to applications currently installed on the device; a notification manager (Notification Manager) is used to control the display and clearing of notification messages; a window manager (Window Manager) is used to manage icons, windows, toolbars, wallpapers and desktop components on the user interface.

[0167] In some embodiments, the activity manager is used to manage the lifecycle of each application and common navigation back functions, such as controlling the exit, opening, and backing of an application. The window manager is used to manage all window programs, such as obtaining the display screen size, determining whether there is a status bar, locking the screen, taking screenshots, and controlling display window changes (such as shrinking the display window, shaking the display, distorting the display, etc.).

[0168] FIG14 is a flow chart of a subtitle display method according to some embodiments. As shown in FIG14 , the subtitle display method provided by an embodiment of the present application includes the following steps:

[0169] S1401: Receive a user's selection operation.

[0170] In some embodiments, the object for receiving the user's selection operation can be any display device that can play audio and video, such as a television, a mobile phone, etc.; the implementation method of receiving the user's selection operation can be: when the display device is a television, the user selects subtitles through a remote control, and the television receives the selection signal of the remote control through a communication interface; when the display device is a mobile phone, the user selects subtitles to be output through the subtitle selection interface, and the communication interface in the mobile phone receives the selection signal.

[0171] S1402: In response to the selection operation, determine one or more subtitles to be output from the subtitles of the multimedia file to be played.

[0172] In some embodiments, the multimedia file to be played can be a multimedia file stored locally on the display device, or a multimedia file downloaded from the network by the display device, for example, a multimedia file provided by video playback software on the display device. Exemplarily, the multimedia file can include a video file, an audio file, and a subtitle file, wherein the audio file includes at least one audio channel, and the subtitle file includes at least one subtitle channel. Exemplarily, the output subtitles can be text subtitles or image subtitles.

[0173] The subtitle file included in the multimedia file is called an embedded subtitle file.

[0174] S1403: Obtain subtitle data of each channel of subtitles to be output.

[0175] In some embodiments, the acquiring of the subtitle data of each channel of the subtitles to be output may be performed by decapsulating and decoding the subtitle file of each channel of the subtitles to be output; or by parsing and acquiring the subtitle file of each channel of the subtitles to be output.

[0176] S1404: Obtain a global clock of a playback pipeline of the multimedia file to be played and synchronous rendering logic corresponding to each channel of subtitles to be output.

[0177] The method for obtaining the global clock of the playback pipeline of the multimedia file to be played includes the following steps 1 and 2:

[0178] Step 1: Obtain the audio clock of the multimedia file to be played.

[0179] In some embodiments, the audio clock of the multimedia file to be played is carried in an audio stream in the multimedia file, and the audio stream is decoded to obtain the audio clock.

[0180] Step 2: Determine the audio clock as the global clock of the playback pipeline of the multimedia file to be played.

[0181] In some embodiments, the global clock provides a monotonically increasing absolute time, ie, the current time of the audio and video playback of the multimedia file, and the display period of each channel of subtitles to be output can be determined by the global clock.

[0182] S1405 : Based on the global clock and the synchronous rendering logic corresponding to each channel of the subtitles to be output, synchronously render the subtitle data of each channel of the subtitles to be output, so as to synchronously display each channel of the subtitles to be output when playing the multimedia file.

[0183] In the above S1405, the synchronous rendering of each channel of subtitle data to be output further includes:

[0184] The identification information of the target subtitle is added to the rendering event corresponding to the target subtitle.

[0185] The target subtitle is any one of the one or more subtitles to be output, and the identification information is used to uniquely identify the target subtitle.

[0186] In some embodiments, it is necessary to determine the identification information corresponding to the target subtitles based on the current rendering event, and then perform synchronous rendering and broadcast. Exemplarily, when synchronously rendering the target subtitles, the global clock can be obtained through the player pipeline, and the runtime of rendering the target subtitles can be obtained. Synchronous processing can be performed by comparing the difference between the display time of rendering the target subtitles and the runtime. If the cumulative sum of the display time of the currently transmitted target subtitles and the frame display duration is within a certain threshold range of the current runtime, the target subtitles are rendered. If it arrives ahead of time, wait, otherwise the frame data is discarded.

[0187] For example, referring to FIG15 , FIG15 is a scene diagram of a subtitle display method according to some embodiments. When the user selects 2-way subtitles to be displayed on a display device, the scene includes: a display screen 1500 of the display device, a subtitle 1 display area 1501, and a subtitle 2 display area 1502. Specifically, the user can also set the area where the subtitles are displayed, as well as the font style, font size, and color of the subtitles.

[0188] As can be seen from the above embodiments, the embodiments of the present application can receive a user's selection operation through a user interface; the processor determines one or more subtitles to be output from the subtitles of the multimedia file to be played in response to the selection operation; then obtains the subtitle data of each subtitle to be output; then obtains the global clock of the playback pipeline of the multimedia file to be played and the synchronous rendering logic corresponding to each subtitle to be output; finally, based on the global clock and the synchronous rendering logic corresponding to each subtitle to be output, the subtitle data of each subtitle to be output is synchronously rendered, so that each subtitle to be output is synchronously displayed when playing the multimedia file. Compared with the related art, only one subtitle can be displayed on the display device. The present application can obtain the global clock of the multimedia file to be played, combined with the synchronous rendering logic of each subtitle to be output, so that at least one subtitle to be output can be played synchronously with the audio and video. Therefore, the embodiments of the present application can display at least one subtitle to be output on the display device, thereby improving the user experience.

[0189] FIG16 is a flow chart of a subtitle display method according to some embodiments. As shown in FIG16 , the subtitle display method provided by an embodiment of the present application includes the following steps:

[0190] S1601: Receive a user's selection operation.

[0191] S1602: In response to the selection operation, determine one or more subtitles to be output from the subtitles of the multimedia file to be played.

[0192] S1603: Obtain the encapsulation file corresponding to the multimedia file to be played.

[0193] In some embodiments, multimedia files are divided into streaming media files and non-streaming media files. For streaming media files, obtaining the encapsulated file corresponding to the multimedia file to be played is performed by parsing the streaming media file protocol to obtain the media segment address and downloading the encapsulated file. For non-streaming media files, the encapsulated file can be directly obtained.

[0194] S1604: Decapsulate the encapsulated file to obtain the elementary stream of the multimedia file to be played.

[0195] The elementary stream of the multimedia file to be played includes at least one elementary stream of subtitles.

[0196] In some embodiments, the elementary stream of the multimedia file to be played further includes: a video stream and at least one audio stream.

[0197] S1605: Decode the subtitle elementary stream of the subtitles to be output in the subtitle elementary stream of the at least one subtitle channel to obtain subtitle data of the subtitles to be output.

[0198] In some embodiments, the subtitle data to be output includes: text content, display duration, and display timestamp of text subtitle data; or bitmap, display duration, and display timestamp of picture subtitle data.

[0199] S1606: Obtain a global clock of a playback pipeline of the multimedia file to be played and synchronous rendering logic corresponding to each channel of subtitles to be output.

[0200] In some embodiments,

[0201] S1607: Based on the global clock and the synchronous rendering logic corresponding to each channel of the subtitles to be output, synchronously render the subtitle data of each channel of the subtitles to be output, so as to synchronously display each channel of the subtitles to be output when playing the multimedia file.

[0202] For example, based on the embodiment described in FIG. 16 , referring to FIG. 17 , FIG. 17 is a schematic diagram of a subtitle display method according to some embodiments. Specifically, FIG. 17 is a schematic diagram of a scenario for implementing a multi-subtitle display method based on the player architecture shown in FIG. 10 . The multimedia files in this scenario include: video files, audio files, and subtitle files. This scenario includes:

[0203] The decapsulator 171 is configured to decapsulate a multimedia file to obtain a corresponding video stream, at least one audio stream, and at least one subtitle stream.

[0204] The buffer 172 is used to buffer the video stream, at least one audio stream and at least one subtitle stream obtained after decapsulation by the decapsulator 171 .

[0205] The buffer queue 173 is used to buffer the video stream, at least one audio stream, and at least one subtitle stream.

[0206] The audio / video selector 174 is configured to select one audio channel from at least one audio channel.

[0207] The multi-subtitle selector 175 is used to select the target subtitle according to the identification information of the target subtitle added in the rendering event.

[0208] Exemplarily, the target subtitles are subtitle 1 and subtitle 2.

[0209] The video decoder 176 is configured to decode the video elementary stream to obtain video data.

[0210] The audio 1 decoder 177 is used to decode the audio channel selected by the audio and video selector 174 to obtain audio data.

[0211] The subtitle 1 decoder 178 is used to decode the elementary stream of subtitle 1 to obtain subtitle 1 data.

[0212] The subtitle 2 decoder 179 is used to decode the subtitle 2 elementary stream to obtain subtitle 1 data.

[0213] The rendering module 1710 is used to synchronously render the video data, audio data, subtitle 1 data, and subtitle 2 data.

[0214] The display module 1711 is used to play audio, display video, subtitle 1, and subtitle 2 on a display device.

[0215] FIG18 is a flowchart of another subtitle display method according to some embodiments. As shown in FIG18 , the subtitle display method provided by an embodiment of the present application includes the following steps:

[0216] S1801: Receive a user's selection operation.

[0217] S1802: In response to the selection operation, determine one or more subtitles to be output from the subtitles of the multimedia file to be played.

[0218] S1803: Obtain at least one external subtitle file of the multimedia file to be played based on the native layer of the display device.

[0219] In some embodiments, the native layer of the display device is a development environment based on the C++ language.

[0220] S1804: Decapsulate the at least one external subtitle file respectively to obtain a subtitle elementary stream corresponding to the at least one external subtitle file.

[0221] S1805: Decode the subtitle elementary stream of the subtitles to be output in the subtitle elementary stream corresponding to the at least one external subtitle file to obtain subtitle data of the subtitles to be output.

[0222] S1806: Obtain a global clock of a playback pipeline of the multimedia file to be played and synchronous rendering logic corresponding to each channel of subtitles to be output.

[0223] S1807: Based on the global clock and a preset synchronous rendering logic, synchronously render the subtitle data of each channel of the subtitles to be output, so as to synchronously display each channel of the subtitles to be output when playing the multimedia file.

[0224] For example, in combination with the embodiment described in FIG. 18 , referring to FIG. 19 , FIG. 19 is a schematic diagram of another subtitle display method according to some embodiments. Specifically, FIG. 19 is a schematic diagram of a scenario for implementing the multi-subtitle display method based on the player architecture shown in FIG. 10 . The streaming media files in this scenario include video files and audio files. This scenario includes:

[0225] The streaming media file decapsulator 1901 is used to decapsulate the streaming media file to obtain the corresponding video stream and at least one audio stream.

[0226] The subtitle file decapsulator 1902 is used to decapsulate the external subtitle file to obtain at least one corresponding subtitle stream.

[0227] The buffer 1903 is used to buffer the decapsulated video stream, at least one audio stream, and at least one subtitle stream.

[0228] The buffer queue 1904 is used to buffer the video stream, at least one audio stream, and at least one subtitle stream.

[0229] The audio and video selector 1905 is used to select one audio channel from at least one audio channel.

[0230] The multi-subtitle selector 1906 is used to select the target subtitle according to the identification information of the target subtitle added in the rendering event.

[0231] Exemplarily, the target subtitles are subtitle 1 and subtitle 2.

[0232] The video decoder 1907 is used to decode the video elementary stream to obtain video data.

[0233] The audio 1 decoder 1908 is used to decode the audio channel selected by the audio and video selector 84 to obtain audio data.

[0234] The subtitle 1 decoder 1909 is used to decode the basic stream of subtitle 1 to obtain subtitle 1 data.

[0235] The subtitle 2 decoder 1910 is used to decode the subtitle 2 elementary stream to obtain subtitle 1 data.

[0236] The rendering module 1911 is used to synchronously render the video data, audio data, subtitle 1 data and subtitle 2 data.

[0237] The display module 1912 is used to play audio, display video, subtitle 1, and subtitle 2 on a display device.

[0238] In the above embodiment, for each channel of subtitle files to be output, in the native layer of the display device, at least one external subtitle file of the multimedia file to be played in the native layer of the display device is obtained, the at least one external subtitle file is decapsulated, a subtitle elementary stream corresponding to the at least one external subtitle file is obtained, and the subtitle elementary stream of the subtitles to be output in the subtitle elementary stream corresponding to the at least one external subtitle file is decoded to obtain subtitle data of the subtitles to be output.

[0239] Then, based on the global clock and the preset synchronous rendering logic, the subtitle data of each channel of the subtitles to be output is synchronously rendered, so that each channel of the subtitles to be output is synchronously displayed when the multimedia file is played, so that at least one channel of subtitles is displayed on the display device, thereby improving the user experience.

[0240] FIG20 is a flowchart of another subtitle display method according to some embodiments. As shown in FIG20 , the subtitle display method provided by an embodiment of the present application includes the following steps:

[0241] S2001: Receive a user's selection operation.

[0242] S2002: In response to the selection operation, determine one or more subtitles to be output from the subtitles of the multimedia file to be played.

[0243] S2003: Acquire at least one external subtitle file of the multimedia file to be played based on the application layer of the display device.

[0244] In some embodiments, the application layer of the display device is a development environment based on the JAVE language.

[0245] S2004: Parse the at least one external subtitle file respectively to obtain subtitle data corresponding to each external subtitle file.

[0246] S2005: Obtain the global clock of the playback pipeline of the multimedia file to be played and the synchronous rendering logic corresponding to each channel of the subtitles to be output.

[0247] S2006: Based on the global clock and the synchronous rendering logic corresponding to each channel of the subtitles to be output, synchronously render the subtitle data of each channel of the subtitles to be output, so as to synchronously display each channel of the subtitles to be output when playing the multimedia file.

[0248] For example, in combination with the embodiment described in FIG. 20 , referring to FIG. 21 , FIG. 21 is a schematic diagram of another subtitle display method according to some embodiments. Specifically, FIG. 21 is a schematic diagram of a scenario for implementing a multi-subtitle display method based on the player architecture shown in FIG. 10 . In this scenario, the subtitles to be output are two-way subtitles 1 and subtitles 2. The subtitle files containing subtitles 1 and subtitles 2 are in the external subtitle file of the application layer. This scenario includes:

[0249] The method for obtaining video and audio can be referred to the description of the embodiment shown in Figure 17, which will not be repeated here.

[0250] The subtitle buffer queue 2108 is used to buffer at least one external subtitle file.

[0251] The parsing module 2109 is configured to parse at least one external subtitle file and obtain subtitle data corresponding to each external subtitle file.

[0252] The multi-subtitle synchronizer 2110 is used to obtain the global clock of the playback pipeline of the multimedia file to be played and the synchronous rendering logic corresponding to each channel of the subtitles to be output.

[0253] The subtitle rendering module 2111 is used to synchronously render the subtitle data of each channel of subtitles to be output based on the global clock and the synchronous rendering logic corresponding to each channel of subtitles to be output, so as to synchronously display each channel of subtitles to be output when playing the multimedia file.

[0254] The display module 2112 is configured to select subtitle 1 and subtitle 2 for rendering and display according to the identification information of the target subtitle added in the rendering event corresponding to the target subtitle.

[0255] In the above embodiment, for each channel of subtitles to be output, the subtitle file is stored in the application layer of the display device. At least one external subtitle file of the multimedia file to be played is obtained by acquiring the application layer of the display device. The at least one external subtitle file is parsed to obtain subtitle data corresponding to each external subtitle file. The global clock of the playback pipeline of the multimedia file to be played and the synchronous rendering logic corresponding to each channel of subtitles to be output are then obtained. The subtitle data of each channel of subtitles to be output is then synchronously rendered. Each channel of subtitles to be output is synchronously displayed when the multimedia file is played, thereby displaying at least one channel of subtitles on the display device and improving the user experience.

[0256] FIG22 is a flowchart of another subtitle display method according to some embodiments. As shown in FIG22 , the subtitle display method provided by an embodiment of the present application includes the following steps:

[0257] S2201: Receive a user's selection operation.

[0258] S2202: In response to the selection operation, determine one or more subtitles to be output from the subtitles of the multimedia file to be played.

[0259] S2203: Obtain the encapsulation file corresponding to the multimedia file to be played.

[0260] S2204: Decapsulate the encapsulated file to obtain elementary streams of the multimedia file to be played, where the elementary streams of the multimedia file to be played include at least one subtitle elementary stream.

[0261] S2205: Decode the subtitle elementary stream of the subtitles to be output in the subtitle elementary stream of the at least one subtitle channel to obtain subtitle data of the subtitles to be output.

[0262] S2206: Obtain a global clock of a playback pipeline of the multimedia file to be played and synchronous rendering logic corresponding to each channel of the subtitles to be output.

[0263] S2207: Based on the global clock and the synchronous rendering logic corresponding to each channel of the subtitles to be output, synchronously render the subtitle data of each channel of the subtitles to be output, so as to synchronously display each channel of the subtitles to be output when playing the multimedia file.

[0264] While executing the above steps S2203-S2207, the embodiment of the present application also executes the following steps S2208-S2212:

[0265] S2208. Obtain at least one external subtitle file of the multimedia file to be played based on the native layer of the display device.

[0266] S2209: Decapsulate the at least one external subtitle file respectively to obtain a subtitle elementary stream corresponding to the at least one external subtitle file.

[0267] S2210: Decode the subtitle elementary stream of the subtitles to be output in the subtitle elementary stream corresponding to the at least one external subtitle file to obtain subtitle data of the subtitles to be output.

[0268] S2211: Obtain a global clock of a playback pipeline of the multimedia file to be played and synchronous rendering logic corresponding to each channel of subtitles to be output.

[0269] S2212: Based on the global clock and the synchronous rendering logic corresponding to each channel of the subtitles to be output, synchronously render the subtitle data of each channel of the subtitles to be output, so as to synchronously display each channel of the subtitles to be output when playing the multimedia file.

[0270] Exemplarily, in combination with the embodiment described in Figure 22 above, referring to Figure 23, Figure 23 is a schematic diagram of another subtitle display method according to some embodiments. Specifically, Figure 23 is a scenario schematic diagram of a multi-subtitle display method based on the player architecture shown in Figure 10. The subtitles to be output are two-way subtitles: subtitle 1 and subtitle a. The subtitle file containing subtitle 1 is in the multimedia file of the native layer, and the subtitle file containing subtitle a is in the external subtitle file of the native layer. The method for the native layer to obtain subtitle 1 can refer to the description of the embodiment scenario schematic diagram shown in Figure 17, and the method for the native layer to obtain subtitle a can be shown in Figure 19, which will not be repeated here.

[0271] FIG24 is a flowchart of another subtitle display method according to some embodiments. As shown in FIG24 , the subtitle display method provided by an embodiment of the present application includes the following steps:

[0272] S2401: Receive a user's selection operation.

[0273] S2402: In response to the selection operation, determine one or more subtitles to be output from the subtitles of the multimedia file to be played.

[0274] S2403: Obtain the encapsulation file corresponding to the multimedia file to be played.

[0275] S2404: Decapsulate the encapsulated file to obtain elementary streams of the multimedia file to be played, where the elementary streams of the multimedia file to be played include at least one subtitle elementary stream.

[0276] S2405: Decode the subtitle elementary stream of the subtitles to be output in the subtitle elementary stream of the at least one subtitle channel to obtain subtitle data of the subtitles to be output.

[0277] S2406: Obtain the global clock of the playback pipeline of the multimedia file to be played and the synchronous rendering logic corresponding to each channel of the subtitles to be output.

[0278] S2407: Based on the global clock and the synchronous rendering logic corresponding to each channel of the subtitles to be output, synchronously render the subtitle data of each channel of the subtitles to be output, so as to synchronously display each channel of the subtitles to be output when playing the multimedia file.

[0279] While executing the above steps S2403-S2407, the embodiment of the present application also executes the following steps S2408-S2410:

[0280] S2408: Acquire at least one external subtitle file of the multimedia file to be played based on the application layer of the display device.

[0281] S2409: Parse the at least one external subtitle file respectively to obtain subtitle data corresponding to each external subtitle file.

[0282] S2410: Based on the global clock and a preset synchronous rendering logic, synchronously render the subtitle data of each channel of the subtitles to be output, so as to synchronously display each channel of the subtitles to be output when playing the multimedia file.

[0283] Exemplarily, in combination with the embodiment described in Figure 24 above, refer to Figure 25, Figure 25 is a schematic diagram of another subtitle display method according to some embodiments. Specifically, Figure 25 is a scenario schematic diagram of a multi-subtitle display method based on the player architecture shown in Figure 10, and the subtitles to be output are two-way subtitles: subtitle 1 and subtitle a. The subtitle file containing subtitle 1 is in the multimedia file of the native layer, and the subtitle file containing subtitle a is in the external subtitle file of the application layer. The method for the native layer to obtain subtitle 1 can refer to the description of the embodiment scenario schematic diagram shown in Figure 17, and then the rendered video and subtitle 1 are transmitted to the application layer, and then displayed synchronously with subtitle a on the display device. The method for the application layer to obtain subtitle a can refer to the description of the embodiment scenario schematic diagram shown in Figure 21, which will not be repeated here.

[0284] As an extension and refinement of the above embodiment, as shown in FIG26 , the subtitle display method provided in the embodiment of the present application further includes the following S2601 and S2602:

[0285] S2601: Receive a deletion operation from the user.

[0286] The deletion operation is used to stop the synchronous display of the first subtitle in the subtitles to be output.

[0287] For example, subtitles 1 and subtitles 2 are currently being displayed on the current display device. When the first subtitle is subtitle 1, the user does not need to display subtitle 2, or wants to switch subtitle 2 to subtitle 3. The user can delete subtitle 1 through the remote control of the display device or the touch button of the user interface.

[0288] Because the subtitle synchronization rendering module does not provide a clock, subtitle decoding modules and rendering modules can be dynamically added and removed. If the user selects a different subtitle during playback, the data in each subtitle decoding module and subtitle synchronization rendering module is directly cleared and removed. Based on the new output elementary stream of the multi-subtitle selector, the corresponding subtitle decoding module and subtitle synchronization rendering module are created, and the connection between the data streams is established.

[0289] S2602: In response to a deletion operation by the user, stop synchronously rendering the first subtitles to stop synchronously displaying the first subtitles when playing the multimedia file.

[0290] In some embodiments, the stopping of synchronous rendering of the first subtitle is to find the first subtitle through the identification information of the target subtitle and stop synchronous rendering of the first subtitle.

[0291] In the above embodiment, by receiving a deletion operation from the user, the synchronous rendering of the first subtitles can be stopped in response to the deletion operation, so as to stop synchronously displaying the first subtitles when playing the multimedia file. The user can customize the deletion of unnecessary subtitles according to current needs, thereby improving the user experience.

[0292] As an extension and refinement of the above embodiment, as shown in FIG27 , the subtitle display method provided in the embodiment of the present application further includes the following steps S2701 and S2702:

[0293] S2701. Receive the user's add operation.

[0294] The adding operation is used to add a second subtitle other than the subtitles to be output for synchronous display.

[0295] For example, subtitles 3 and 4 are currently displayed on the current display device, and the second subtitle is subtitle 5. At this time, the user wants to display subtitle 5 on the display device, or wants to switch subtitle 4 to subtitle 5. The user can add subtitle 5 to display through the remote control or touch button of the user interface of the display device, or delete subtitle 4 first and then add subtitle 5.

[0296] S2702: Acquire subtitle data of the second subtitle and synchronous rendering logic corresponding to the second subtitle, and synchronously render the subtitle data of the second subtitle according to the global clock and the synchronous rendering logic corresponding to the second subtitle, so as to add synchronous display of the second subtitle.

[0297] In the above embodiment, by receiving the user's adding operation, the subtitle data of the second subtitle and the synchronous rendering logic corresponding to the second subtitle are obtained, and the subtitle data of the second subtitle is synchronously rendered according to the global clock and the synchronous rendering logic corresponding to the second subtitle, so as to add the synchronous display of the second subtitle. The user can customize and add the subtitles required for display according to current needs, thereby improving the user experience.

[0298] FIG28 is a schematic diagram of the structure of an electronic device according to some embodiments. The device can implement the audio processing method described in any embodiment of the present disclosure. The device specifically includes the following: a memory 2801 and a processor 2802.

[0299] The memory 2801 is configured to store computer instructions;

[0300] The processor 2802 is connected to the memory 2801 and is configured to perform the following steps when executing the computer instructions:

[0301] If it is determined that the multimedia file to be played contains multiple audio tracks, parsing the multimedia file to be played to obtain the multiple audio tracks of the multimedia file to be played;

[0302] Determining at least two target audio tracks from the multiple audio tracks, and establishing a correspondence between each target audio track and each playback headphone;

[0303] Decoding each target audio track based on a preset decoder corresponding to each target audio track to obtain a pulse modulation code corresponding to each target audio track;

[0304] Based on the pulse modulation code corresponding to each target audio track and the corresponding relationship between each target audio track and each playback earphone, the target audio track is played through the playback earphone corresponding to each target audio track.

[0305] As an optional implementation of the embodiment of the present disclosure, the processor 2802 is further configured to obtain parameter information corresponding to each target audio track; and establish the preset decoder corresponding to each target audio track based on the parameter information corresponding to each target audio track.

[0306] As an optional implementation of the embodiment of the present disclosure, the processor 2802 includes a hardware decoder and a soft decoder; the parameter information includes: audio sampling rate, number of channels and bit rate;

[0307] The processor 2802 is specifically configured to perform a product operation on the audio sampling rate, number of channels, and bit rate of each target audio track to obtain a product operation result for each target audio track; establish the hard decoder for the target audio track corresponding to the maximum product operation result, and establish the soft decoder for the other target audio tracks.

[0308] As an optional implementation of the embodiment of the present disclosure, the processor 2802 is further configured to synchronously play the video and subtitles included in the multimedia file to be played based on a target synchronization clock when playing the target audio track through the playback headphones corresponding to each target audio track.

[0309] As an optional implementation of the embodiment of the present disclosure, the processor 2802 is further configured to decode elementary streams corresponding to the video and subtitles contained in the multimedia file to be played, to obtain initial data corresponding to the video and subtitles respectively;

[0310] The processor 2802 is specifically configured to synchronously play the video and subtitles included in the to-be-played multimedia file based on the target synchronization clock and the initial data corresponding to the video and the subtitles, respectively.

[0311] As an optional implementation of the embodiment of the present disclosure, the processor 2802 is further used to determine the audio clock corresponding to each target audio track; and determine a target synchronization clock among multiple audio clocks, wherein the target synchronization clock is used to synchronously play the video, subtitles and at least two target audio tracks contained in the multimedia file to be played.

[0312] As an optional implementation of the embodiment of the present disclosure, the processor 2802 is specifically configured to determine a first target audio track corresponding to a maximum product operation result based on the product operation results of each target audio track; and use the audio clock of the first target audio track as the target synchronization clock.

[0313] As an optional implementation of the embodiment of the present disclosure, the processor 2802 is specifically configured to, when establishing the corresponding preset decoder for each target audio track, determine the last second target audio track for which the preset decoder is established; and use the audio clock of the second target audio track as the target synchronization clock.

[0314] As an optional implementation of the embodiment of the present disclosure, the processor 2802 is further configured to switch the target audio track being played when receiving a switching instruction input by the user.

[0315] In this way, the present embodiment parses the multimedia file to be played, obtains the multiple audio tracks of the multimedia file to be played, determines at least two target audio tracks from the multiple audio tracks, and establishes a corresponding relationship between each target audio track and each playback headphone, decodes each target audio track based on a preset decoder corresponding to each target audio track, and obtains a pulse modulation code corresponding to each target audio track, and plays the target audio track through the playback headphone corresponding to each target audio track based on the pulse modulation code corresponding to each target audio track and the corresponding relationship between each target audio track and each playback headphone. In the above process, the multiple audio tracks to be played can be determined according to different user needs, and the corresponding relationship between each target audio track and each playback headphone can be established, so that different playback headphones can play the multiple audio tracks contained in the same multimedia file to be played at the same time, thereby achieving the playback of multiple audio tracks, meeting the audio playback needs of different users for the multimedia file to be played, allowing each user to listen to the audio they need, and improving the user experience.

[0316] Figure 29 is a schematic diagram of the structure of a smart device according to some embodiments. As shown in Figure 29, the smart device includes a processor 2901 and a memory (also referred to as a storage device) 2902. The number of processors 2901 in the smart device can be one or more, and Figure 29 uses one processor 2901 as an example. The processor 2901 and the storage device 2902 in the smart device can be connected via a bus or other means, and Figure 29 uses a bus connection as an example.

[0317] Storage device 2902, as a computer-readable storage medium, can be used to store software programs, computer-executable programs, and modules, such as the program instructions / modules corresponding to the audio processing method in the embodiments of the present disclosure. Processor 2901 executes the software programs, instructions, and modules stored in storage device 2902 to execute various functional applications and data processing of the smart device, thereby implementing the audio processing method provided in the embodiments of the present disclosure.

[0318] The storage device 2902 may primarily include a program storage area and a data storage area. The program storage area may store an operating system and at least one application required for a function; the data storage area may store data created based on the use of the terminal, etc. Furthermore, the storage device 2902 may include high-speed random access memory and may also include non-volatile memory, such as at least one disk storage device, flash memory device, or other non-volatile solid-state memory device. In some instances, the storage device 2902 may further include a memory remotely located relative to the processor 2901, and such remote memory may be connected to the smart device via a network. Examples of the aforementioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.

[0319] The smart device provided in this embodiment can be used to execute the audio processing method provided in any of the above embodiments, and has corresponding functions and beneficial effects.

[0320] An embodiment of the present disclosure provides a computer-readable non-volatile storage medium, which stores a computer program. When the computer program is executed by a processor, it implements the various processes of any of the above methods and can achieve the same technical effects. To avoid repetition, it will not be described here.

[0321] The computer-readable storage medium may be a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.

[0322] In some embodiments, the embodiments of the present application provide a computer program product, which, when executed on a computer, enables the computer to implement any of the above methods.

[0323] For ease of explanation, the above description has been made in conjunction with specific embodiments. However, the above discussion of some embodiments is not intended to be exhaustive or to limit the embodiments to the specific forms disclosed above. Based on the above teachings, various modifications and variations can be obtained. The selection and description of the above embodiments are intended to better explain the principles and practical applications, so that those skilled in the art can better use the embodiments and various different variations of the embodiments suitable for specific use considerations.

Claims

1. An audio processing method, comprising: If it is determined that the multimedia file to be played contains multiple audio tracks, parsing the multimedia file to be played to obtain the multiple audio tracks of the multimedia file to be played; Determining at least two target audio tracks from the multiple audio tracks, and establishing a correspondence between each target audio track and each playback headphone; Decoding each target audio track based on a preset decoder corresponding to each target audio track to obtain a pulse modulation code corresponding to each target audio track; Based on the pulse modulation code corresponding to each target audio track and the corresponding relationship between each target audio track and each playback earphone, the target audio track is played through the playback earphone corresponding to each target audio track.

2. The method according to claim 1, before decoding each target audio track based on a preset decoder corresponding to each target audio track, further comprising: Obtain parameter information corresponding to each target audio track; Based on the parameter information corresponding to each target audio track, the preset decoder corresponding to each target audio track is established.

3. The method according to claim 2, wherein the preset decoder comprises a hard decoder and a soft decoder; The parameter information includes: audio sampling rate, number of channels and bit rate; The step of establishing the preset decoder corresponding to each target audio track based on the parameter information corresponding to each target audio track includes: Performing a product operation on the audio sampling rate, the number of channels, and the bit rate of each target audio track to obtain a product operation result of each target audio track; For the target audio track corresponding to the maximum product operation result, the hard decoder is established, and for the other target audio tracks, the soft decoder is established.

4. The method according to claim 1, further comprising: When the target audio track is played through the playback earphones corresponding to each target audio track, the video and subtitles included in the to-be-played multimedia file are played synchronously based on the target synchronization clock.

5. The method according to claim 4, wherein the step of synchronously playing the video and subtitles contained in the multimedia file to be played based on the target synchronization clock comprises: Decoding the elementary code streams corresponding to the video and subtitles contained in the multimedia file to be played to obtain initial data corresponding to the video and subtitles; Based on the target synchronization clock and the initial data corresponding to the video and the subtitles respectively, the video and subtitles included in the to-be-played multimedia file are synchronously played.

6. The method according to claim 4, before parsing the elementary streams corresponding to the video and subtitles contained in the multimedia file to be played to obtain the initial data corresponding to the video and subtitles, further comprising: Determine the audio clock corresponding to each target audio track; A target synchronization clock is determined among multiple audio clocks, wherein the target synchronization clock is used to synchronously play the video, subtitles and at least two target audio tracks contained in the multimedia file to be played.

7. The method according to claim 6, wherein determining a target synchronization clock from a plurality of audio clocks comprises: Determining a first target audio track corresponding to a maximum product operation result based on the product operation result of each target audio track; The audio clock of the first target audio track is used as the target synchronization clock.

8. The method according to claim 6, wherein determining a target synchronization clock from a plurality of audio clocks comprises: When establishing the corresponding preset decoder for each target audio track, determining the last second target audio track corresponding to the preset decoder to be established; The audio clock of the second target audio track is used as the target synchronization clock.

9. The method according to claim 1, further comprising: When a switching instruction input by the user is received, the target audio track currently being played is switched.

10. The method according to any one of claims 4 to 8, wherein when playing the target audio track through the playback earphone corresponding to each target audio track, playing the subtitles contained in the multimedia file to be played comprises: Receive the user's selection operation; In response to the selection operation, determining one or more subtitles to be output from the subtitles of the multimedia file to be played; Obtaining subtitle data of each channel of subtitles to be output; Obtaining a global clock of a playback pipeline of the multimedia file to be played and synchronization rendering logic corresponding to each channel of subtitles to be output; Based on the global clock and the synchronous rendering logic corresponding to each channel of the subtitles to be output, the subtitle data of each channel of the subtitles to be output are synchronously rendered respectively, so as to synchronously display each channel of the subtitles to be output when playing the multimedia file to be played.

11. The method according to claim 10, wherein obtaining the subtitle data of each channel of the subtitles to be output comprises: Obtaining the package file corresponding to the multimedia file to be played; Decapsulating the encapsulated file to obtain an elementary stream of the multimedia file to be played, wherein the elementary stream of the multimedia file to be played includes a subtitle elementary stream of at least one subtitle; Decode the subtitle elementary stream of the subtitles to be output in the subtitle elementary stream of the at least one subtitle to obtain subtitle data of the subtitles to be output.

12. The method according to claim 10, wherein obtaining the subtitle data of each channel of the subtitles to be output comprises: Acquire at least one external subtitle file of the multimedia file to be played based on the native layer of the display device; Decapsulating the at least one external subtitle file respectively to obtain a subtitle elementary stream corresponding to the at least one external subtitle file; The subtitle elementary stream of the subtitles to be output in the subtitle elementary stream corresponding to the at least one external subtitle file is decoded to obtain subtitle data of the subtitles to be output.

13. The method according to claim 10, wherein obtaining the subtitle data of each channel of the subtitles to be output comprises: Acquiring at least one external subtitle file of the multimedia file to be played based on the application layer of the display device; The at least one external subtitle file is parsed respectively to obtain subtitle data corresponding to each external subtitle file.

14. The method according to claim 13, wherein the synchronous rendering of the subtitle data of each channel of the subtitles to be output is performed based on the global clock and the synchronous rendering logic corresponding to each channel of the subtitles to be output, so as to synchronously display each channel of the subtitles to be output when playing the multimedia file to be played, comprising: Based on the global clock and the preset synchronous rendering logic, the subtitle data of each channel of the subtitles to be output are synchronously rendered, so that each channel of the subtitles to be output are synchronously displayed when the multimedia file is played.

15. The method according to claim 10, after obtaining the global clock of the playback pipeline of the multimedia file to be played and the synchronous rendering logic corresponding to each channel of the subtitles to be output, and before synchronously rendering the subtitle data of each channel of the subtitles to be output based on the global clock and the synchronous rendering logic corresponding to each channel of the subtitles to be output, the method further comprises: Adding identification information of the target subtitle to the rendering event corresponding to the target subtitle; The target subtitle is any one of the one or more subtitles to be output, and the identification information is used to uniquely identify the target subtitle.

16. The method according to claim 10, wherein obtaining the global clock of the playback pipeline of the multimedia file to be played comprises: Obtaining the audio clock of the multimedia file to be played; The audio clock is determined as a global clock of a playback pipeline of the multimedia file to be played.

17. The method according to any one of claims 10 to 16, further comprising: receiving a deletion operation from a user, wherein the deletion operation is used to stop synchronous display of a first subtitle in the subtitles to be output; In response to a deletion operation by the user, synchronous rendering of the first subtitles is stopped, so as to stop synchronously displaying the first subtitles when playing the multimedia file.

18. The method according to any one of claims 10 to 16, further comprising: receiving an adding operation from a user, wherein the adding operation is used to add a second subtitle other than the subtitles to be output for synchronous display; Acquire subtitle data of the second subtitle and synchronous rendering logic corresponding to the second subtitle, and synchronously render the subtitle data of the second subtitle according to the global clock and the synchronous rendering logic corresponding to the second subtitle to add synchronous display of the second subtitle.

19. An electronic device comprising: a memory configured to store computer instructions; The processor, connected to the memory, is configured to perform the following steps when executing the computer instructions: If it is determined that the multimedia file to be played contains multiple audio tracks, parsing the multimedia file to be played to obtain the multiple audio tracks of the multimedia file to be played; Determining at least two target audio tracks from the multiple audio tracks, and establishing a correspondence between each target audio track and each playback headphone; Decoding each target audio track based on a preset decoder corresponding to each target audio track to obtain a pulse modulation code corresponding to each target audio track; Based on the pulse modulation code corresponding to each target audio track and the corresponding relationship between each target audio track and each playback earphone, the target audio track is played through the playback earphone corresponding to each target audio track.