Audio track switching methods, devices, computing equipment, and storage media
By using a preset mapping table to correct the identification information during audio track switching, the problem of playback interruption caused by the uncertainty of audio and video track identification is solved, and seamless switching and stable playback of audio tracks are achieved.
Patent Information
- Application Number
- CN202310492609.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-05-04
- Publication Date
- 2025-11-14
- Estimated Expiration
- 2043-05-04
AI Technical Summary
In existing technologies, the uncertainty of audio and video track identifiers can lead to interruptions or errors in media program playback, especially in audio track switching scenarios where errors are more likely to occur.
By acquiring track data and identification information from the media encoding data, comparing them using a preset mapping table, correcting any inconsistent identification information, and storing the data in the corresponding data queue, the consistency of media types is ensured, enabling seamless switching of audio tracks.
It ensured the normal playback of media programs, avoided errors during audio track switching, achieved seamless switching, prevented stuttering and audio dropouts, and improved the user experience.
Smart Images

Figure CN116528018B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of media playback technology, specifically to an audio track switching method, apparatus, computing device, and storage medium. Background Technology
[0002] The media player obtains the media data stream and plays the media content through operations such as demultiplexing and decoding. Specifically, the player plays the media program by parsing the identifiers of the audio and video tracks in the media data stream, and it is necessary to maintain the consistency of the audio and video tracks to ensure uninterrupted and error-free playback. In addition, to meet different user audio needs, multiple audio tracks are compressed into the media data stream, and an audio track switching function is provided.
[0003] However, the identification of audio and video tracks is uncertain. For example, changes in the network can cause changes in the identification of audio and video tracks, or changes in the encapsulation format of the media data stream can cause changes in the identification of audio and video tracks. Both of these can lead to interruption or errors in media program playback, and errors are more likely to occur in application scenarios where audio tracks are switched. Summary of the Invention
[0004] The purpose of this application is to provide an audio track switching method, apparatus, computing device, and storage medium to solve the problem that the uncertainty of audio and video track identification in the prior art leads to media program playback interruption or errors.
[0005] According to one aspect of this application, an audio track switching method is provided, comprising:
[0006] Obtain data and identification information for multiple tracks in the media encoding data;
[0007] If the identification information and media type of multiple tracks do not conform to the preset mapping table, the identification information of multiple tracks is corrected, and the data of multiple tracks are stored in the data queues corresponding to the corrected identification information of multiple tracks respectively.
[0008] The data for multiple tracks includes audio data from multiple audio tracks and video data from multiple video tracks. A preset mapping table is used to record the mapping relationship between identification information and media type.
[0009] In response to a trigger operation on the target audio track switching element among multiple audio track switching elements, the current audio track is switched to the target audio track;
[0010] Decode and consume the audio data in the data queue of the target audio track.
[0011] Optionally, the method further includes:
[0012] If the identification information and media type of multiple tracks are determined to match the preset mapping table, the data of multiple tracks are stored in the data queues corresponding to the identification information of the multiple tracks respectively.
[0013] Optionally, the identification information of any track is determined according to the order in which the data of that track is encapsulated into the media encoding data.
[0014] Optionally, the method further includes:
[0015] Query at least one target identifier in the preset mapping table that has a mapping relationship with any media type;
[0016] If the identification information of at least one track corresponding to the media type is inconsistent with at least one target identification information, then it is determined that the identification information of multiple tracks and the media type do not conform to the preset mapping table.
[0017] The modification of the identification information of multiple tracks further includes: replacing the identification information of at least one track corresponding to the media type with information consistent with at least one target identification information.
[0018] Optionally, the method further includes:
[0019] Multiple initial identification information of multiple tracks is mapped to media types and recorded in a preset mapping table;
[0020] Multiple data queues are constructed based on the initial identification information of multiple orbits.
[0021] Optionally, decoding and consuming the audio data in the data queue of the target audio track further includes:
[0022] The audio data in the target audio track's data queue is read into the audio decoder for decoding, and the decoded audio data is saved to the audio consumption queue for consumption.
[0023] Optionally, the method further includes:
[0024] Retrieve multi-audio track identification information from media encoding data, and display multiple audio track switching elements based on the multi-audio track identification information.
[0025] Optionally, displaying multiple audio track switching elements based on multi-audio track identification information further includes:
[0026] The multi-audio track identification information is passed through to the interaction layer so that the interaction layer can parse the multi-audio track identification information to display multiple audio track switching elements.
[0027] Optionally, the multi-audio track identification information is constructed based on the live supplemental enhancement information.
[0028] According to another aspect of this application, an audio track switching device is provided, comprising:
[0029] The acquisition module is suitable for acquiring data and identification information of multiple tracks in media encoding data;
[0030] The correction module is suitable for correcting the identification information of multiple tracks when it is determined that the identification information and media type of multiple tracks do not conform to the preset mapping table;
[0031] The storage module is suitable for storing data from multiple tracks into data queues corresponding to the corrected identification information of the multiple tracks respectively;
[0032] The switching module is adapted to respond to the trigger operation of the target audio track switching element among multiple audio track switching elements, switch the current audio track to the target audio track, and decode and consume the audio data in the data queue of the target audio track.
[0033] Optionally, the storage module is further adapted to: if the identification information and media type of multiple tracks match a preset mapping table, store the data of multiple tracks into the data queue corresponding to the identification information of the multiple tracks respectively.
[0034] Optionally, the identification information of any track is determined according to the order in which the data of that track is encapsulated into the media encoding data.
[0035] Optionally, the correction module is further adapted to:
[0036] Query at least one target identifier in the preset mapping table that has a mapping relationship with any media type;
[0037] If the identification information of at least one track corresponding to the media type is inconsistent with at least one target identification information, then it is determined that the identification information of multiple tracks and the media type do not conform to the preset mapping table.
[0038] Replace the identification information of at least one track corresponding to the media type with the identification information of at least one target track.
[0039] Optionally, the device further includes: a recording module, adapted to map and record multiple initial identification information of multiple tracks to media types in a preset mapping table;
[0040] The building module is suitable for constructing multiple data queues based on the initial identification information of multiple tracks.
[0041] Optionally, the switching module is further adapted to:
[0042] The audio data in the target audio track's data queue is read into the audio decoder for decoding, and the decoded audio data is saved to the audio consumption queue for consumption.
[0043] Optionally, the device further includes a display module, adapted to acquire multi-audio track identification information in media encoding data and display multiple audio track switching elements based on the multi-audio track identification information.
[0044] Optionally, the display module is further adapted to: pass the multi-audio track identification information to the interaction layer, so that the interaction layer can parse the multi-audio track identification information to display multiple audio track switching elements.
[0045] Optionally, the multi-audio track identification information is constructed based on the live supplemental enhancement information.
[0046] According to another aspect of this application, a computing device is provided, comprising: a processor, a memory, a communication interface, and a communication bus, wherein the processor, the memory, and the communication interface communicate with each other through the communication bus;
[0047] The memory is used to store at least one executable instruction, which causes the processor to perform the operation corresponding to the audio track switching method described above.
[0048] According to another aspect of this application, a computer storage medium is provided, wherein the storage medium stores at least one executable instruction that causes a processor to perform an operation corresponding to the audio track switching method described above.
[0049] According to the audio track switching method, apparatus, computing device, and storage medium of this application, by comparing a preset mapping table with the track identification information and media type, and if it is determined that the track identification information and media type do not conform to the preset mapping table, the identification information of multiple tracks is corrected. This method masks situations where track identification information is incorrect due to specific reasons, ensuring that the media type of the data stored in the data queue remains consistent, thereby guaranteeing normal playback of media programs and preventing errors during audio track switching. Furthermore, when requesting playback of multiple audio track segments for the first time, the initial identification information and media type mapping relationship of each track are recorded as a basis for masking track order errors in subsequent processes, and data queues corresponding to each initial identification information are created. Subsequently, each track is obtained through demultiplexing. When checking the identification information, it determines whether the identification information and media type of each track match the recorded identification information and media type. If not, the identification information of each track is corrected, and the data of each track is stored in the data queue corresponding to the corrected identification information. This ensures that video data is always stored in the data queue used to store video data during the first playback, and audio data is always stored in the data queue used to store audio data during the first playback. This ensures that the playback of the upper-layer live content is uninterrupted and error-free, and can shield the problem of uncertainty in track identification. Furthermore, when the user switches audio tracks in the operation panel, the data reading channel of the audio decoder is changed to the data queue corresponding to the target audio track, avoiding stuttering and audio dropouts when switching audio tracks, and achieving seamless audio switching.
[0050] The above description is only an overview of the technical solution of this application. In order to better understand the technical means of this application and to implement it in accordance with the contents of the specification, and to make the above and other objects, features and advantages of this application more obvious and understandable, the following are specific embodiments of this application. Attached Figure Description
[0051] Various other advantages and benefits will become apparent to those skilled in the art upon reading the following detailed description of preferred embodiments. The accompanying drawings are for illustrative purposes only and are not intended to limit the scope of this application. Furthermore, the same reference numerals denote the same parts throughout the drawings. In the drawings:
[0052] Figure 1 A flowchart of the video audio track switching method provided in an embodiment of this application is shown;
[0053] Figure 2 A flowchart of a video audio track switching method provided in another embodiment of this application is shown;
[0054] Figure 3 A schematic diagram of media content containing multiple audio segments is shown in an embodiment of this application;
[0055] Figure 4a A schematic diagram of the playback interface when playing multiple audio segments is shown in an embodiment of this application;
[0056] Figure 4b A schematic diagram of the playback interface when playing multiple audio segments is shown in an embodiment of this application;
[0057] Figure 4c A schematic diagram of the playback interface when playing multiple audio segments is shown in an embodiment of this application;
[0058] Figure 5 A schematic diagram of the media content playback process in an embodiment of this application is shown;
[0059] Figure 6 A schematic diagram of the media content playback process in an embodiment of this application is shown;
[0060] Figure 7 A schematic diagram of the audio track switching device provided in an embodiment of this application is shown;
[0061] Figure 8 A schematic diagram of the structure of a computing device provided in an embodiment of this application is shown. Detailed Implementation
[0062] Exemplary embodiments of the present application will now be described in more detail with reference to the accompanying drawings. While exemplary embodiments of the present application are shown in the drawings, it should be understood that the present application may be implemented in various forms and should not be limited to the embodiments set forth herein. Rather, these embodiments are provided so that this application will be thorough and complete, and will fully convey the scope of the present application to those skilled in the art.
[0063] First, the terms and concepts involved in one or more embodiments of this application will be explained.
[0064] Media encoding data encapsulates the audio stream encoding of the audio track and the video stream encoding of the video track, with each stream occupying one track.
[0065] Media types include audio and video. Audio data is the data obtained from recording sound, while video data is the data obtained from recording video.
[0066] Figure 1 A flowchart of an audio track switching method provided in an embodiment of this application is shown. This method can be applied to any device with computing power, such as... Figure 1 As shown, the method includes the following steps:
[0067] Step S110: Obtain data and identification information of multiple tracks in the media encoding data.
[0068] The media encoding data includes audio data from multiple audio tracks and video data from multiple video tracks. The number of video tracks can be one or more. The media encoding data is obtained by encapsulating the data from multiple tracks. The media encoding data can be live stream encoding data or video-on-demand stream encoding data.
[0069] The live stream encoded data can be in HLS or FLV format. The video encoding format can be H264, HEVC, etc., and the audio encoding format is AAC, including LC and HE.
[0070] The track identification information is related to the order in which the track data is encapsulated into the media encoding data. For example, a specified value is used as the identification information of the track corresponding to the first piece of data encapsulated into the media encoding data. After that, for each additional piece of data, a fixed value is added to the identification information of the previous track to obtain the identification information of the track corresponding to the newly added data.
[0071] Step S120: If it is determined that the identification information and media type of multiple tracks do not conform to the preset mapping table, the identification information of multiple tracks is corrected, and the data of multiple tracks are stored in the data queues corresponding to the corrected identification information of multiple tracks.
[0072] The preset mapping table records the mapping relationship between identification information and media types, and there is also a one-to-one correspondence between identification information and data queues. For example, the mapping relationship between identification information and media types is preset, and data queues corresponding to each identification information are pre-built. For instance, if the mapping relationship between video type and identification information 0 is recorded, and the mapping relationship between audio type and identification information 1 and 2 is recorded, then when playing media content, the video decoder reads video data from the data queue corresponding to identification information 0 for decoding to display the video image, and the audio decoder reads audio data from the data queue corresponding to identification information 1 or 2 for decoding and playback of sound.
[0073] During media playback, issues such as network instability and changes in live stream addresses can occur. If these issues arise during the playback of multiple audio track segments, the identification information of each track will change. For a media player to play media programs correctly, it's crucial to maintain consistent track identification. The media type corresponding to a track is the media type of the track's data. Assuming two audio data streams are first encapsulated into media encoding data, and the video data is encapsulated into media encoding data later, the identification information for the two audio tracks is 0 and 1, while the identification information for the video track is 2. In this case, the two audio data streams will be stored in the data queues corresponding to identification information 0 and 1, while the video data will be stored in the data queue corresponding to identification information 2. The video decoder will then read the audio data from the data queue corresponding to identification information 0, and the audio decoder will read the video data from the data queue corresponding to identification information 2, resulting in both audio and video failing to play correctly, leading to playback errors in the media content.
[0074] Based on this, in this embodiment of the application, after obtaining the identification information of each track in the media encoding data through demultiplexing processing, it is compared with a preset mapping table to determine whether the identification information of multiple tracks and the media type conform to the preset mapping table. If they do not conform, the identification information of each track is corrected, and the track data is stored in the data queue corresponding to the corrected identification information of the track.
[0075] Step S130: In response to the triggering operation of the target audio track switching element among multiple audio track switching elements, the current audio track is switched to the target audio track.
[0076] When playing a segment containing multiple audio tracks, multiple audio track switching elements corresponding to the multiple audio tracks are displayed on the video screen. When the user clicks on any of these audio track switching elements, the data queue sent to the audio decoder is modified, switching the current audio track to the target audio track. The target audio track is the audio track corresponding to the target audio track switching element.
[0077] Step S140: Decode and consume the audio data in the data queue of the target audio track.
[0078] The audio data in the data queue corresponding to the target audio track is decoded and consumed, thereby playing the sound of the target audio track and achieving the effect of switching audio tracks.
[0079] According to the audio track switching method of this application embodiment, by comparing the preset mapping table with the track identification information and media type, and when it is determined that the track identification information and media type do not conform to the preset mapping table, the identification information of multiple tracks is corrected. In this way, the situation where the track identification information is incorrect due to specific reasons is shielded, which can ensure that the media type of the data stored in the data queue is always consistent, thereby ensuring the normal playback of media programs and avoiding errors during audio track switching.
[0080] Figure 2 A flowchart of an audio track switching method according to another embodiment of this application is shown. This method can be applied to any device with computing power, such as... Figure 2 As shown, the method includes the following steps:
[0081] Step S210: Map multiple initial identification information of multiple tracks to media types and record them in a preset mapping table; construct multiple data queues based on multiple initial identification information of multiple tracks.
[0082] In live streaming scenarios, as long as there is a playable address, the user experience of the upper layer can be guaranteed to be uninterrupted. At the same time, the track identifiers of different live stream formats are uncertain. That is, after the playback of multiple audio track segments begins, if there is a live stream interruption, network instability, change of live stream address, etc., it will cause a second request to play the multiple audio track segments, or even multiple requests to play. Once a new request is made, the identifier information of each track may be inconsistent with the identifier information of each track during the first playback, which will cause the media content to fail to play or make errors.
[0083] Based on this, the embodiments of this application record the identification information of each track during the first playback to mask the problem of track identification errors in subsequent processes. During the first playback, the requested media encoding data is demultiplexed to obtain the data and related fields of multiple tracks, such as initial identification information, media type, etc., and the mapping relationship between multiple initial identification information of multiple tracks and media type is recorded.
[0084] Simultaneously, based on the initial identification information of multiple tracks, data queues are constructed, each corresponding to a specific initial identification information. Each data queue stores the data for its corresponding track, establishing a correspondence between the data queues and the tracks through the identification information. For example, when requesting multiple audio track segments for the first time, the video track's identification information is 0, while the audio tracks' identification information is 1 and 2 respectively. The mapping relationship between video type and identification information 0, and between audio type and identification information 1 and 2, is recorded. Furthermore, data queue 0 corresponding to identification information 0, and data queues 1 and 2 corresponding to identification information 1 and 2 respectively, are constructed. The data for the track corresponding to identification information 0 (i.e., the video data of the video track) is stored in data queue 0, while the data for the tracks corresponding to identification information 1 and 2 (i.e., the audio data of the audio tracks) are stored in data queues 1 and 2 respectively. Additionally, depending on the media type of the data, the video decoder reads data from data queue 0 for decoding to play the video, and the audio decoder reads audio data from data queue 1 or 2 for decoding to play the sound.
[0085] Step S220: Obtain data and identification information of multiple tracks in the media encoding data.
[0086] The player continuously retrieves media encoding data from the server. If the media encoding data contains audio data from multiple audio tracks, it indicates that the corresponding media content is a multi-audio track segment. Multi-track compression is performed in advance, and the resulting media encoding data is encapsulated. By demultiplexing the media encoding data, the data for each track and its identification information are obtained.
[0087] Step S230: Determine whether the identification information and media type of multiple tracks match the preset mapping table.
[0088] If it is determined that the identification information and media type of multiple tracks do not conform to the preset mapping table, proceed to step S240; if it is determined that the identification information and media type of multiple tracks conform to the preset mapping table, proceed to step S250.
[0089] Specifically, the system queries a preset mapping table for at least one target identifier that has a mapping relationship with any media type. If the identifier of at least one track corresponding to that media type is inconsistent with at least one target identifier, then it is determined that the identifiers of multiple tracks and the media type do not conform to the preset mapping table. Conversely, if the identifier of at least one track corresponding to that media type is consistent with at least one target identifier, then it is determined that the identifiers of multiple tracks and the media type conform to the preset mapping table.
[0090] For example, a preset mapping table records the mapping relationship between video type and identifier 0, and the mapping relationship between audio type and identifiers 1 and 2. During the playback of multiple audio segments in a live stream, network instability causes a switch to a different encapsulation format of the live stream data. At this time, the identifier of the video track is 2, and the identifiers of the two audio tracks are 0 and 1 respectively. The identifier corresponding to the video type is 2, while the identifier mapped in the preset mapping table is 1. The identifier corresponding to the audio type is 0 and 1, while the identifier mapped in the preset mapping table is 1 and 2. Since the identifiers corresponding to both the audio and video types are inconsistent with the identifiers recorded in the preset mapping table, it is determined that they do not conform to the preset mapping table.
[0091] Step S240: Correct the identification information of multiple tracks and store the data of multiple tracks into the data queues corresponding to the corrected identification information of multiple tracks.
[0092] If it is determined that the identification information and media type of multiple tracks do not conform to the preset mapping table, the identification information of multiple tracks will be corrected according to the preset mapping table.
[0093] Specifically, the identifier information of at least one track corresponding to any media type is replaced with at least one target identifier information. Continuing with the example above, for tracks corresponding to video data, their identifier information is replaced with 0, that is, the identifier information corresponding to the track with identifier information 2 is replaced with 0. For tracks corresponding to audio data, the identifier information of the track with identifier information 0 is replaced with 2. The replaced identifier information of the video type tracks is consistent with the identifier information mapped to the video type recorded in the preset mapping table, and the replaced identifier information of multiple audio type tracks is consistent with the multiple identifier information mapped to the audio type recorded in the preset mapping table.
[0094] Subsequently, for each track, its data is stored in the data queue corresponding to its corrected identifier information. In this way, video data is always stored in the data queue built during the first playback, and audio data is always stored in the data queue built during the first playback. This ensures that the video decoder always reads video data and the audio decoder always reads audio data, thereby guaranteeing uninterrupted and error-free playback of the upper-layer media content and masking the uncertainty of track identifiers.
[0095] Step S250: Store the data of multiple tracks into the data queues corresponding to the identification information of the multiple tracks respectively.
[0096] If it is determined that the identification information of multiple tracks matches the media type in a preset mapping table, then for any track, its data is stored in the data queue corresponding to its identification information.
[0097] Step S260: In response to the triggering operation of the target audio track switching element among multiple audio track switching elements, the current audio track is switched to the target audio track.
[0098] Specifically, when playing a multi-audio-track segment, multiple audio track switching elements are displayed on the page. More specifically, multi-audio-track identifier information is added at the beginning of the segment, and when this identifier is obtained, multiple audio track switching elements are displayed accordingly.
[0099] Optionally, SEI (Supplemental Enhancement Information) is used to construct multi-audio track identification information. In the player, a callback function is used to obtain the SEI information. When the callback SEI information is multi-audio track identification information, the player passes the multi-audio track identification information to the interaction layer. Specifically, this is achieved by adding a bridge in the player. The interaction layer receives the multi-audio track identification information, parses it to obtain information such as the number of audio tracks and audio track names, and displays audio track switching buttons corresponding to each audio track according to the pre-configured interactive content.
[0100] Figure 3 The illustration shows a schematic diagram of media content containing multiple audio segments in an embodiment of this application, such as... Figure 3 As shown, the area within the dashed box represents a multi-audio track segment, while the areas before and after the dashed box represent single-audio track segments. For live streams, the multi-audio track segments are unpredictable. When a multi-audio track segment is reached, an SEI message is added to the live stream to trigger the audio track switching function. For on-demand streams, the position and duration of the multi-audio track segments are known. An SEI message is added in advance at an absolute time to trigger the audio track switching function at a set time. This is manifested as multiple audio track switching buttons. Clicking any of these buttons allows switching between audio tracks. When the multi-audio track segment ends, the audio track switching function terminates, including removing the multiple audio track switching buttons and automatically switching back to the default audio track, without affecting the live streaming and viewing experience of other single-track programs. By switching audio tracks within the same media program, users can consume the same video content with different audio versions.
[0101] Step S270: Read the audio data in the data queue of the target audio track into the audio decoder for decoding, and save the decoded audio data to the audio consumption queue for consumption.
[0102] The audio decoder is context-agnostic; it can decode as long as the audio data format is consistent. Therefore, in this embodiment, when switching audio by calling the relevant interface, the audio decoder's data reading channel is changed to the data queue corresponding to the target audio track selected by the user. The audio decoder decodes the read audio data and saves the decoded data to the audio consumption queue for consumption to play the sound of the target audio track. In other words, this application enables the switching operation via clicking buttons in the interaction layer, allowing control of the playback of a specified track in the audio track playback layer.
[0103] In existing technologies, a common approach is to use one audio track per audio decoder. For example, when switching audio tracks, the previous audio decoder is destroyed, and then the decoder for the new audio track is restarted. This can lead to noticeable stuttering, audio dropouts, slow switching speeds, and audio-visual asynchrony. In this embodiment, a single audio decoder is used, and audio track switching is achieved by changing the decoder's data read channel. This method enables seamless audio switching, avoiding problems such as slow switching speeds, stuttering, audio dropouts, and audio-visual asynchrony in audio track switching scenarios. It is particularly suitable for live streams with low latency and undetectable multi-audio-track content.
[0104] In an alternative approach, the method further includes: constructing an internal table to record the identification information of multiple audio tracks and their positions in the internal table, and simultaneously marking the identification information of the currently playing audio track and its position in the internal table.
[0105] In an optional approach, steps S210-S250 are implemented in a proxy layer within the player. This proxy layer can be a processing layer within the demultiplexing layer. By abstracting a low-level proxy layer within the player to handle various exceptions, the uninterrupted user experience of the upper layers of the player is ensured. The proxy layer constructs a structure of media type and identification information called stream maps (i.e., a preset mapping table). It also corrects the identification information of each track for the upper-layer consumer to read data. Furthermore, the processing logic and callback logic of the SEI can also be added to this proxy layer. This method ensures that both the APP client and the web client can achieve the functions of parsing media data from multiple audio tracks and switching audio tracks through the player.
[0106] Figure 4a A schematic diagram of the playback interface when playing multiple audio segments is shown in an embodiment of this application. When the multi-audio track identification information in the media encoding data is obtained through demultiplexing, a prompt element 41 is presented on the page to indicate that the currently playing segment contains multiple audio tracks.
[0107] Figure 4bThe diagram shows a playback interface when playing multiple audio segments in an embodiment of this application. After the user clicks on the prompt element 41, multiple audio track switching elements 42 are displayed.
[0108] Figure 4c The diagram shows a playback interface when playing multiple audio segments in an embodiment of this application. When the user clicks the audio track switching element 42, the currently playing audio track is switched to the audio track corresponding to the audio track switching element 42. The audio data of the data queue corresponding to the audio track is read, decoded and consumed to play the sound of the audio track. At the same time, the other audio track switching elements can be hidden. Figure 4c This corresponds to playing the sound of audio track V1, while audio track switching elements V2 and V3 are hidden.
[0109] If an audio track segment follows the end of a multi-audio track segment, then the multiple audio track switching elements are removed, and the system switches to the default audio track. Optionally, a preset identifier is added to the media encoding data based on the end position of the multi-audio track segment to indicate the end of the multi-audio track segment. This preset identifier is passed to the interaction layer, and the interaction layer removes the multiple audio track switching elements in response to the preset identifier.
[0110] Figure 5 This application illustrates a schematic diagram of the media content playback process in an embodiment of the present application, such as... Figure 5 As shown, the media encoding data of multiple audio tracks is separated into multiple video data and multiple audio data through a proxy layer. The video decoder and audio decoder read the video data and audio data for decoding, thereby playing the media program. It should be noted that the video data can also contain video data from multiple video tracks. In this case, multiple data queues are constructed to store the video data. Under certain circumstances, video track switching can also be supported. The implementation method for video track switching is similar to that for audio track switching.
[0111] Figure 6 This application illustrates a schematic diagram of the media content playback process in an embodiment of the present application, such as... Figure 6As shown, the media encoded data obtained from the edge server 61 is processed by the demultiplexing layer 62 to obtain a video queue 63 (i.e., a data queue for storing video data) and an audio queue group 66. The audio queue group 66 contains multiple audio queues (i.e., data queues for storing audio data) and is used to manage the control and consumption of multiple audio tracks. The video decoder 64 reads the video data from the video queue 63 and decodes it. The decoded video data is saved to the video consumption queue 65 for consumption. The audio decoder 67 reads the audio data from the audio queue of the currently selected audio track and decodes it. The decoded audio data is saved to the audio consumption queue 68 for consumption. The audio and video synchronization module 69 synchronizes the video and audio, and the media program is played through precise seek by the player. While playing the audio data of the currently selected audio track, the audio data of other audio tracks is synchronously discarded, i.e., the data queue is cleared.
[0112] According to the audio track switching method of this application embodiment, on the one hand, when requesting the playback of multiple audio track segments for the first time, the initial identification information of each track and the mapping relationship between the media type are recorded as the basis for shielding track identification errors in subsequent processes, and each data queue corresponding to each initial identification information is created; when the identification information of each track is subsequently obtained through demultiplexing, it is determined whether the identification information and media type of each track match the recorded identification information and media type. If not, the identification information of each track is corrected, and the data of each track is stored in the data queue corresponding to the corrected identification information, thereby ensuring that video data is always stored in the data queue for storing video data built during the first playback, and audio data is always stored in the data queue for storing audio data built during the first playback. This ensures that the playback of the upper-layer live content is uninterrupted and error-free, and can shield the problem of track identification uncertainty; on the other hand, when the user switches audio tracks in the operation panel, the data reading channel of the audio decoder is changed to the data queue corresponding to the target audio track being switched, avoiding stuttering, slow switching speed, audio-visual asynchrony, etc. during audio track switching, and achieving seamless audio switching.
[0113] Figure 7 A schematic diagram of the audio track switching device provided in an embodiment of this application is shown. Figure 7 As shown, the device includes:
[0114] Acquisition module 71 is adapted to acquire data and identification information of multiple tracks in media encoding data;
[0115] The correction module 72 is adapted to correct the identification information of multiple tracks when it is determined that the identification information and media type of multiple tracks do not conform to the preset mapping table;
[0116] The storage module 73 is adapted to store the data of multiple tracks into data queues corresponding to the corrected identification information of the multiple tracks respectively;
[0117] The switching module 74 is adapted to switch the current audio track to the target audio track in response to a trigger operation on the target audio track switching element among multiple audio track switching elements; and to decode and consume the audio data in the data queue of the target audio track.
[0118] In an alternative embodiment, the storage module 73 is further adapted to: if it is determined that the identification information and media type of multiple tracks conform to a preset mapping table, store the data of the multiple tracks into the data queue corresponding to the identification information of the multiple tracks respectively.
[0119] In one alternative approach, the identification information of any track is determined according to the order in which the data of that track is encapsulated into the media encoding data.
[0120] In one alternative embodiment, the correction module 72 is further adapted to:
[0121] Query at least one target identifier in the preset mapping table that has a mapping relationship with any media type;
[0122] If the identification information of at least one track corresponding to the media type is inconsistent with at least one target identification information, then it is determined that the identification information of multiple tracks and the media type do not conform to the preset mapping table.
[0123] Replace the identification information of at least one track corresponding to the media type with the identification information of at least one target track.
[0124] In one alternative embodiment, the device further includes: a recording module adapted to record multiple initial identification information of multiple tracks in a preset mapping table, mapping multiple media types;
[0125] The building module is suitable for constructing multiple data queues based on the initial identification information of multiple tracks.
[0126] In one alternative embodiment, the switching module 74 is further adapted to:
[0127] The audio data in the target audio track's data queue is read into the audio decoder for decoding, and the decoded audio data is saved to the audio consumption queue for consumption.
[0128] In one alternative embodiment, the device further includes a display module adapted to acquire multi-audio track identification information from media encoding data and display multiple audio track switching elements based on the multi-audio track identification information.
[0129] In one alternative approach, the display module is further adapted to: pass multi-audio track identification information to the interaction layer, so that the interaction layer can parse the multi-audio track identification information to display multiple audio track switching elements.
[0130] In one alternative approach, the multi-audio track identification information is constructed based on live supplemental enhancement information.
[0131] Through the above methods, on the one hand, when requesting the playback of multiple audio track segments for the first time, the initial identification information of each track and the mapping relationship between the media type are recorded. This serves as the basis for masking track order errors in subsequent processes, and data queues corresponding to each initial identification information are created. When the identification information of each track is subsequently obtained through demultiplexing, it is determined whether the identification information and media type of each track match the previously recorded identification information and media type. If not, the identification information of each track is corrected, and the data of each track is stored in the data queue corresponding to the corrected identification information. This ensures that video data is always stored in the data queue used to store video data built during the first playback, and audio data is always stored in the data queue used to store audio data built during the first playback. This ensures that the playback of the upper-layer live content is uninterrupted and error-free, and masks the uncertainty of track identification. On the other hand, when the user switches audio tracks in the operation panel, the data reading channel of the audio decoder is changed to the data queue corresponding to the target audio track being switched, avoiding stuttering, slow switching speed, and audio-visual asynchrony during audio track switching, thus achieving seamless audio switching.
[0132] This application provides a non-volatile computer storage medium storing at least one executable instruction that can execute the audio track switching method in any of the above method embodiments.
[0133] Figure 8 The diagram shows a structural schematic of a computing device provided in an embodiment of this application. The specific embodiments of this application do not limit the specific implementation of the computing device.
[0134] like Figure 8 As shown, the computing device may include: a processor 802, a communications interface 804, a memory 806, and a communications bus 808.
[0135] The processor 802, communication interface 804, and memory 806 communicate with each other via communication bus 808. Communication interface 804 is used to communicate with other network elements such as clients or other servers. The processor 802 executes program 810, specifically performing the relevant steps in the above-described embodiment of the audio track switching method for computing devices.
[0136] Specifically, program 810 may include program code that includes computer operation instructions.
[0137] The processor 802 may be a central processing unit (CPU), an application-specific integrated circuit (ASIC), or one or more integrated circuits configured to implement the embodiments of this application. The computing device includes one or more processors, which may be processors of the same type, such as one or more CPUs; or they may be processors of different types, such as one or more CPUs and one or more ASICs.
[0138] Memory 806 is used to store program 810. Memory 806 may include high-speed RAM memory, and may also include non-volatile memory, such as at least one disk storage device.
[0139] The algorithms or displays provided herein are not inherently related to any particular computer, virtual system, or other device. Various general-purpose systems can also be used in conjunction with the teachings herein. The required structure for constructing such systems is apparent from the above description. Furthermore, the embodiments of this application are not directed to any particular programming language. It should be understood that the content of this application described herein can be implemented using various programming languages, and the above description of specific languages is for the purpose of disclosing the best mode of implementation of this application.
[0140] Numerous specific details are set forth in the specification provided herein. However, it will be understood that embodiments of this application may be practiced without these specific details. In some instances, well-known methods, structures, and techniques have not been shown in detail so as not to obscure the understanding of this specification.
[0141] Similarly, it should be understood that, in order to simplify this application and aid in understanding one or more of the various aspects of the application, features of the embodiments of this application are sometimes grouped together in a single embodiment, figure, or description thereof in the above description of exemplary embodiments of this application. However, this disclosure method should not be construed as reflecting an intention that the claimed application requires more features than are expressly recited in each claim. Rather, as reflected in the claims, the application aspect lies in fewer than all features of the single embodiment disclosed above. Therefore, the claims following the detailed description are hereby expressly incorporated into that detailed description, wherein each claim itself is a separate embodiment of this application.
[0142] Those skilled in the art will understand that modules in the device of the embodiments can be adaptively changed and placed in one or more devices different from that embodiment. Modules, units, or components in the embodiments can be combined into a single module, unit, or component, and further, they can be divided into multiple sub-modules, sub-units, or sub-components. Except where at least some of such features and / or processes or units are mutually exclusive, any combination can be used to combine all features disclosed in this specification (including the accompanying claims, abstract, and drawings) and all processes or units of any method or device so disclosed. Unless expressly stated otherwise, each feature disclosed in this specification (including the accompanying claims, abstract, and drawings) may be replaced by an alternative feature that serves the same, equivalent, or similar purpose.
[0143] Furthermore, those skilled in the art will understand that although some embodiments herein include certain features included in other embodiments but not others, combinations of features from different embodiments are intended to be within the scope of this application and form different embodiments. For example, in the claims, any of the claimed embodiments can be used in any combination.
[0144] The various component embodiments of this application can be implemented in hardware, or as software modules running on one or more processors, or a combination thereof. Those skilled in the art will understand that microprocessors or digital signal processors (DSPs) can be used in practice to implement some or all of the functions of some or all of the components according to the embodiments of this application. This application can also be implemented as a device or apparatus program (e.g., a computer program and computer program product) for performing part or all of the methods described herein. Such an implementation of this application can be stored on a computer-readable medium, or can be in the form of one or more signals. Such signals can be downloaded from an Internet website, provided on a carrier signal, or provided in any other form.
[0145] It should be noted that the above embodiments are illustrative of this application and not restrictive, and those skilled in the art can devise alternative embodiments without departing from the scope of the appended claims. In the claims, any reference signs placed between parentheses should not be construed as limiting the claims. The word "comprising" does not exclude the presence of elements or steps not listed in the claims. The word "a" or "an" preceding an element does not exclude the presence of a plurality of such elements. This application can be implemented by means of hardware comprising several different elements and by means of a suitably programmed computer. In the unit claims enumerating several means, several of these means may be embodied by the same item of hardware. The use of the words first, second, and third, etc., does not indicate any order. These words can be interpreted as names. The steps in the above embodiments, unless otherwise specified, should not be construed as limiting the order of execution.
Claims
1. An audio track switching method, comprising: Obtain data and identification information for multiple tracks in the media encoding data; If it is determined that the identification information and media type of the multiple tracks do not conform to the preset mapping table, the identification information of the multiple tracks is corrected, and the data of the multiple tracks are stored in the data queues corresponding to the corrected identification information of the multiple tracks respectively. The media types are divided into audio types and video types; the data of the multiple tracks includes audio data of multiple audio tracks and video data of video tracks; the preset mapping table is used to record the mapping relationship between identification information and media types. In response to a trigger operation on the target audio track switching element among multiple audio track switching elements, the current audio track is switched to the target audio track; The audio data in the data queue of the target audio track is decoded and consumed.
2. The method according to claim 1, wherein, The method further includes: If the identification information and media type of the multiple tracks are determined to match the preset mapping table, the data of the multiple tracks are stored in the data queues corresponding to the identification information of the multiple tracks respectively.
3. The method according to claim 1, wherein, The identification information of any track is determined according to the order in which the data of that track is encapsulated into the media encoding data.
4. The method according to claim 1, wherein, The method further includes: Query at least one target identifier in the preset mapping table that has a mapping relationship with any media type; If the identification information of at least one track corresponding to the media type is inconsistent with the at least one target identification information, then it is determined that the identification information of the multiple tracks and the media type do not conform to the preset mapping table; The modification of the identification information of the plurality of tracks further includes: replacing the identification information of at least one track corresponding to the media type with information consistent with the at least one target identification information.
5. The method according to any one of claims 1-4, wherein, The method further includes: The initial identification information of the multiple tracks is mapped to the media type and recorded in a preset mapping table; Multiple data queues are constructed based on the initial identification information of the multiple orbits.
6. The method according to any one of claims 1-4, wherein, The decoding and consumption of audio data in the data queue of the target audio track further includes: The audio data in the data queue of the target audio track is read into the audio decoder for decoding, and the decoded audio data is saved to the audio consumption queue for consumption.
7. The method according to any one of claims 1-4, wherein, The method further includes: Obtain the multi-audio track identification information from the media encoding data, and display multiple audio track switching elements based on the multi-audio track identification information.
8. The method according to claim 7, wherein, The step of displaying multiple audio track switching elements based on the multi-audio track identification information further includes: The multi-audio track identification information is passed through to the interaction layer so that the interaction layer can parse the multi-audio track identification information to display multiple audio track switching elements.
9. The method according to claim 8, wherein, The multi-audio track identification information is constructed based on the live broadcast supplementary enhancement information.
10. An audio track switching device, comprising: The acquisition module is suitable for acquiring data and identification information of multiple tracks in media encoding data; The correction module is adapted to correct the identification information of the multiple tracks when it is determined that the identification information and media type of the multiple tracks do not conform to a preset mapping table; The storage module is suitable for storing data from multiple tracks into data queues corresponding to the corrected identification information of the multiple tracks respectively; The media types are divided into audio types and video types; the data of the multiple tracks includes audio data of multiple audio tracks and video data of video tracks; the preset mapping table is used to record the mapping relationship between identification information and media types. The switching module is adapted to respond to a trigger operation on a target audio track switching element among multiple audio track switching elements, switch the current audio track to the target audio track, and decode and consume the audio data in the data queue of the target audio track.
11. A computing device, comprising: The processor, memory, communication interface, and communication bus are provided, wherein the processor, memory, and communication interface communicate with each other via the communication bus. The memory is used to store at least one executable instruction that causes the processor to perform the operation corresponding to the audio track switching method as described in any one of claims 1-9.
12. A computer storage medium storing at least one executable instruction that causes a processor to perform an operation corresponding to the audio track switching method as described in any one of claims 1-9.
13. A computer program product comprising at least one executable instruction that causes a processor to perform an operation corresponding to the audio track switching method as described in any one of claims 1-9.
Citation Information
Patent Citations
Audio track switching method and system of multimedia player and corresponding player and equipment
CN104505109A
Audio and video synchronization method and device and storage medium
CN109729403A