Audio and video playing method and device, electronic equipment and storage medium

By pre-acquisitioning and decoding local files of audio and video, the rapid start of audio and video is achieved, solving the problem of slow audio and video playback speed and improving user experience.

CN120264069APending Publication Date: 2025-07-04TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410011092.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-01-03
Publication Date
2025-07-04

AI Technical Summary

Technical Problem

Among the existing audio and video playback methods, the audio and video playback speed is slow, resulting in poor user experience.

Method used

During the process of playing the target audio and video, the front-end local audio and video files in the audio and video to be played are pre-acquisced, and audio and video track data is extracted and decoded, and audio and video frame sequences are obtained, and audio and video rendering and frame sequence synchronization are performed when switching to playing the audio and video to be played.

Benefits of technology

It reduces the waiting time for playing audio and video to be played, shortens the start time, and improves the start speed and user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120264069A_ABST
    Figure CN120264069A_ABST
Patent Text Reader

Abstract

The invention discloses an audio and video playing method and device, electronic equipment and a storage medium. The embodiment of the invention is applied to various scenes such as cloud technology, artificial intelligence, intelligent traffic, auxiliary driving and the like. The method comprises the following steps: in a process of playing a target audio / video, acquiring a local audio / video file corresponding to an audio / video to be played; respectively extracting audio track data and video track data from the local audio and video files; performing audio decoding on the audio track data to obtain an audio frame sequence; performing video decoding on the video track data to obtain a video frame sequence; and in response to switching to play the to-be-played audio and video, performing audio and video rendering and frame sequence synchronization processing on the video frame sequence and the audio frame sequence to realize playing of the to-be-played audio and video. In the application, the local audio and video file of the to-be-played audio and video is pre-decoded in advance in the process of playing the target audio and video, the audio frame sequence and the video frame sequence are obtained, and the playing speed of the audio and video is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of audio - video technology, and more specifically, to an audio - video playing method, apparatus, electronic device, and storage medium. Background Art

[0002] In the application scenarios of audio - video, the audio - video application running on the terminal can be used to play audio - video, and the start - broadcast duration of the audio - video directly affects the user experience.

[0003] Currently, the audio - video application on the terminal can obtain the audio - video file corresponding to the to - be - played audio - video from the audio - video server in real - time, and then perform real - time decoding on the real - time obtained audio - video file through the audio - video decoder on the terminal to obtain the decoding result. Finally, the audio - video application on the terminal uses the player on the terminal to perform audio - video rendering according to the decoding result obtained by real - time decoding to realize the playing of the to - be - played audio - video.

[0004] However, when using this method to play the to - be - played audio - video, the start - broadcast speed of the audio - video is slow, resulting in a poor user experience. Summary of the Invention

[0005] In view of this, embodiments of the present application propose an audio - video playing method, apparatus, electronic device, and storage medium.

[0006] In a first aspect, embodiments of the present application provide an audio - video playing method, the method including: during the process of playing a target audio - video, pre - obtaining a partial audio - video file corresponding to the foremost audio - video in the to - be - played audio - video; performing audio - video track data extraction processing on the partial audio - video file to respectively extract audio track data and video track data; performing audio decoding on the audio track data to obtain an audio frame sequence; performing video decoding on the video track data to obtain a video frame sequence; in response to switching to playing the to - be - played audio - video, performing audio - video rendering and frame sequence synchronization processing on the video frame sequence and the audio frame sequence to realize the playing of the to - be - played audio - video.

[0007] In a second aspect, embodiments of the present application provide an audio - video playing apparatus, the apparatus including: an obtaining module, configured to pre - obtain a partial audio - video file corresponding to the foremost audio - video in the to - be - played audio - video during the process of playing a target audio - video; an extraction module, configured to perform audio - video track data extraction processing on the partial audio - video file to respectively extract audio track data and video track data; an audio decoding module, configured to perform audio decoding on the audio track data to obtain an audio frame sequence; a video decoding module, configured to perform video decoding on the video track data to obtain a video frame sequence; a rendering module, configured to, in response to switching to playing the to - be - played audio - video, perform audio - video rendering and frame sequence synchronization processing on the video frame sequence and the audio frame sequence to realize the playing of the to - be - played audio - video.

[0008] Optionally, the obtaining module is further configured to pre-obtain the file header information corresponding to the audio-video to be played during the playing of the target audio-video; correspondingly, the extracting module is further configured to determine the first track information of the audio track and the second track information of the video track in the partial audio-video file according to the file header information; and perform track data analysis on the partial audio-video file according to the first track information and the second track information to obtain audio track data and video track data.

[0009] Optionally, the extracting module is further configured to perform track data extraction on the partial audio-video file according to the first track information and the second track information to obtain initial audio track data and initial video track data; and perform synchronization processing on the initial audio track data and the initial video track data according to the timestamp information of the initial audio track data and the timestamp information of the initial video track data to obtain audio track data and video track data.

[0010] Optionally, the obtaining module is further configured to obtain the audio-video service address corresponding to the audio-video to be played from the cache, where the audio-video service address includes the address of the server corresponding to the audio-video to be played and the address of the content delivery network that sends the content of the audio-video to be played; the audio-video service address is obtained by domain name resolution during a historical time period; and during the playing of the target audio-video, pre-obtain the partial audio-video file corresponding to the frontmost audio-video in the audio-video to be played according to the audio-video service address.

[0011] Optionally, the obtaining module is further configured to obtain the network access conditions of multiple preset content delivery networks for the location of the client; the preset content delivery networks are the content delivery networks between the server corresponding to the audio-video to be played and the client; determine the target content delivery network that meets the preset network quality requirements from the multiple preset content delivery networks according to the network access conditions; and replace the address of the content delivery network in the audio-video service address with the address of the target content delivery network.

[0012] Optionally, when applied to the client, the audio decoding module is further configured to construct an audio structure according to the audio track data; and perform audio decoding on the audio track data according to the target audio decoder corresponding to the system environment information of the client to obtain an audio frame sequence.

[0013] Optionally, the audio decoding module is further configured to suspend the thread tasks in the client that are not related to the target audio decoder if there are thread tasks not related to the target audio decoder.

[0014] Optionally, the audio decoding module is further configured to cache the audio decoding result corresponding to the audio structure in real time during the audio decoding of the audio structure by the target audio decoder; after the cached audio decoding result reaches the preset audio data volume threshold, assemble the audio data according to the cached decoding result to obtain an audio frame sequence.

[0015] Optionally, the video decoding module is further configured to determine an initial video frame sequence according to the video track data; read the first key video frame in the initial video frame sequence through the target video decoder corresponding to the system environment information of the client; in response to obtaining the first key video frame, decompress the initial video frame sequence through the target video decoder according to the first key video frame to obtain a video frame sequence.

[0016] Optionally, the device further includes an allocation module, configured to determine a target number according to the maximum number of player instances displayed on the playback page, where the target number does not exceed the sum of the maximum number and 2, and the target number is not less than the maximum number; create the target number of player instances and add the target number of player instances to the player instance queue; before playing the target audio-visual content, obtain the player instance at the head of the queue in the player instance queue and allocate it to the target audio-visual content; during the process of playing the target audio-visual content through the player instance allocated to the target audio-visual content, obtain the player instance at the head of the queue in the player instance queue and allocate it to the audio-visual content to be played; correspondingly, the rendering module is further configured to, in response to switching to play the audio-visual content to be played, perform audio-visual rendering and frame sequence synchronization processing on the video frame sequence and the audio frame sequence through the player instance allocated to the audio-visual content to be played, so as to implement the playback of the audio-visual content to be played.

[0017] Optionally, the allocation module is further configured to, if the player instance allocated to the target audio-visual content is not within the playback page, add the player instance allocated to the target audio-visual content to the head of the player instance queue.

[0018] In a third aspect, an embodiment of the present application provides an electronic device, including a processor and a memory; computer-readable instructions are stored on the memory, and when the computer-readable instructions are executed by the processor, the above method is implemented.

[0019] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, on which computer-readable instructions are stored, and when the computer-readable instructions are executed by a processor, the above method is implemented.

[0020] In a fifth aspect, an embodiment of the present application provides a computer program product or a computer program including computer instructions, and when the computer instructions are executed by a processor, the above method is implemented.

[0021] An audio - video playing method, device, electronic device and storage medium provided by an embodiment of the present application. In the present application, during the process of playing a target audio - video, a local audio - video file corresponding to the foremost audio - video in the to - be - played audio - video is pre - obtained, and audio - video track data extraction processing and decoding are performed on the local audio - video file to obtain an audio frame sequence and a video frame sequence. It realizes that during the process of playing the target audio - video, the local audio - video file of the to - be - played audio - video is pre - downloaded and pre - decoded in advance (that is, before playing the to - be - played audio - video, the downloaded local audio - video file is decoded in advance using a decoder), to obtain an audio frame sequence and a video frame sequence. Thus, when switching to play the to - be - played audio - video, directly in response to switching to play the to - be - played audio - video, audio - video rendering and frame sequence synchronization processing are performed based on the video frame sequence and the audio frame sequence to realize the playing of the to - be - played audio - video, without the need to obtain the local audio - video file corresponding to the to - be - played audio - video and decode the local audio - video file of the to - be - played audio - video when it is needed to play the to - be - played audio - video. Therefore, when it is needed to play the to - be - played audio - video, the process of downloading and decoding the local audio - video file of the to - be - played audio - video is omitted, reducing the waiting time for playing the to - be - played audio - video, shortening the start - up duration of the to - be - played audio - video, increasing the start - up speed of the to - be - played audio - video, and enhancing the start - up user experience. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for description in the embodiments. Obviously, the following drawings are only some embodiments of the present application. For those skilled in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0023] Figure 1 A schematic diagram showing an application scenario applicable to the embodiments of the present application;

[0024] Figure 2 A flowchart showing an audio - video playing method proposed in an embodiment of the present application;

[0025] Figure 3 A schematic diagram showing a process of updating an audio - video service address according to the address of a target content distribution network in the embodiments of the present application;

[0026] Figure 4 A schematic diagram showing a working process of a demultiplexer in the embodiments of the present application;

[0027] Figure 5 A flowchart showing an audio - video playing method proposed in another embodiment of the present application;

[0028] Figure 6Shows a schematic diagram of a playback page in an embodiment of the present application;

[0029] Figure 7 Shows a schematic diagram of another playback page in an embodiment of the present application;

[0030] Figure 8 Shows a schematic diagram of the decoding process of an audio track data and a video track data in an embodiment of the present application;

[0031] Figure 9 Shows a schematic diagram of the switching process of a player instance in an embodiment of the present application;

[0032] Figure 10 Shows a schematic diagram of an audio - video playback process in an embodiment of the present application;

[0033] Figure 11 Shows a block diagram of an audio - video playback device proposed in an embodiment of the present application;

[0034] Figure 12 Shows a block diagram of an electronic device for executing the audio - video playback method according to an embodiment of the present application. Detailed implementation manners

[0035] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. According to the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without making creative efforts belong to the scope of protection of the present application.

[0036] In the following description, the terms "first / second" only distinguish similar objects and do not represent a specific order for the objects. It can be understood that "first / second" can be interchanged with a specific order or sequence when allowed, so that the embodiments of the present application described here can be implemented in an order other than that illustrated or described here.

[0037] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the technical field to which this application belongs. The terms used herein are only for the purpose of describing the embodiments of the present application and are not intended to limit the present application.

[0038] It should be noted that: "a plurality of" mentioned in this article refers to two or more. "And / or" describes the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. The character " / " generally represents an "or" relationship between the associated objects before and after.

[0039] In the embodiments of the present application, the term "module" or "unit" refers to a computer program with a predetermined function or a part of a computer program, which works together with other related parts to achieve a predetermined goal, and can be fully or partially implemented by using software, hardware (such as a processing circuit or a memory), or a combination thereof. Similarly, one processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be a part of an overall module or unit that includes the function of that block or unit.

[0040] The present application discloses an audio-video playing method, device, electronic device, and storage medium, which are related to various scenarios such as cloud technology, artificial intelligence, intelligent transportation, and assisted driving.

[0041] As Figure 1 shown, the application scenarios applicable to the embodiments of the present application include a terminal 20 and a server 10, and the terminal 20 and the server 10 are communicatively connected through a wired network or a wireless network. The terminal 20 can be a smart phone, a tablet computer, a laptop computer, a desktop computer, a smart home appliance, a vehicle-mounted terminal, an aircraft, a wearable device terminal, a virtual reality device, and other terminal devices that can perform page display, or run other applications that can call audio-video applications (such as instant messaging applications, shopping applications, search applications, game applications, forum applications, map traffic applications, etc.).

[0042] The server 10 can be an independent physical server, or a server cluster or a distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN (Content Delivery Network), and big data and artificial intelligence platforms. The server 10 can be used to provide services for applications running on the terminal 20.

[0043] Among them, the terminal 20 can obtain a target audio-video from the server 10, and during the process of playing the target audio-video, the terminal 20 obtains a local audio-video file corresponding to the frontmost audio-video in the to-be-played audio-video from the server 10. Then, the terminal 20 extracts audio track data and video track data from the local audio-video file; performs audio decoding on the audio track data to obtain an audio frame sequence; performs video decoding on the video track data to obtain a video frame sequence; then, in response to switching to playing the to-be-played audio-video, the terminal 20 performs audio-video rendering and frame sequence synchronization processing based on the video frame sequence and the audio frame sequence to implement the playing of the to-be-played audio-video.

[0044] In this embodiment, the terminal 20 may be installed with a client, and the terminal 20 implements part or all of the audio-video playing method through the client.

[0045] Please refer to Figure 2 , Figure 2 which shows a flowchart of an audio-video playing method proposed in an embodiment of the present application. The electronic device may be the terminal 20 in Figure 1 , and the method includes:

[0046] S110. During the process of playing a target audio-video, pre-obtain a local audio-video file corresponding to the foremost audio-video in the to-be-played audio-video.

[0047] The electronic device may be installed with a client, and the client may refer to an application program, a mini-program, or a web client that can implement audio-video playing. For example, the client may be a video application that can directly play audio-video, or the client may be a video website that can directly play audio-video, or the client may be a video mini-program in a shopping software.

[0048] The audio-video played in the mini-program is also called mini-program audio-video, and the playing of the mini-program audio-video depends on the <video>Component, in the applet <video>The underlying implementation of the component depends on the operating system and the native video player of the device. When used in a mini program <video>When using the component, relevant interfaces and functions of the underlying native video player are actually called through the applet framework.

[0049] In some embodiments, the client can respond to a play operation for a target audio-video (such as an operation to open a play page in the client or an operation to search for the name of the target audio-video on the play page of the client, etc.). The client obtains the target audio-video from the server. After obtaining the target audio-video, the client plays the target audio-video. During the playback of the target audio-video, the client obtains the local audio-video file corresponding to the foremost audio-video in the to-be-played audio-video, thereby realizing the pre-acquisition of the local audio-video file corresponding to the foremost audio-video in the to-be-played audio-video.

[0050] The to-be-played audio-video refers to the audio-video waiting to be played. In this way, when the target audio-video finishes playing or the user manually switches the target audio-video, the client can switch to play the to-be-played audio-video. Among them, the to-be-played audio-video can be one or more, and no specific limitation is made here. In some embodiments, the client in the terminal may include a video play queue, which includes multiple video identifiers. The client can sequentially play the audio-video indicated by each video identifier in the order of the video identifiers in the video play queue. Of course, in some embodiments, the play order of the audio-video can also be switched according to the user's needs. The audio-video indicated by the N video identifiers after the video identifier of the target audio-video in the video play queue (i.e., the audio-video that has not been played yet) can all be used as the to-be-played audio-video in this application, where N is a positive integer and N can be set according to actual needs. It can be understood that the larger N is, the greater the pre-decoding pressure on the terminal.

[0051] Among them, the server providing the target audio-video and the server providing the to-be-played audio-video can be the same or different. For example, the server pushes 10 audio-videos to the client in the form of a feed stream (Feed is a format of information, and the server uses it to deliver the audio-video to the client). The 10 audio-videos are played in sequence according to the play order. When playing the first audio-video, the first audio-video can be used as the target audio-video. Correspondingly, the second to tenth audio-videos can all be used as the to-be-played audio-video.

[0052] In some embodiments, the to-be-played audio-video can be an audio-video related to the target audio-video. For example, when the target audio-video is a TV drama (or movie), the to-be-played audio-video is the trailer or analysis of the TV drama or movie, etc. The to-be-played audio-video can also be an audio-video not related to the target audio-video. For example, when the target audio-video is an audio-video for learning English, the to-be-played audio-video is an audio-video of dancing.

[0053] In some possible scenarios, the server can determine multiple candidate audio-visuals for the client based on the account information of the account logged in by the client and the historical video playback records. Among the multiple candidate audio-visuals, the currently playing audio-visual is the target audio-visual, and the audio-visual to be played after the target audio-visual is the to-be-played audio-visual. For example, if the server determines through the account information and historical video playback records in the client that the user's interests are in anime and games, then the target audio-visual can be a game video, and the to-be-played audio-visual can be an anime audio-visual.

[0054] In addition, it should be noted that in the embodiments of the present application, the acquisition and application of the foregoing account information, historical video playback records, and other information all require the user's permission or consent, and the collection, use, processing, and storage of information such as the user's work type, education level, age, life stage (such as whether married), work location, and favorite products in the account information need to comply with the regulations of the region where the user is located.

[0055] In some embodiments, during the playback of the target audio-visual, the start data corresponding to the to-be-played audio-visual can be pre-acquired. The start data corresponding to the to-be-played audio-visual at least includes the partial audio-visual file corresponding to the audio-visual at the very front end of the to-be-played audio-visual, so as to achieve the purpose of pre-acquiring the partial audio-visual file corresponding to the audio-visual at the very front end of the to-be-played audio-visual.

[0056] Among them, the audio-visual at the very front end of the to-be-played audio-visual can be an audio-visual with a fixed duration or a fixed ratio. For example, the audio-visual at the very front end can be the audio-visual of the first 5s of the to-be-played audio-visual. Another example is that the audio-visual at the very front end can be the audio-visual of the first 1 / 10 of the to-be-played audio-visual. Among them, the partial audio-visual file can also include the advertisements and the screens corresponding to the transition animations.

[0057] In this embodiment, the partial audio-visual file is not all of the data of the to-be-played audio-visual. The data volume of the partial audio-visual file is small, so that the partial audio-visual file can be quickly acquired, the acquisition efficiency of the partial audio-visual file is improved, and the situation where the start speed of the to-be-played audio-visual is slow due to the slow acquisition of the partial audio-visual file is avoided.

[0058] In some embodiments, before S110, the method may further include: obtaining the audio-visual service address corresponding to the to-be-played audio-visual from the cache. The audio-visual service address includes the address of the server corresponding to the to-be-played audio-visual and the address of the content delivery network for sending the content of the to-be-played audio-visual. The audio-visual service address is obtained by domain name resolution in a historical time period. Correspondingly, S110 may include: during the playback of the target audio-visual, according to the audio-visual service address, pre-acquiring the partial audio-visual file corresponding to the audio-visual at the very front end of the to-be-played audio-visual.

[0059] Among them, the server refers to a server that can provide audio and video services (at least including the target audio and video and the audio and video to be played). Content Delivery Network (CDN): A distributed network service that is connected to the server and is responsible for quickly and reliably transmitting the audio and video provided by the server to the terminal. It is worth mentioning that the address of the server corresponding to the audio and video to be played in the audio and video service address is an IP address, which is obtained through domain name resolution in a historical period.

[0060] The historical period refers to the period before playing the audio and video to be played, which can be 1 minute, 2 minutes, etc. For example, when the client is a small program in an application, the historical period can be within 1 minute after starting the application where the client is located. Another example is that when the client is an application or a web client, the historical period can be within 1 minute after the previous time the client exits.

[0061] During the historical period, the client can associate and store the domain name address with the IP address determined by domain name resolution. Further, during the historical period, it can also determine the address of the content delivery network passed between the domain name address and the address where the client is located, and associate and store the domain name address, the IP address determined by domain name resolution, and the address of the content delivery network passed between the domain name address and the network address where the client is located in the cache. The client can send a domain name resolution request to a DNS (Domain Name System) server, and the DNS server performs domain name resolution according to the domain name resolution request to determine the IP address corresponding to the domain name address. During the process of playing the target audio and video, the start data of the audio and video to be played is obtained through the address of the server corresponding to the audio and video to be played in the audio and video service address and the address of the content delivery network. Since the address information obtained through historical domain name resolution is stored in the cache, if the domain name address of the server providing the audio and video to be played has been historically resolved, there is no need to perform domain name resolution for the audio and video to be played, which can improve the acquisition efficiency of the start data of the audio and video to be played and reduce the time-consuming of domain name resolution.

[0062] The DNS (Domain Name System) is a service on the Internet. As a distributed database that maps domain names and IP addresses to each other, it enables people to access the Internet more conveniently. DNS uses UDP port 53. Currently, the limit for the length of each level of domain name is 63 characters, and the total length of the domain name cannot exceed 253 characters. DNS is a system on the Internet for resolving the naming of online machines. Just like you need to know how to get to a friend's house before visiting, when a host on the Internet wants to access another host, it must first obtain its address. The IP address in TCP / IP consists of four numbers separated by ". " (here taking the IPv4 address as an example, and the IPv6 address is the same). It is always less convenient to remember than a name. Therefore, the domain name system is adopted to manage the correspondence between names and IPs.

[0063] In this embodiment, when the client needs to obtain the audio-video to be played, it obtains the start data of the audio-video to be played by obtaining the audio-video service address providing the audio-video to be played from the cache, without the need to first resolve the audio-video service address when obtaining the start data of the audio-video to be played, achieving the pre-resolution and acquisition of the audio-video service address, saving the resolution time of the audio-video service address, and improving the acquisition efficiency of the start data of the audio-video to be played.

[0064] In some embodiments, after obtaining the audio-video service address corresponding to the audio-video to be played from the cache, the method further includes: obtaining the network access conditions of multiple pre-set content distribution networks for the location of the client; the pre-set content distribution network is the content distribution network between the server corresponding to the audio-video to be played and the client; determining a target content distribution network that meets the pre-set network quality requirements from the multiple pre-set content distribution networks according to the network access conditions; and replacing the address of the content distribution network in the audio-video service address with the address of the target content distribution network. It can be understood that for any pre-set content distribution network, the pre-set content distribution network is located between the server providing the audio-video service to be played and the client receiving the audio-video to be played, and the pre-set content distribution network can send the audio-video provided by the server (at least including the audio-video to be played) to the client.

[0065] The access conditions of the pre-set content distribution network at least include the access peak period, access speed, access success rate, and hit rate of the pre-set content distribution network. The access conditions of the pre-set content distribution network can also include the start time period of a specified program, etc. The specified program can be, for example, a variety show with a large viewership or a short update of a TV drama. The location of the client refers to the location of the electronic device on which the client is installed. For example, the location of the client is b Street in City A, etc.

[0066] The preset network quality requirements may include at least one of the following: the acquisition period of the audio - video to be played is not during the access peak period, the access speed is greater than the access speed threshold, the access success rate is higher than the access success rate threshold, and the hit rate is higher than the hit rate threshold. When the access situation of the preset content delivery network also includes the broadcast period of a specified program, the preset network quality requirements may also include that the acquisition period of the audio - video to be played is not during the broadcast period of the specified program. When there are multiple preset content delivery networks that meet the preset network quality requirements, one preset content delivery network with the highest hit rate, the largest access success rate, and the largest access speed can be selected from the multiple preset content delivery networks that meet the preset network quality requirements as the target content delivery network.

[0067] After determining the target content delivery network, replace the address of the content delivery network in the audio - video service address with the address of the target content delivery network, which realizes the real - time update of the audio - video service address in the cache according to the network situation of the content delivery network, so as to quickly obtain the start - broadcast data of the audio - video according to the address of the target content delivery network in the audio - video service address next time. Since the network quality of the target content delivery network is relatively high, the efficiency of obtaining the start - broadcast data can be improved.

[0068] In this embodiment, the process of updating the audio - video service address according to the address of the target content delivery network can be as Figure 3 shown. Based on the user's playback operation (such as the operation of opening the playback page of the application), create a playback page for playing the audio - video. While playing the target audio - video on the playback page, send an audio - video request to download the start - broadcast data of the audio - video to be played. Among them, before downloading the start - broadcast data of the audio - video to be played, pre - resolve the audio - video service address of the audio - video to be played through the foregoing method, and during the process of playing the audio - video, perform real - time detection on the CDN to obtain the access situation of the CDN, and give feedback according to the access situation of the CDN to update the audio - video service address corresponding to the audio - video to be played.

[0069] S120. Perform audio - video track data extraction processing on the partial audio - video file to extract audio track data and video track data respectively.

[0070] As described above, during the process of playing the target audio - video, pre - obtain the partial audio - video file corresponding to the front - most audio - video in the audio - video to be played. The audio data corresponding to the front - most audio - video can be directly separated from the partial audio - video file as the audio track data. At the same time, the video data corresponding to the front - most audio - video is separated from the partial audio - video file as the video track data, so as to realize the audio - video track data extraction processing on the partial audio - video file and obtain the audio track data and the video track data.

[0071] In some embodiments, during the playback of the target audio-video, the file header information corresponding to the to-be-played audio-video may also be pre-obtained; the file header information may include information such as the encapsulation format, the number of tracks, and metadata of the partial audio-video file, and the file header information is used to indicate information such as the structure and attributes of the partial audio-video file. At this time, S120 may further include: determining the first track information of the audio track and the second track information of the video track in the partial audio-video file according to the file header information; performing track data analysis on the partial audio-video file according to the first track information and the second track information to obtain audio track data and video track data.

[0072] Among them, the track information may include information such as the encoding format, duration, and bit rate of the track data to be extracted. In the present application, for the convenience of distinction, the track information corresponding to the audio track is referred to as the first track information, and the track information corresponding to the video track is referred to as the second track information. For example, the first track information may include information such as the encoding format, duration, and bit rate of the audio track data to be extracted, and the second track information may include information such as the encoding format, duration, and bit rate of the video track data to be extracted.

[0073] In some embodiments, the start data corresponding to the to-be-played audio-video pre-obtained may also include file header information. During the playback of the target audio-video, the start data corresponding to the to-be-played audio-video is obtained, and then the file header information is extracted from the start data corresponding to the to-be-played audio-video to achieve pre-obtaining the file header information corresponding to the to-be-played audio-video.

[0074] In this embodiment, the partial audio-video file included in the start data is not all of the to-be-played audio-video data, and the data volume of the start data is small, so that the start data can be obtained quickly, the efficiency of obtaining the start data is improved, and the situation that the start speed of the to-be-played audio-video is slow due to the slow acquisition of the start data is avoided.

[0075] In this embodiment, a demultiplexer may be constructed before the audio-video playback is executed. The demultiplexer may include a reading module, an analysis module, and an extraction module, and each to-be-played audio-video may correspond to a demultiplexer. First, the file header information in the start data of the to-be-played audio-video may be read through the reading module, and then the analysis module analyzes information such as the encoding format, duration, and bit rate of the track data of each track (audio track and video track) according to the read file header information to obtain the first track information and the second track information. After that, the extraction module performs audio track data analysis on the partial audio-video file according to the first track information to obtain audio track data, and at the same time, the extraction module performs video track data analysis on the partial audio-video file according to the second track information to obtain video track data.

[0076] Optionally, in some embodiments, the foregoing analysis of the track data of the local audio-visual file according to the first track information and the second track information to obtain the audio track data and the video track data may include: extracting the track data of the local audio-visual file according to the first track information and the second track information to obtain the initial audio track data and the initial video track data; synchronizing the initial audio track data and the initial video track data according to the timestamp information of the initial audio track data and the timestamp information of the initial video track data to obtain the audio track data and the video track data.

[0077] Correspondingly, in this embodiment, the demultiplexer may further include a synchronization module. At this time, the extraction module analyzes the audio track data in the local audio-visual file according to the first track information to obtain the initial audio track data. At the same time, the extraction module analyzes the video track data in the local audio-visual file according to the second track information to obtain the initial video track data. Then, the synchronization module synchronizes the initial audio track data and the initial video track data according to the timestamp information of the initial audio track data and the timestamp information of the initial video track data to obtain the audio track data and the video track data, so that the audio track data and the video track data can be synchronized during playback, avoiding the situation of out-of-sync audio and video.

[0078] It can be understood that when there are multiple audio-visual files to be played, each audio-visual file to be played can correspond to its own demultiplexer, and multiple demultiplexers can be parallel, so that the track data of the start data of each audio-visual file to be played can be extracted simultaneously, obtaining the audio track data and the video track data of each audio-visual file to be played. Thus, the processing speed of the start data of the audio-visual files to be played is greatly improved, and the start speed of the audio-visual files to be played is increased.

[0079] In addition, each module in the demultiplexer can be asynchronous, that is, when multiple audio-visual files to be played share a demultiplexer, after any module in the demultiplexer completes the processing of the start data of one audio-visual file to be played, it can process the start data of other audio-visual files to be played, without waiting for all modules in the demultiplexer to complete the processing of the previous audio-visual file to be played and then processing the next audio-visual file to be played, thus realizing the reuse of the demultiplexer and improving the decoding efficiency of the audio-visual files to be played.

[0080] For example, after the reading module completes the extraction of the file header information of the audio-visual file a1 to be played, during the process of the analysis module obtaining the track information according to the file header information of the audio-visual file a1 to be played, the reading module in the idle state can extract the file header information of the audio-visual file a2 to be played.

[0081] In this embodiment, the working process of the demultiplexer is as Figure 4 As shown in Figure 4 , the reading module of the demultiplexer reads the file header information in the start data of the audio-visual content to be played, and then the analysis module analyzes the track information based on the read file header information (analyzing information such as the coding format, duration, and bit rate of the track data of each track (audio track and video track)) to obtain the first track information and the second track information. Then, the extraction module analyzes the audio track data in the local audio-visual file according to the first track information to obtain the initial audio track data. At the same time, the extraction module analyzes the video track data in the local audio-visual file according to the second track information to obtain the initial video track data. Finally, the synchronization module synchronizes the initial audio track data and the initial video track data according to the timestamp information of the initial audio track data and the timestamp information of the initial video track data to obtain the audio track data and the video track data.

[0082] S130: Perform audio decoding on the audio track data to obtain an audio frame sequence; perform video decoding on the video track data to obtain a video frame sequence.

[0083] After obtaining the audio track data and the video track data, an audio decoder can be used to decode the audio track data to obtain an audio frame sequence. The audio frame sequence includes multiple audio frames, and the audio frames in the audio frame sequence are arranged in the playback order. At the same time, a video decoder can be used to decode the video track data to obtain a video frame sequence. The video frame sequence includes multiple video frames, and the video frames in the video frame sequence are arranged in the playback order. Among them, the audio decoder and the video decoder are determined according to the system environment information of the electronic device (such as a terminal) where the client is located. The system environment information can indicate the decoders (audio decoder, video decoder) supported by the operating system in the electronic device. The audio decoder can be AAC LC, fdk-aac, HE-AACv2, etc. For example, since different operating systems, such as the Android system and the IOS system, may support different audio decoders, and the Android system and the IOS system may support different video decoders, based on the system environment information of the electronic device, the audio decoder and the video decoder supported by the operating system of the current electronic device can be determined, and audio decoding can be performed through the audio decoder supported by the operating system of the current electronic device, and video decoding can be performed through the video decoder supported by the operating system of the current electronic device. In this way, it is possible to avoid the need to reselect a decoder due to decoding through a decoder (audio decoder, video decoder) not supported by the current operating system.

[0084] In some embodiments, S130 may include: constructing an audio structure according to the audio track data; performing audio decoding on the audio structure according to the target audio decoder corresponding to the system environment information of the client to obtain an audio frame sequence; constructing a video structure according to the video track data; performing video decoding on the video structure according to the target video decoder corresponding to the system environment information of the client to obtain a video frame sequence. Among them, the audio structure and the video structure may refer to the PCM structure.

[0085] PCM (Pulse Code Modulation) is one of the encoding methods for digital communication. The main process is to sample analog signals such as voice and images at regular intervals to make them discrete, and at the same time round the sampled values to the nearest integer according to the quantization unit, and represent the amplitude of the sampled pulse with a group of binary codes.

[0086] An audio structure can be constructed based on the audio track data to store audio data information through the audio structure. The audio structure may include a file name (the name or identifier of the audio-visual to be played), a path (the cache directory of the audio structure corresponding to the audio-visual data to be played in the electronic device), a sampling rate, a frame rate, etc. Similarly, a video structure can be constructed based on the video track data to store video data information through the video structure. The video structure may include a file name (the name or identifier of the video to be played), a path (the cache directory of the video structure corresponding to the video data to be played in the electronic device), a sampling rate, a frame rate, etc.

[0087] After obtaining the audio-visual structure and the video structure, the audio-visual structure can be sent to the target audio decoder and the video structure can be sent to the target video decoder through FFmpeg (FFmpeg is an open-source computer program that can be used to record, convert digital audio and video, and convert them into streams), so as to decode the audio structure through the target audio decoder to obtain an audio frame sequence, and decode the video structure through the target video decoder to obtain a video frame sequence.

[0088] In this embodiment, according to the target audio decoder corresponding to the system environment information of the client, the audio structure is audio-decoded to obtain an audio frame sequence, including: in the process of the target audio decoder performing audio decoding on the audio structure, the audio decoding result corresponding to the audio structure is cached in real time; after the cached audio decoding result reaches the preset audio data volume threshold, the audio data is assembled according to the cached decoding result to obtain an audio frame sequence. Similarly, according to the target video decoder corresponding to the system environment information of the client, the video structure is video-decoded to obtain a video frame sequence, including: in the process of the target video decoder performing video decoding on the video structure, the video decoding result corresponding to the video structure is cached in real time; after the cached video decoding result reaches the preset video data volume threshold, the video data is assembled according to the cached decoding result to obtain a video frame sequence. Among them, the preset audio data volume threshold and the preset video data volume threshold can be set based on demand, and this application does not limit it.

[0089] The decoding process of the audio structure by the target audio decoder can be divided into two steps: decoding the audio structure to obtain the audio decoding result, and assembling the audio decoding result to obtain the audio frame sequence; similarly, the decoding process of the video structure by the target video decoder can be divided into two steps: decoding the video structure to obtain the video decoding result, and assembling the video decoding result to obtain the video frame sequence.

[0090] In some embodiments, the performance of the electronic device can be comprehensively considered to determine the preset audio data volume threshold and the preset video data volume threshold. For example, if the performance of the electronic device is high, the preset audio data volume threshold and the preset video data volume threshold can both be large. For another example, if the performance of the electronic device is low, the preset audio data volume threshold and the preset video data volume threshold can both be small.

[0091] When the cached audio decoding results reach the corresponding preset audio data volume threshold, data assembly is performed according to the cached audio decoding results to obtain the final audio frame sequence. At the same time, when the cached video decoding results reach the corresponding preset video data volume threshold, data assembly is performed according to the cached video decoding results to obtain the final video frame sequence. By reasonably setting the preset audio data volume threshold and the preset video data volume threshold, the situation in which the long assembly waiting time and low decoding efficiency are caused by subsequent data assembly when there are too many cached decoding results can be avoided, and the situation in which the assembly failure is caused by insufficient data when the cached decoding is too small, that is, the subsequent data assembly is avoided, thereby improving the decoding success rate and efficiency.

[0092] In some embodiments, video decoding is performed on video track data to obtain a video frame sequence, including: determining an initial video frame sequence according to the video track data; reading the first key video frame in the initial video frame sequence through a target video decoder corresponding to the system environment information of the client; in response to obtaining the first key video frame, decompressing the initial video frame sequence by the target video decoder according to the first key video frame to obtain a video frame sequence.

[0093] Among them, a key frame can refer to an I frame. An I frame (I frame) is also called an intra picture. An I frame is usually the first frame of each GOP (a video compression technology used by MPEG). After being moderately compressed, it serves as a random access reference point and can be regarded as an image. During the MPEG encoding process, part of the video frame sequence is compressed into I frames; part is compressed into P frames; and part is compressed into B frames. The I-frame method is an intra-frame compression method, also known as the "key frame" compression method. The I-frame method is a compression technology based on the Discrete Cosine Transform (DCT), and this algorithm is similar to the JPEG compression algorithm.

[0094] In this embodiment, a video structure can be constructed according to the video track data, and then video decoding is performed on the video structure by the target video decoder. During the process of the target video decoder performing video decoding on the video structure, the video decoding result corresponding to the video structure is cached in real time; after the cached video decoding result reaches a preset video data volume threshold, video data is assembled according to the cached decoding result to obtain an initial video frame sequence.

[0095] After obtaining the initial video frame sequence by decoding the video structure through the target video decoder, continue to read the first key video frame (that is, the first key frame) in the initial video frame sequence through the target video decoder, and then, decompress the video frame sequence by the target video decoder according to the first key video frame to obtain a video frame sequence.

[0096] After obtaining the first key frame, read the video frame sequence through the target video decoder and decompress the read video frame sequence to obtain a playback video frame sequence, where the data volume of the read video frame sequence = sampling rate * bit depth * number of channels * time, where time refers to the duration of the local audio-visual file.

[0097] The bit depth refers to the fact that when recording the colors of a digital image, a computer actually represents it using the bit depth required for each pixel. The reason a computer can display colors is that it uses a counting unit called "bit" to record the data representing the colors. When this data is recorded in the computer in a certain arrangement, it forms a computer file of a digital image. "Bit" is the smallest unit in a computer's memory, and it is used to record the color value of each pixel. The richer the colors of an image, the more "bits" there are. The number of bits used for each pixel in the computer is the "bit depth".

[0098] In some embodiments, before performing audio decoding on the audio structure according to the target audio decoder corresponding to the system environment information of the client to obtain an audio frame sequence, the method further includes: if there are thread tasks unrelated to the target audio decoder, suspending the thread tasks unrelated to the target audio decoder in the client. Similarly, before performing video decoding on the video structure according to the target video decoder corresponding to the system environment information of the client to obtain a video frame sequence, the method further includes: if there are thread tasks unrelated to the target video decoder, suspending the thread tasks unrelated to the target video decoder in the client.

[0099] Among them, the thread task can refer to a thread in the client. A thread is the smallest unit that the operating system can perform operation scheduling on. It is contained within a process and is the actual operating unit within the process. A thread refers to a single sequential control flow within a process. Multiple threads can be concurrent within a process, and each thread executes different tasks in parallel.

[0100] Before decoding the audio structure and the video structure, stop other thread tasks unrelated to the target audio encoder and the target video decoder, so that more resources can be freed up for the target audio decoder and the target video decoder to use, enabling the target audio decoder and the target video decoder to have sufficient computing resources for decoding, thereby achieving efficient and fast decoding, improving the decoding efficiency of the audio - video to be played, and further improving the start - up efficiency of the audio - video to be played.

[0101] After the decoding of the start - up data of the audio - video to be played is completed, the suspended other thread tasks unrelated to the target audio encoder and the target video decoder can be restarted so that the client can continue to normally run other thread tasks unrelated to the target audio encoder and the target video decoder.

[0102] S140. In response to switching to playing the audio - video to be played, perform audio - video rendering and frame sequence synchronization processing on the video frame sequence and the audio frame sequence to achieve the playback of the audio - video to be played.

[0103] After obtaining the video frame sequence and the audio frame sequence, the system can wait for the user to manually switch to play the audio-visual content to be played or automatically switch to play the audio-visual content to be played after the target audio-visual content finishes playing.

[0104] After switching to play the audio-visual content to be played, the client directly uses the player to perform audio-visual rendering on the video frame sequence and the audio frame sequence, displays the video frames in the video frame sequence on the screen of the electronic device, and passes the audio frames in the audio frame sequence to the audio output device. At the same time, during the rendering process, frame sequence synchronization processing is required for the audio frame sequence and the video frame sequence to ensure audio-visual synchronization during the playback of the audio-visual content to be played, realize the playback of the audio-visual content to be played, and ensure the quick start of the audio-visual content to be played. Among them, frame sequence synchronization processing can be performed according to the timestamp information of each audio frame in the audio frame sequence and the timestamp information of each video frame in the video frame sequence, so that the video frames and audio frames with the same timestamp information are played simultaneously.

[0105] In addition, during the process of using the player to render the video frame sequence and the audio frame sequence, rendering can be performed through the first track information corresponding to the audio track data and the second track information corresponding to the video track data, display the video frames in the video frame sequence on the screen of the electronic device, and pass the audio frames in the audio frame sequence to the audio output device to realize the playback of the audio-visual content to be played.

[0106] For example, audio rendering is performed according to the duration and bit rate in the first track information corresponding to the audio track data. Similarly, video rendering is performed according to the duration and bit rate in the second track information corresponding to the video track data to realize the playback of the audio-visual content to be played.

[0107] In this embodiment, during the playback of the target audio - video, a local audio - video file corresponding to the foremost audio - video in the to - be - played audio - video is pre - fetched, and audio - video track data extraction processing and decoding are performed on the local audio - video file to obtain an audio frame sequence and a video frame sequence. This realizes pre - downloading and pre - decoding of the local audio - video file of the to - be - played audio - video in advance during the playback of the target audio - video (that is, before playing the to - be - played audio - video, the downloaded local audio - video file is decoded in advance using a decoder), obtaining the audio frame sequence and the video frame sequence. Thus, when switching to play the to - be - played audio - video, directly in response to switching to play the to - be - played audio - video, audio - video rendering and frame sequence synchronization processing are performed based on the video frame sequence and the audio frame sequence to realize the playback of the to - be - played audio - video, without the need to obtain the local audio - video file corresponding to the to - be - played audio - video and decode the local audio - video file of the to - be - played audio - video when it is needed to play the to - be - played audio - video. Therefore, when it is needed to play the to - be - played audio - video, the process of downloading and decoding the local audio - video file of the to - be - played audio - video is omitted, reducing the waiting time for playing the to - be - played audio - video, shortening the start - up duration of the to - be - played audio - video, increasing the start - up speed of the to - be - played audio - video, and enhancing the start - up user experience.

[0108] Meanwhile, the start - up data includes the local audio - video file corresponding to the foremost audio - video in the to - be - played audio - video. Only the local audio - video file needs to be processed to obtain the audio frame sequence and the video frame sequence, without the need to process the entire to - be - played audio - video, reducing the amount of data processing. Thus, the audio frame sequence and the video frame sequence can be obtained quickly, thereby improving the start - up efficiency.

[0109] Please refer to Figure 5 , Figure 5 which shows a flowchart of an audio - video playback method proposed in another embodiment of the present application. The electronic device can be the Figure 1 terminal 20 in. The method includes:

[0110] S210. Determine the target number according to the maximum number of player instances displayed on the playback page; create the target number of player instances and add the target number of player instances to the player instance queue.

[0111] Among them, the target number does not exceed the sum of the maximum number and 2, and the target number is not less than the maximum number. For example, as Figure 6 shown, the player instance displayed on the playback page 50 is the player instance 501 (the part surrounded by the thick solid line frame). At this time, the maximum number of player instances displayed on the playback page 50 is 1, then the maximum value of the target number is 1 + 2 = 3. Of course, the target number can also be 2. Another example, as Figure 7 As shown, in the playback page 60, the player instances displayed are player instance 601 (the part enclosed by the first thick solid line box from top to bottom), player instance 602 (the part enclosed by the second thick solid line box from top to bottom), and player instance 603 (the part enclosed by the third thick solid line box from top to bottom). At this time, the maximum number of player instances displayed in the playback page 60 is 3. Then, the maximum value of the target number is 3 + 2 = 5, and the minimum value of the target number is 3.

[0112] After determining the target number, add the created player instances to the player instance queue, and the player instances in the player instance queue are arranged in order.

[0113] A player instance can refer to an example of a process for playing audio and video. The player of the electronic device can run the player instance to play the audio and video corresponding to the player instance. For example, the player instance corresponding to the audio and video s1 is t1. The player runs the player instance t1, and the player instance t1 runs the video frame sequence and audio frame sequence of the audio and video s1 to achieve the playback of the audio and video s1.

[0114] S220. Before playing the target audio and video, obtain the player instance located at the head of the queue in the player instance queue and assign it to the target audio and video.

[0115] Before playing the target audio and video, obtain the player instance located at the head of the queue from the player instance queue and assign it to the target audio and video. At this time, the number of player instances in the player instance queue is reduced by one. For example, the number of player instances in the created player instance queue is 5. After obtaining the player instance located at the head of the queue from the player instance queue and assigning it to the target audio and video, the number of player instances in the player instance queue is 4, and the second player instance in the original player instance queue becomes the head of the new player instance queue.

[0116] S230. During the process of playing the target audio and video, obtain the local audio and video file; perform audio and video track data extraction processing on the local audio and video file to extract the audio track data and video track data respectively; perform audio decoding on the audio track data to obtain the audio frame sequence; perform video decoding on the video track data to obtain the video frame sequence.

[0117] Among them, the description of S230 refers to the description of S110 - S130 above and will not be elaborated here.

[0118] S240. During the process of playing the target audio and video through the player instance assigned to the target audio and video, obtain the player instance located at the head of the queue in the player instance queue and assign it to the audio and video to be played.

[0119] During the playback of the target audio - video, based on the video playback queue described above, player instances can be sequentially allocated to each to - be - played audio - video according to the order of the video identifiers of the to - be - played audio - videos in the video playback queue. It can be understood that after a player instance is allocated to an audio - video, the corresponding player instance is taken out from the player instance queue.

[0120] For example, the number of player instances in the created player instance queue is 5. After obtaining the player instance at the head of the queue from the player instance queue and allocating it to the target audio - video, the number of player instances in the player instance queue becomes 4, and the second player instance in the original player instance queue becomes the head of the new player instance queue. During the playback of the target audio - video, when allocating a player instance to the to - be - played audio - video, the player instance allocated to the to - be - played audio - video is the player instance at the head of the player instance queue including the above - mentioned 4 player instances. At this time, the number of player instances in the player instance queue is 3.

[0121] In some embodiments, during the playback of the target audio - video, the number of to - be - played audio - videos that need to be pre - decoded can be restricted by the number of player instances in the player instance queue. That is, during the playback of the target audio - video, if the number of player instances in the player instance queue is K, then the number of to - be - played audio - videos that need to be pre - decoded is K. In this way, it can be avoided that a large amount of computing overhead and a large amount of memory are occupied due to pre - decoding a large number of audio - videos in advance.

[0122] S250. In response to switching to play the to - be - played audio - video, perform audio - video rendering and frame sequence synchronization processing on the video frame sequence and the audio frame sequence through the player instance allocated to the to - be - played audio - video, so as to realize the playback of the to - be - played audio - video.

[0123] After allocating a player instance to the to - be - played audio - video, the player instance allocated to the to - be - played audio - video can be run through the player of the electronic device, so that the player instance allocated to the to - be - played audio - video performs audio - video rendering and frame sequence synchronization processing based on the video frame sequence and the audio frame sequence, and plays the to - be - played audio - video on the playback page.

[0124] In this embodiment, the decoding process is as Figure 8 As shown, a PCM structure: an audio structure and a video structure is constructed based on the audio track data and the video track data; then, a decoder: a target audio decoder and a target video decoder is obtained according to the system environment information; then, other thread tasks of the client are traversed to determine whether there are thread tasks unrelated to the decoder (the target audio decoder and the target video decoder). If so, the thread tasks unrelated to the decoder are suspended, and then the step of decoding the first key frame is executed. If not, the step of decoding the first key frame is directly executed. After decoding the first key frame, the video frame sequence is decompressed according to the first key frame to obtain the playing video frame sequence. Finally, the audio and video are drawn through the playing video frame sequence and the audio frame sequence to play the audio and video to be played. Finally, the previously suspended thread tasks are restarted.

[0125] In some embodiments, after step S250, the method may further include: if the player instance allocated for the target audio and video is not within the playing page, add the player instance allocated for the target audio and video to the head of the player instance queue. When the player instance allocated for the target audio and video is not within the playing page and it is determined that the target audio and video is no longer being played, the player instance allocated for the target audio and video can be retrieved and added to the head of the player instance queue, so that the player instance can be reused by the subsequent played audio and video without the need to re-create a player instance.

[0126] For example, the audio and video are pushed to the client in the form of a feed stream (Feed is a format of information, and the server delivers the audio and video to the client through it). The maximum number of player instances displayed on the playing page is 1, the determined target number is 3, and the created player instance queue also includes 3 player instances. As Figure 9 shown in a of , the player instance allocated for the paused audio and video d1 is e1 (outside the playing page 70), the player instance allocated for the playing audio and video d2 displayed on the playing page 70 is e2, and the player instance allocated for the paused audio and video d3 is e3 (outside the playing page 70). At this time, the number of player instances in the player instance queue is 0. Among them, the filled color being gray indicates that the allocated player instance is the player instance e1, the filled color being white indicates that the allocated player instance is the player instance e2, and the filled color being dark gray indicates that the allocated player instance is the player instance e3.

[0127] During the process of the user swiping up the screen, the situation of the playing page 70 of the electronic device is as Figure 9 shown in b of . The player instance allocated for the audio and video d2 is about to slide out of the playing page 70, and the player instance allocated for the audio and video d3 is about to enter the playing page 70. When the player instance allocated for the audio and video d2 completely slides out of the playing page 70, the player instance allocated for the audio and video d3 completely enters the playing page 70, as Figure 9 As shown in c in , at this time, the playback of the audio-video d2 is paused, and the audio-video d3 is played through the player instance e3; at the same time, it is determined that the player instance allocated to the audio-video d1 is not within the playback page. At this time, even if the user swipes down the screen, the audio-video that first appears within the playback page 70 is also the player instance allocated to d2. Therefore, the player instance of the audio-video d1 can be retrieved and placed at the head of the player instance queue. At this time, the player instance queue only includes one player instance e1. It is determined that the audio-video that may continue to be played after the audio-video is d4, and the player instance e1 is allocated to the audio-video d4. At this time, the number of player instances in the player instance queue becomes 0 again.

[0128] In this embodiment, the playback process of the audio-video to be played is as Figure 10 shown. The user initiates a playback operation on the client side. The client of the electronic device creates a playback page according to the playback operation. The client prepares the playback page (which may include playing advertisements, entering the page, and transition animations, etc.), and then obtains the player kernel according to the system environment information of the electronic device where the client is located (including many algorithms related to the picture and sound quality such as decoding, buffering, and frequency conversion. Among them, in this embodiment, the main things obtained are the video decoding algorithm and the audio decoding algorithm, that is, determining the target video decoder and the target audio decoder), realizes the creation of the player kernel, and then creates player instances to obtain a player instance queue. After that, a request for obtaining the audio-video can be sent to the server.

[0129] The server issues the start-up data of the audio-video according to the audio-video request. The client parses the start-up data to obtain the file header information, and extracts the audio track data and the video track data according to the file header information. Finally, the client decodes according to the audio track data and the video track data to obtain the audio frames and the playback video frames. Finally, the player instance corresponding to the audio-video is rendered through the player, realizing the fast start-up of the audio-video.

[0130] In this embodiment, according to the size of the playback page, the target quantity is determined, and a player instance queue including the target quantity of player instances is created. Before playing the audio-video, a corresponding player instance is allocated to each audio-video, realizing the pre-creation of the player instances required for the audio-video playback, and there is no need to create player instances when playing the audio-video, saving the player instance creation time. At the same time, the created player instances can be reused, without creating a player instance for each audio-video, reducing the creation overhead of the player instances and improving the start-up efficiency.

[0131] Please refer to Figure 11 , Figure 11 which shows a block diagram of an audio-video playback device proposed in an embodiment of the present application. The device 1000 includes:

[0132] An obtaining module 1010, configured to pre-obtain a local audio-video file corresponding to the foremost audio-video in the to-be-played audio-video during the process of playing a target audio-video;

[0133] An extraction module 1020, configured to perform audio-video track data extraction processing on the local audio-video file to respectively extract audio track data and video track data;

[0134] An audio decoding module 1030, configured to perform audio decoding on the audio track data to obtain an audio frame sequence;

[0135] A video decoding module 1040, configured to perform video decoding on the video track data to obtain a video frame sequence;

[0136] A rendering module 1050, configured to, in response to switching to playing the to-be-played audio-video, perform audio-video rendering and frame sequence synchronization processing on the video frame sequence and the audio frame sequence to implement playing of the to-be-played audio-video.

[0137] Optionally, the obtaining module 1010 is further configured to pre-obtain file header information corresponding to the to-be-played audio-video during the process of playing the target audio-video; correspondingly, the extraction module 1020 is further configured to obtain the file header information from the start data; determine first track information of an audio track and second track information of a video track in the local audio-video file according to the file header information; and perform track data analysis on the local audio-video file according to the first track information and the second track information to obtain the audio track data and the video track data.

[0138] Optionally, the extraction module 1020 is further configured to perform track data extraction on the local audio-video file according to the first track information and the second track information to obtain initial audio track data and initial video track data; and perform synchronization processing on the initial audio track data and the initial video track data according to the timestamp information of the initial audio track data and the timestamp information of the initial video track data to obtain the audio track data and the video track data.

[0139] Optionally, the obtaining module 1010 is further configured to obtain an audio-video service address corresponding to the to-be-played audio-video from a cache, where the audio-video service address includes an address of a server corresponding to the to-be-played audio-video and an address of a content delivery network that sends the content of the to-be-played audio-video; the audio-video service address is obtained by performing domain name resolution in a historical time period; and during the process of playing the target audio-video, pre-obtain a local audio-video file corresponding to the foremost audio-video in the to-be-played audio-video according to the audio-video service address.

[0140] Optionally, the obtaining module 1010 is further configured to obtain the network access conditions of multiple preset content distribution networks for the location of the client; the preset content distribution network is the content distribution network between the server corresponding to the audio-video to be played and the client; determine a target content distribution network that meets the preset network quality requirements from the multiple preset content distribution networks according to the network access conditions; replace the address of the content distribution network in the audio-video service address with the address of the target content distribution network.

[0141] Optionally, applied to the client, the audio decoding module 1030 is further configured to construct an audio structure according to the audio track data; perform audio decoding on the audio track data according to the target audio decoder corresponding to the system environment information of the client to obtain an audio frame sequence.

[0142] Optionally, the audio decoding module 1030 is further configured to suspend the thread tasks in the client that are not related to the target audio decoder if there are thread tasks not related to the target audio decoder.

[0143] Optionally, the audio decoding module 1030 is further configured to cache the audio decoding results corresponding to the audio structure in real time during the process of audio decoding of the audio structure by the target audio decoder; after the cached audio decoding results reach the preset audio data volume threshold, assemble the audio data according to the cached decoding results to obtain an audio frame sequence.

[0144] Optionally, the video decoding module 1040 is further configured to determine an initial video frame sequence according to the video track data; read the first key video frame in the initial video frame sequence through the target video decoder corresponding to the system environment information of the client; in response to obtaining the first key video frame, decompress the initial video frame sequence by the target video decoder according to the first key video frame to obtain a video frame sequence.

[0145] Optionally, the device further includes an allocation module, configured to determine a target number according to the maximum number of player instances displayed on the play page, where the target number does not exceed the sum of the maximum number and 2, and the target number is not less than the maximum number; create the target number of player instances and add the target number of player instances to the player instance queue; before playing the target audio-video, obtain the player instance at the head of the queue in the player instance queue and allocate it to the target audio-video; during the process of playing the target audio-video through the player instance allocated to the target audio-video, obtain the player instance at the head of the queue in the player instance queue and allocate it to the audio-video to be played; correspondingly, the rendering module 1050 is further configured to, in response to switching to play the audio-video to be played, perform audio-video rendering and frame sequence synchronization processing on the video frame sequence and the audio frame sequence through the player instance allocated to the audio-video to be played, so as to realize the playing of the audio-video to be played.

[0146] Optionally, the allocation module is further configured to add the player instance allocated for the target audio - video to the head of the player instance queue if the player instance allocated for the target audio - video is not within the playback page.

[0147] It should be noted that the device embodiments in this application correspond to the foregoing method embodiments. The specific principles in the device embodiments can be referred to in the content of the foregoing method embodiments, and will not be elaborated here.

[0148] Figure 12 The block diagram of an electronic device for executing the audio - video playback method according to an embodiment of the present application is shown. The electronic device may be Figure 1 the terminal 20, etc. It should be noted that Figure 12 the computer system 1200 of the electronic device shown is only an example, and should not impose any limitations on the functions and usage scope of the embodiments of the present application.

[0149] As Figure 12 shown, the computer system 1200 includes a central processing unit (CPU) 1201, which can perform various appropriate actions and processes according to the program stored in the read - only memory (ROM) 1202 or the program loaded from the storage section 1208 into the random access memory (RAM) 1203, such as executing the method in the above - mentioned embodiment. In the RAM 1203, various programs and data required for system operation are also stored. The CPU 1201, ROM 1202, and RAM 1203 are connected to each other through a bus 1204. The input / output (I / O) interface 1205 is also connected to the bus 1204.

[0150] The following components are connected to the I / O interface 1205: an input section 1206 including a keyboard, a mouse, etc.; an output section 1207 including, for example, a cathode ray tube (CRT), a liquid crystal display (LCD), etc., and a speaker, etc.; a storage section 1208 including a hard disk, etc.; and a communication section 1209 including a network interface card such as a LAN (Local Area Network) card, a modem, etc. The communication section 1209 performs communication processing via a network such as the Internet. A drive 1210 is also connected to the I / O interface 1205 as required. A removable medium 1211, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 1210 as required so that a computer program read therefrom is installed into the storage section 1208 as required.

[0151] Specifically, according to an embodiment of the present application, the processes described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present application includes a computer program product that includes a computer program carried on a computer-readable medium, and the computer program includes program code for performing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from a network via the communication section 1209, and / or installed from the removable medium 1211. When the computer program is executed by a central processing unit (CPU) 1201, various functions defined in the system of the present application are performed.

[0152] It should be noted that the computer-readable medium shown in the embodiments of the present application can be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. A computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of a computer-readable storage medium can include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM), a flash memory, an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present application, a computer-readable storage medium can be any tangible medium that contains or stores a program, and this program can be used by or in conjunction with an instruction execution system, apparatus, or device. In the present application, a computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. A computer-readable signal medium can also be any computer-readable medium other than a computer-readable storage medium, and this computer-readable medium can send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on a computer-readable medium can be transmitted using any appropriate medium, including but not limited to: wireless, wired, etc., or any suitable combination of the above.

[0153] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present application. Among them, each block in the flowchart or block diagram can represent a module, a program segment, or a part of code, and the above module, program segment, or part of code contains one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than that marked in the accompanying drawings. For example, two consecutive blocks shown can actually be executed substantially in parallel, and they can sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram or flowchart, and the combination of blocks in the block diagram or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.

[0154] The units involved in the embodiments described in this application can be implemented in software or in hardware, and the described units can also be provided in a processor. Among them, the names of these units do not constitute a limitation to the unit itself in certain cases.

[0155] As another aspect, the present application also provides a computer-readable storage medium, which may be included in the electronic device described in the above embodiments; or may exist alone without being assembled into the electronic device. The above computer-readable storage medium carries computer-readable instructions, and when the computer-readable storage instructions are executed by a processor, the methods in any of the above embodiments are implemented.

[0156] According to one aspect of the embodiments of the present application, there is provided a computer program product or a computer program, the computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. The processor of the electronic device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the electronic device executes the methods in any of the above embodiments.

[0157] It should be noted that although several modules or units of the devices for action execution are mentioned in the above detailed description, this division is not mandatory. In fact, according to the embodiments of the present application, the features and functions of two or more of the above-mentioned modules or units can be embodied in one module or unit. Conversely, the features and functions of one module or unit described above can be further divided and embodied by multiple modules or units.

[0158] Through the description of the above embodiments, those skilled in the art can easily understand that the example embodiments described herein can be implemented by software or by a combination of software and necessary hardware. Therefore, the technical solutions according to the embodiments of the present application can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (such as a CD-ROM, a USB flash drive, a mobile hard disk, etc.) or on a network, including several instructions to enable an electronic device (such as a personal computer, a server, a touch terminal, or a network device, etc.) to execute the methods according to the embodiments of the present application.

[0159] Other embodiments of the present application will readily occur to those skilled in the art upon considering the specification and practicing the embodiments disclosed herein. The present application is intended to cover any variations, uses, or adaptations of the present application, which follow the general principles of the present application and include known common general knowledge or conventional technical means in the technical field not disclosed in the present application. It should be understood that the present application is not limited to the exact structures described above and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present application is only limited by the appended claims.

[0160] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application and are not intended to limit them. Although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments or equivalently replace some of the technical features. These modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present application.< / video> < / video> < / video>

Claims

1. An audio-video playing method, characterized in that, The method includes: During the process of playing the target audio - video, pre - obtain the local audio - video file corresponding to the foremost audio - video in the to - be - played audio - video; Perform audio - video track data extraction processing on the local audio - video file to respectively extract audio track data and video track data; Perform audio decoding on the audio track data to obtain an audio frame sequence; Perform video decoding on the video track data to obtain a video frame sequence; In response to switching to playing the to - be - played audio - video, perform audio - video rendering and frame sequence synchronization processing on the video frame sequence and the audio frame sequence to realize the playing of the to - be - played audio - video.

2. The method according to claim 1, characterized in that The method further includes: During the process of playing the target audio - video, pre - obtain the file header information corresponding to the to - be - played audio - video; The performing audio - video track data extraction processing on the local audio - video file to respectively extract audio track data and video track data includes: According to the file header information, determine the first track information of the audio track and the second track information of the video track in the local audio - video file; According to the first track information and the second track information, perform track data analysis on the local audio - video file to obtain audio track data and video track data.

3. The method according to claim 2, wherein The according to the first track information and the second track information, performing track data analysis on the local audio - video file to obtain audio track data and video track data includes: According to the first track information and the second track information, perform track data extraction on the local audio - video file to obtain initial audio track data and initial video track data; According to the timestamp information of the initial audio track data and the timestamp information of the initial video track data, perform synchronization processing on the initial audio track data and the initial video track data to obtain audio track data and video track data.

4. The method according to claim 1, wherein Before the pre - obtaining the local audio - video file corresponding to the foremost audio - video in the to - be - played audio - video during the process of playing the target audio - video, the method further includes: Obtain the audio - video service address corresponding to the to - be - played audio - video from the cache, where the audio - video service address includes the address of the server corresponding to the to - be - played audio - video and the address of the content delivery network that sends the content of the to - be - played audio - video; the audio - video service address is obtained through domain name resolution in a historical time period; The pre - obtaining the local audio - video file corresponding to the foremost audio - video in the to - be - played audio - video during the process of playing the target audio - video includes: During the process of playing the target audio - video, according to the audio - video service address, pre - obtain the local audio - video file corresponding to the foremost audio - video in the to - be - played audio - video.

5. The method according to claim 4, wherein Applied to the client, after obtaining the audio - video service address corresponding to the to - be - played audio - video from the cache, the method further includes: Obtain the network access conditions of multiple preset content delivery networks for the location of the client; the preset content delivery networks are the content delivery networks between the server corresponding to the to - be - played audio - video and the client; Determine a target content delivery network that meets the preset network quality requirements from the multiple preset content delivery networks according to the network access situation; Replace the address of the content delivery network in the audio-visual service address with the address of the target content delivery network.

6. The method according to claim 1, characterized in that, Applied to a client, the audio decoding of the audio track data to obtain an audio frame sequence includes: Construct an audio structure according to the audio track data; Perform audio decoding on the audio structure according to the target audio decoder corresponding to the system environment information of the client to obtain an audio frame sequence.

7. The method according to claim 6, wherein Before performing audio decoding on the audio structure according to the target audio decoder corresponding to the system environment information of the client to obtain an audio frame sequence, the method further includes: If there are thread tasks unrelated to the target audio decoder, suspend the thread tasks unrelated to the target audio decoder in the client.

8. The method according to claim 6, wherein The performing audio decoding on the audio structure according to the target audio decoder corresponding to the system environment information of the client to obtain an audio frame sequence includes: During the process of the target audio decoder performing audio decoding on the audio structure, cache the audio decoding result corresponding to the audio structure in real time; After the cached audio decoding result reaches the preset audio data volume threshold, assemble the audio data according to the cached decoding result to obtain an audio frame sequence.

9. The method according to claim 1, wherein The video decoding of the video track data to obtain a video frame sequence includes: Determine an initial video frame sequence according to the video track data; Read the first key video frame in the initial video frame sequence through the target video decoder corresponding to the system environment information of the client; In response to obtaining the first key video frame, decompress the initial video frame sequence according to the first key video frame through the target video decoder to obtain a video frame sequence.

10. The method according to claim 1, characterized in that, Before prefetching the local audio-visual file corresponding to the most front-end audio-visual in the to-be-played audio-visual during the process of playing the target audio-visual, the method further includes: Determine a target number according to the maximum number of player instances displayed on the play page, where the target number does not exceed the sum of the maximum number and 2, and the target number is not less than the maximum number; Create a target number of player instances and add the target number of player instances to the player instance queue; Before playing the target audio-visual, obtain the player instance at the head of the queue in the player instance queue and allocate it to the target audio-visual; During the process of playing the target audio-visual through the player instance allocated to the target audio-visual, obtain the player instance at the head of the queue in the player instance queue and allocate it to the to-be-played audio-visual; The responding to switching to playing the to-be-played audio-visual and performing audio-visual rendering and frame sequence synchronization processing on the video frame sequence and the audio frame sequence to realize the playing of the to-be-played audio-visual includes: In response to switching to play the to-be-played audio-video, perform audio-video rendering and frame sequence synchronization processing on the video frame sequence and the audio frame sequence through the player instance allocated for the to-be-played audio-video, so as to realize the playing of the to-be-played audio-video.

11. The method according to claim 10, characterized in that, After performing audio-video rendering and frame sequence synchronization processing on the video frame sequence and the audio frame sequence through the player instance allocated for the to-be-played audio-video in response to switching to play the to-be-played audio-video, so as to realize the playing of the to-be-played audio-video, the method further includes: If the player instance allocated for the target audio-video is not within the playing page, add the player instance allocated for the target audio-video to the head of the player instance queue.

12. An audio and video playback device, characterized in that, The device includes: An acquisition module, configured to pre-acquire a partial audio-video file corresponding to the foremost audio-video in the to-be-played audio-video during the playing of the target audio-video; An extraction module, configured to perform audio-video track data extraction processing on the partial audio-video file to respectively extract audio track data and video track data; An audio decoding module, configured to perform audio decoding on the audio track data to obtain an audio frame sequence; A video decoding module, configured to perform video decoding on the video track data to obtain a video frame sequence; A rendering module, configured to, in response to switching to play the to-be-played audio-video, perform audio-video rendering and frame sequence synchronization processing on the video frame sequence and the audio frame sequence, so as to realize the playing of the to-be-played audio-video.

13. An electronic device, characterized in that, including: a processor; a memory, on which computer-readable instructions are stored, and when the computer-readable instructions are executed by the processor, the method according to any one of claims 1-11 is implemented.

14. A computer-readable storage medium, characterized in that, On which computer-readable instructions are stored, and when the computer-readable instructions are executed by the processor, the method according to any one of claims 1-11 is implemented.

15. A computer program product or a computer program, characterized in that, including computer instructions, characterized in that when the computer instructions are executed by the processor, the method according to any one of claims 1-11 is implemented.