Artificial Intelligence-Based Video Data Processing Method, Apparatus, Device, and Medium

By determining target video frames and GOPs based on playback time and decode states, parallel decoding is employed to enhance decoding efficiency and meet user demands for accelerated video playback.

CN115484460BActive Publication Date: 2025-07-15BEIJING BAIDU NETCOM SCI & TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211123362.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-09-15
Publication Date
2025-07-15
Estimated Expiration
2042-09-15

AI Technical Summary

Technical Problem

The prior art cannot meet the user's needs for accelerated playback during video playback, and the traditional time-domain dimension one-frame decoding method is inefficient.

Method used

By determining the video frame and GOP set at the target playback time, the target decoder set is used for parallel decoding, and the GOP set to be decoded flexibly determines the GOP set to be decoded, improving the decoding efficiency.

Benefits of technology

It has achieved improvement in video decoding efficiency, supports fast playback of video frames, and meets users' accelerated playback needs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115484460B_ABST
    Figure CN115484460B_ABST
Patent Text Reader

Abstract

The present disclosure provides a video data processing method, apparatus, device, and medium based on artificial intelligence. In the field of artificial intelligence, it specifically relates to cloud computing, cloud storage, and distributed storage technologies and can be applied in intelligent cloud scenarios. The specific implementation solution is as follows: Determine a target video frame corresponding to a target playback moment from a video to be played, and use the candidate group of pictures (GOP) to which the target video frame belongs as the target GOP, and use the candidate GOP set to which the target GOP belongs as the target GOP set; Determine a GOP set to be decoded from at least one candidate GOP set corresponding to the video to be played according to the first decoding state of the target video frame and / or the second decoding state of the target GOP set; Perform parallel decoding on the GOP set to be decoded based on a target decoder set to obtain video frames to be played. Through the above technical solution, the efficiency of video decoding can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of artificial intelligence, specifically to cloud computing, cloud storage, and distributed storage technologies, and can be applied in intelligent cloud scenarios. Background Art

[0002] A variety of videos have become an indispensable part of people's daily life entertainment. The transmission and playback of videos require encoding and decoding processes. Video decoding is the process of using a specific method to restore the digitally encoded video data to its represented video content, or to convert an electrical pulse signal into the video information, data, etc. it represents. Only after decoding can the video be normally displayed to the user.

[0003] Currently, during the playback of videos, mainly decoding is performed frame by frame in the time domain dimension and then played. This way of decoding according to the order of the video stream cannot meet the user's need to accelerate video playback, so there is an urgent need for improvement. Summary of the Invention

[0004] The present disclosure provides a method, apparatus, device, and medium for processing video data based on artificial intelligence.

[0005] According to one aspect of the present disclosure, there is provided a method for processing video data based on artificial intelligence, the method comprising:

[0006] Determine a target video frame corresponding to a target playback moment from a video to be played, and use the candidate group of pictures (GOP) to which the target video frame belongs as the target GOP, and use the candidate GOP set to which the target GOP belongs as the target GOP set;

[0007] Determine a GOP set to be decoded from at least one candidate GOP set corresponding to the video to be played according to the first decoding state of the target video frame and / or the second decoding state of the target GOP set;

[0008] Perform parallel decoding on the GOP set to be decoded based on a target decoder set to obtain video frames to be played.

[0009] According to another aspect of the present disclosure, there is provided a device for processing video data based on artificial intelligence, the device comprising:

[0010] A target GOP set determination module, configured to determine a target video frame corresponding to a target playback moment from a video to be played, and use the candidate group of pictures (GOP) to which the target video frame belongs as the target GOP, and use the candidate GOP set to which the target GOP belongs as the target GOP set;

[0011] A GOP set to be decoded determination module, configured to determine a GOP set to be decoded from at least one candidate GOP set corresponding to the video to be played according to the first decoding state of the target video frame and / or the second decoding state of the target GOP set;

[0012] A video to be played determination module, configured to perform parallel decoding on the GOP set to be decoded based on a target decoder set to obtain video frames to be played.

[0013] According to another aspect of the present disclosure, there is provided an electronic device, which includes:

[0014] At least one processor; and

[0015] A memory communicatively connected to the at least one processor; wherein,

[0016] The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the artificial intelligence-based video data processing method according to any embodiment of the present disclosure.

[0017] According to another aspect of the present disclosure, there is provided a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause a computer to execute the artificial intelligence-based video data processing method according to any embodiment of the present disclosure.

[0018] According to another aspect of the present disclosure, there is provided a computer program product, including a computer program, where the computer program, when executed by a processor, implements the artificial intelligence-based video data processing method according to any embodiment of the present disclosure.

[0019] According to the technology of the present disclosure, the decoding efficiency of a video can be improved.

[0020] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. Description of the Drawings

[0021] The drawings are used to better understand the solution and do not constitute a limitation to the present disclosure. Among them:

[0022] Figure 1 is a flowchart of an artificial intelligence-based video data processing method provided according to an embodiment of the present disclosure;

[0023] Figure 2 is a flowchart of another artificial intelligence-based video data processing method provided according to an embodiment of the present disclosure;

[0024] Figure 3 is a flowchart of yet another video data processing method based on artificial intelligence provided according to an embodiment of the present disclosure;

[0025] Figure 4 is a flowchart of still another video data processing method based on artificial intelligence provided according to an embodiment of the present disclosure;

[0026] Figure 5 is a schematic structural diagram of a video data processing apparatus based on artificial intelligence provided according to an embodiment of the present disclosure;

[0027] Figure 6 is a block diagram of an electronic device for implementing the video data processing method based on artificial intelligence according to an embodiment of the present disclosure. Detailed implementation manners

[0028] The following makes an explanation of exemplary embodiments of the present disclosure with reference to the accompanying drawings. Various details of the embodiments of the present disclosure are included to facilitate understanding, and they should be considered merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described here without departing from the scope and spirit of the present disclosure. Similarly, for the sake of clarity and conciseness, the description below omits the description of well-known functions and structures.

[0029] It should be noted that the terms "first", "second", "target", "candidate", etc. in the specification and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects and do not necessarily need to be used to describe a specific order or sequence. It should be understood that such used data can be interchanged under appropriate circumstances so that the embodiments of the present invention described here can be implemented in an order other than those illustrated or described here. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products, or devices.

[0030] In addition, it should also be noted that in the technical solution of the present invention, the processing of collection, storage, use, processing, transmission, provision, and disclosure of the to-be-played video, etc. complies with the provisions of relevant laws and regulations and does not violate public order and good customs.

[0031] Figure 1The flowchart shows a method for processing video data based on artificial intelligence according to an embodiment of the present disclosure. This embodiment is applicable to the situation of how to process video data. This method can be executed by a video data processing device, which can be implemented in software and / or hardware and integrated into an electronic device with video data processing capabilities, such as a mobile terminal. As Figure 1 shown, the method for processing video data based on artificial intelligence in this embodiment may include:

[0032] S101, determine a target video frame corresponding to a target playback moment from the video to be played, use the candidate picture group GOP to which the target video frame belongs as the target GOP, and use the candidate GOP set to which the target GOP belongs as the target GOP set.

[0033] In this embodiment, the video to be played refers to the video that needs to be played, that is, the video that needs to be decoded; further, the video to be played includes at least one candidate group of pictures (GOP) set, that is, the candidate GOP set; the candidate GOP set includes at least one candidate group of pictures (GOP); each candidate GOP includes at least one video frame. A candidate GOP refers to a group of video frames starting with an I frame. A candidate GOP set refers to at least one candidate GOP that can be independently decoded.

[0034] The target playback moment refers to the moment when it is desired to play the video to be played; optionally, the target playback moment can be the initial playback moment of the video to be played; further, it can also be a certain playback moment on the playback progress bar of the video to be played selected by the user.

[0035] The target video frame refers to the video frame in the video to be played corresponding to the target playback moment. The target GOP refers to the candidate GOP to which the target video frame belongs. The target GOP set refers to the candidate GOP set to which the target GOP belongs.

[0036] Specifically, the target playback moment can be determined, and then the target video frame corresponding to the target playback moment is determined from the video to be played, and the candidate GOP to which the target video frame belongs is used as the target GOP, and the candidate GOP set to which the target GOP belongs is used as the target GOP set.

[0037] S102, determine a GOP set to be decoded from at least one candidate GOP set corresponding to the video to be played according to the first decoding state of the target video frame and / or the second decoding state of the target GOP set.

[0038] In this embodiment, the first decoding state refers to the decoding state of the target video frame, which may include undecoded, decoding, and decoded. The second decoding state refers to the decoding state of the target GOP set, which may include undecoded, decoding, and decoded.

[0039] The GOP set to be decoded refers to a candidate GOP set that can be decoded in parallel simultaneously; it should be noted that the number of GOP sets to be decoded can be one or more.

[0040] Optionally, according to the determination rule of the GOP set to be decoded corresponding to the decoding state, based on the first decoding state of the target video frame and / or the second decoding state of the target GOP set, the GOP set to be decoded can be determined from at least one candidate GOP set corresponding to the video to be played.

[0041] S103, perform parallel decoding on the GOP set to be decoded based on the target decoder set to obtain the video frames to be played.

[0042] In this embodiment, the target decoder set refers to a decoder set that can be decoded in parallel simultaneously, including at least two target decoders.

[0043] The video to be played refers to the decoded video frames, which can be used for playing.

[0044] Specifically, the GOP set to be decoded can be input into the target decoders of the target decoder set in sequence for parallel decoding to obtain the video to be played.

[0045] The technical solution provided by the embodiments of the present disclosure determines the target video frame corresponding to the target playback moment from the video to be played, takes the candidate GOP to which the target video frame belongs as the target GOP, and takes the candidate GOP set to which the target GOP belongs as the target GOP set. Then, according to the first decoding state of the target video frame and / or the second decoding state of the target GOP set, the GOP set to be decoded is determined from at least one candidate GOP set corresponding to the video to be played, and then parallel decoding is performed on the GOP set to be decoded based on the target decoder set to obtain the video frames to be played. Compared with the prior art that decodes frame by frame in the time domain dimension, the above technical solution of the present disclosure determines the GOP set to be decoded based on the target video frame corresponding to the target playback moment and the decoding state of the target GOP, making the determination of the GOP set to be decoded more flexible; at the same time, parallel decoding is performed on the GOP set to be decoded by the target decoder set, improving the video decoding efficiency.

[0046] Based on the above embodiment, as an optional manner of the present disclosure, the video frames to be played can also be stored in a cache queue for playback display.

[0047] Specifically, the video frame to be played can be stored in a cache queue so that when playing the video frame to be played corresponding to the target playing moment, it can be quickly obtained from the cache queue, thereby improving the video playing speed and enabling the user to quickly see the expected video.

[0048] Furthermore, the video frames to be played in the cache queue can be deleted and updated according to the memory length of the cache queue, as well as the cache duration and / or playing situation of the video frames to be played. Herein, the cache duration refers to the storage duration of the video frame to be played in the cache queue; the longer the cache duration of the video frame to be played obtained by decoding first.

[0049] An optional method is to delete and update the video frames to be played in the cache queue according to the memory length of the cache queue and the cache duration of the video frames to be played. Specifically, since the memory length of the cache queue is fixed, when the remaining memory length in the cache queue is insufficient, the video frames to be played with a relatively long cache duration can be deleted to ensure that the newly decoded video frames to be played can be stored in the cache queue.

[0050] Another optional method is to delete and update the video frames to be played in the cache queue according to the memory length of the cache queue and the playing situation of the video frames to be played. Specifically, since the memory length of the cache queue is fixed, when the remaining memory length in the cache queue is insufficient, the played video frames to be played at the front of the cache queue can be deleted to ensure that the newly decoded video frames to be played can be stored in the cache queue.

[0051] Another optional method is to delete and update the video frames to be played in the cache queue according to the memory length of the cache queue, as well as the cache duration and playing situation of the video frames to be played. Specifically, since the memory length of the cache queue is fixed, when the remaining memory length in the cache queue is insufficient, the video frames to be played in the cache queue can be deleted and updated by combining the cache duration and playing situation of the video frames to be played.

[0052] It can be understood that storing the video frames to be played in the cache queue can improve the playing efficiency of the video. At the same time, deleting and updating the cache queue according to the queue length of the cache queue, as well as the cache duration and playing situation of the video frames to be played, can ensure that the video frames to be played stored in the cache queue are the newly decoded ones.

[0053] Based on the above embodiments, as an alternative of the present disclosure, determining the GOP set to be decoded from at least one candidate GOP set corresponding to the video to be played according to the second decoding state of the target GOP set may be that if the second decoding state of the target GOP set is decoded, select a first number of GOP sets to be decoded from at least one candidate GOP set corresponding to the video to be played; wherein, the first number is less than or equal to the number of decoders in the target decoder set. Specifically, it may be to determine the first undecoded candidate GOP set from at least one candidate GOP set corresponding to the video to be played, and starting from this candidate GOP set, determine a first number of GOP sets to be decoded from at least one candidate GOP set corresponding to the video to be played.

[0054] Based on the above embodiments, as an alternative of the present disclosure, determining the GOP set to be decoded from at least one candidate GOP set corresponding to the video to be played according to the second decoding state of the target GOP set may be that if the second decoding state of the target GOP set is undecoded, use the target GOP set as the starting GOP set to be decoded, and select a first number of GOP sets to be decoded from at least one candidate GOP set; wherein, the first number is less than or equal to the number of decoders in the target decoder set.

[0055] Specifically, if the second decoding state of the target GOP set is undecoded, use the target GOP set as the starting GOP set to be decoded, and select a first number of GOP sets to be decoded from at least one candidate GOP set. For example, if the video frames to be played correspond to 8 candidate GOP sets denoted as {GOP1, GOP2, GOP3, GOP4, GOP5, GOP6, GOP7, GOP8}, the target GOP set is the 4th candidate GOP set, i.e., GOP4, and the number of decoders in the target decoder set is 4, then GOP4, GOP5, GOP6, and GOP7 are used as the GOP sets to be decoded respectively. Particularly, if the target GOP set is the 6th candidate GOP set, i.e., GOP6, and the number of decoders in the target decoder set is 4, then GOP6 and GOP7 are used as the GOP sets to be decoded respectively.

[0056] It can be understood that when the target GOP set is in the undecoded state, regardless of whether other candidate GOP sets before the target GOP set are being decoded, directly jump to the target GOP set to perform parallel decoding on the target GOP set and the candidate GOP sets after it, so as to quickly obtain the desired video frames to be played.

[0057] Figure 2It is a flowchart of another video data processing method based on artificial intelligence provided according to an embodiment of the present disclosure. On the basis of the above embodiment, this embodiment further optimizes "determining a GOP set to be decoded from at least one candidate GOP set corresponding to a video to be played according to the first decoding state of a target video frame and the second decoding state of a target GOP set" and "performing parallel decoding on the GOP set to be decoded based on a target decoder set to obtain the video to be played", and provides an alternative implementation scheme. As Figure 2 shown, the video data processing method based on artificial intelligence in this embodiment may include:

[0058] S201, determining a target video frame corresponding to a target playing moment from the video to be played, taking the candidate GOP to which the target video frame belongs as the target GOP, and taking the candidate GOP set to which the target GOP belongs as the target GOP set.

[0059] S202, if the first decoding state of the target video frame is undecoded and the second decoding state of the target GOP set is being decoded, determining whether there is a GOP set before the target GOP set from other GOP sets decoded in parallel with the target GOP set.

[0060] Specifically, if the first decoding state of the target video frame is undecoded and the second decoding state of the target GOP set is being decoded, that is, the target GOP set is being decoded but the target video frame has not been decoded yet, determining whether there is a GOP set before the target GOP set from other GOP sets decoded in parallel with the target GOP set.

[0061] S203, if there is, taking this GOP set as the GOP set to be released.

[0062] Specifically, if there is a GOP set before the target GOP set, taking this GOP set as the GOP set to be released.

[0063] S204, based on the GOP set to be released, selecting a second number of candidate decoded GOP sets from at least one candidate GOP set corresponding to the video to be played.

[0064] Specifically, based on the number of the GOP sets to be released, selecting a second number of candidate GOP sets with the second decoding state being undecoded from at least one candidate GOP set corresponding to the video to be played as the candidate decoded GOP sets. Wherein the second number is less than or equal to the number of the GOP sets to be released.

[0065] S205, taking the candidate decoded GOP sets and other GOP sets after the target GOP set decoded in parallel with the target GOP set as the GOP sets to be decoded.

[0066] Specifically, the candidate decoded GOP set, as well as other GOP sets after the target GOP set that are decoded in parallel with the target GOP set, can be used as the GOP sets to be decoded. Delay

[0067] S206. Determine the target decoder corresponding to the GOP set to be released from the target decoder set as the decoder to be released.

[0068] Specifically, determine the target decoder corresponding to the GOP set to be released from the target decoder set as the decoder to be released.

[0069] S207. Use the decoder to be released to perform parallel decoding on the candidate decoded GOP set.

[0070] Specifically, use the decoder to be released to perform parallel decoding on the candidate decoded GOP set.

[0071] S208. Use the target decoders in the target decoder set other than the decoder to be released to continue performing parallel decoding on the target GOP set and other GOP sets after the target GOP set that are decoded in parallel with the target GOP set.

[0072] Meanwhile, use the target decoders in the target decoder set other than the decoder to be released to continue performing parallel decoding on the target GOP set and other GOP sets after the target GOP set that are decoded in parallel with the target GOP set.

[0073] It should be noted that S207 and S208 are performed simultaneously.

[0074] As a specific example, the 8 candidate GOP sets corresponding to the video frames to be played are recorded as {GOP1, GOP2, GOP3, GOP4, GOP5, GOP6, GOP7, GOP8}, the target GOP set is the 4th candidate GOP set, namely GOP4, and the number of decoders in the target decoder set is 4, namely decoder1, decoder2, decoder3 and decoder4, which decode GOP2, GOP3, GOP4 and GOP5 in parallel respectively. Then GOP2 and GOP3 are used as releasable GOP sets, and the releasable decoders corresponding to the releasable GOP sets are decoder1 and decoder2; at this time, the undecoded candidate GOP sets are GOP6, GOP7, GOP8, so GOP6 and GOP7 are selected as candidate decoding GOP sets, and the GOP set to be decoded is {GOP4, GOP5, GOP6, GOP7}. Decoder1 and decoder2 are used to decode GOP6 and GOP7 in parallel, and decoder3 and decoder4 are used to decode GOP4 and GOP5 in parallel.

[0075] In another specific example, if the video frame to be played corresponds to 8 candidate GOP sets recorded as {GOP1, GOP2, GOP3, GOP4, GOP5, GOP6}, and the target GOP set is the fourth candidate GOP set, namely GOP4, and the number of decoders in the target decoder set is 4, namely decoder1, decoder2, decoder3 and decoder4, respectively, which decode GOP2, GOP3, GOP4 and GOP5 in parallel, then GOP2 and GOP3 are used as releasable GOP sets, and the releasable decoders corresponding to the releasable GOP sets are decoder1 and decoder2; at this time, the undecoded candidate GOP set is only GOP6, so GOP6 is selected as the candidate decoding GOP set, and the GOP set to be decoded is {GOP4, GOP5, GOP6}. Decoder1 is used to decode GOP6, and decoder3 and decoder4 are continued to decode GOP4 and GOP5 in parallel.

[0076] In the technical solution of the embodiment of the present disclosure, the target video frame corresponding to the target playback moment is determined from the video to be played, and the candidate GOP to which the target video frame belongs is used as the target GOP, and the candidate GOP set to which the target GOP belongs is used as the target GOP set. Then, if the first decoding state of the target video frame is undecoded and the second decoding state of the target GOP set is in decoding, it is determined whether there is a GOP set before the target GOP set among the other GOP sets decoded in parallel with the target GOP set. If so, this GOP set is used as the GOP set to be released; based on the GOP set to be released, the second number of candidate decoded GOP sets are selected from at least one candidate GOP set corresponding to the video to be played, and the candidate decoded GOP sets and the other GOP sets after the target GOP set decoded in parallel with the target GOP set are used as the GOP sets to be decoded. Further, the target decoder corresponding to the GOP set to be released is determined from the target decoder set as the decoder to be released, and the decoder to be released is used to perform parallel decoding on the candidate decoded GOP sets, and at the same time, the target decoders other than the decoder to be released in the target decoder set are used to continue to perform parallel decoding on the target GOP set and the other GOP sets after the target GOP set decoded in parallel with the target GOP set. In the above technical solution, when there is a candidate GOP set decoded in parallel before the target GOP set, that is, the GOP set to be released, the target decoder corresponding to the GOP set to be released is released to continue decoding the subsequent undecoded candidate GOP sets, which can make full use of the target decoder and improve the decoding efficiency at the same time.

[0077] Figure 3 It is a flowchart of another video data processing method based on artificial intelligence provided according to an embodiment of the present disclosure. On the basis of the above embodiment, this embodiment further elaborates on the determination of the candidate GOP set. As Figure 3 shown, the video data processing method based on artificial intelligence in this embodiment may include:

[0078] S301, determine at least two candidate GOPs in the video to be played and the dependency relationship between the at least two candidate GOPs.

[0079] In this embodiment, the dependency relationship refers to whether there is a dependency relationship between video frames in two adjacent candidate GOPs during the decoding process. For example, if the first frame of a candidate GOP is an IDR frame, then the candidate GOP can be decoded independently, that is, the candidate GOP does not depend on other candidate GOPs before or after it during the decoding process; if the first frame of a candidate GOP is an I frame, and the B frames in the candidate GOP need to depend on the decoded images in the previous candidate GOP of the candidate GOP during the decoding process, then the candidate GOP is an open-GOP, that is, the candidate GOP cannot be decoded independently and needs to depend on the previous candidate GOP during the decoding process.

[0080] Specifically, the video to be played is parsed to obtain at least two candidate GOPs of the video to be played and the dependency relationship between the at least two candidate GOPs.

[0081] S302. According to the dependency relationship between the at least two candidate GOPs, the at least two candidate GOPs are grouped to obtain at least one candidate GOP set.

[0082] Specifically, according to the dependency relationship between the at least two candidate GOPs, the at least two candidate GOPs can be grouped to obtain at least one candidate GOP set. For example, the video to be played includes three candidate GOPs, namely IDR-GOP1, IDR-GOP2, and IDR-GOP3, and these three candidate GOPs can all be decoded independently, then each candidate GOP is a candidate GOP set.

[0083] Another example is that the video to be played includes three candidate GOPs, namely IDR-GOP1, IDR-GOP2, and I-GOP3, and I-GOP3 is an open-GOP, then the number of candidate GOP sets is two, which are {IDR-GOP1} and {IDR-GOP2, I-GOP3} respectively.

[0084] S303. Determine the target video frame corresponding to the target playback moment from the video to be played, and use the candidate GOP to which the target video frame belongs as the target GOP, and use the candidate GOP set to which the target GOP belongs as the target GOP set.

[0085] S304. According to the first decoding state of the target video frame and / or the second decoding state of the target GOP set, determine the GOP set to be decoded from the at least one candidate GOP set corresponding to the video to be played.

[0086] S305. Based on the target decoder set, perform parallel decoding on the GOP set to be decoded to obtain the video frames to be played.

[0087] In the technical solution of the embodiment of the present disclosure, at least two candidate GOPs in the video to be played are determined, as well as the dependency relationship between at least two candidate GOPs. Then, according to the dependency relationship between at least two candidate GOPs, at least two candidate GOPs are grouped to obtain at least one candidate GOP set. After that, the target video frame corresponding to the target playback moment is determined from the video to be played, and the candidate GOP to which the target video frame belongs is used as the target GOP, and the candidate GOP set to which the target GOP belongs is used as the target GOP set. Furthermore, according to the first decoding state of the target video frame and / or the second decoding state of the target GOP set, the GOP set to be decoded is determined from at least one candidate GOP set corresponding to the video to be played, and the GOP set to be decoded is decoded in parallel based on the target decoder set to obtain the video frames to be played. In the above technical solution, grouping is performed according to the dependency relationship between candidate GOP sets to determine the candidate GOP set, which provides guarantee for subsequent parallel decoding of the video.

[0088] Figure 4 It is a flowchart of another video data processing method based on artificial intelligence provided by an embodiment of the present disclosure. On the basis of the above embodiment, this embodiment further elaborates on the determination process of the target decoder set. As Figure 4 shown, the video data processing method based on artificial intelligence in this embodiment may include:

[0089] S401, determine the number of parallel decodings according to the kernel resources of the local device.

[0090] Among them, the kernel resources include at least one of the following: the number of CPU cores, the compression ratio, the resolution, and the number of GPU cards.

[0091] An optional method is to determine the number of parallel decodings according to the number of CPU cores, the compression ratio, and the resolution. For example, when the compression ratio is less than or equal to a set value (such as H.264), the number of parallel decodings is determined according to the resolution and the number of CPU cores. Specifically, when the resolution is less than or equal to 360P, the number of CPU cores is used as the number of parallel decodings; when the resolution is less than or equal to 720P, the result of dividing the number of cores by 2 is used as the number of parallel decodings; when the resolution is less than or equal to 1080P, the result of dividing the number of cores by 3 is used as the number of parallel decodings; in other cases of resolution, the result of dividing the number of cores by 4 is used as the number of parallel decodings.

[0092] For another example, when the compression ratio is greater than a set value (such as H.264), the number of parallel decodings is determined according to the resolution and the number of CPU cores. Specifically, when the resolution is less than or equal to 360P, the number of CPU cores is used as the number of parallel decodings; when the resolution is less than or equal to 720P, the result of dividing the number of cores by 3 is used as the number of parallel decodings; in the case of other resolutions, the result of dividing the number of cores by 4 is used as the number of parallel decodings.

[0093] Another optional method is that the number of parallel decodings can be determined according to the number of GPU cards. Specifically, the number of GPU cards can be used as the number of parallel decodings.

[0094] Another optional method is that the number of parallel decodings can also be customized according to the actual needs of the user.

[0095] S402. Create at least two target decoders according to the number of parallel decodings to obtain a set of target decoders.

[0096] Specifically, the number of parallel decodings of target decoders can be created according to the number of parallel decodings to obtain a set of target decoders. It should be noted that the target decoder in the present disclosure is not specifically limited and can be any decoder capable of video decoding.

[0097] S403. Determine the target video frame corresponding to the target playback moment from the video to be played, and use the candidate GOP to which the target video frame belongs as the target GOP, and use the set of candidate GOPs to which the target GOP belongs as the set of target GOPs.

[0098] S404. Determine the set of GOPs to be decoded from at least one set of candidate GOPs corresponding to the video to be played according to the first decoding state of the target video frame and / or the second decoding state of the set of target GOPs.

[0099] S405. Perform parallel decoding on the set of GOPs to be decoded based on the set of target decoders to obtain the video frames to be played.

[0100] In the technical solution of the embodiment of the present disclosure, the number of parallel decoders is determined according to the kernel resources of the local device, and at least two target decoders are created according to the number of parallel decoders to obtain a set of target decoders. Then, the target video frame corresponding to the target playback moment is determined from the video to be played, and the candidate GOP to which the target video frame belongs is used as the target GOP, and the set of candidate GOPs to which the target GOP belongs is used as the target GOP set. Furthermore, according to the first decoding state of the target video frame and / or the second decoding state of the target GOP set, the GOP set to be decoded is determined from at least one set of candidate GOPs corresponding to the video to be played, and the GOP set to be decoded is decoded in parallel based on the set of target decoders to obtain the video frames to be played. In the above technical solution, the number of parallel decoders is determined by local resources to determine the set of target decoders, so that the local resources are fully utilized, thus laying a foundation for improving the decoding efficiency of the video.

[0101] Figure 5 It is a schematic structural diagram of a video data processing device based on artificial intelligence provided by an embodiment of the present disclosure. This embodiment is applicable to the situation of how to process video data. The device can be implemented in software and / or hardware, and can be integrated into an electronic device with video data processing functions, such as a mobile terminal. As Figure 5 shown, the video data processing device 500 based on artificial intelligence in this embodiment may include:

[0102] A target GOP set determination module 501, configured to determine a target video frame corresponding to a target playback moment from the video to be played, use the candidate group of pictures (GOP) to which the target video frame belongs as the target GOP, and use the set of candidate GOPs to which the target GOP belongs as the target GOP set;

[0103] A GOP set to be decoded determination module 502, configured to determine the GOP set to be decoded from at least one set of candidate GOPs corresponding to the video to be played according to the first decoding state of the target video frame and / or the second decoding state of the target GOP set;

[0104] A video frame to be played determination module 503, configured to perform parallel decoding on the GOP set to be decoded based on the set of target decoders to obtain the video frames to be played.

[0105] The technical solution provided by the embodiments of the present disclosure determines a target video frame corresponding to a target playback moment from a video to be played, uses the candidate GOP to which the target video frame belongs as the target GOP, and uses the set of candidate GOPs to which the target GOP belongs as the target GOP set. Then, according to the first decoding state of the target video frame and / or the second decoding state of the target GOP set, a GOP set to be decoded is determined from at least one set of candidate GOPs corresponding to the video to be played. Furthermore, the GOP set to be decoded is decoded in parallel based on the target decoder set to obtain the video frames to be played. Compared with the prior art solution of decoding frame by frame in the time domain, the present disclosure determines the GOP set to be decoded based on the target video frame corresponding to the target playback moment and the decoding state of the target GOP, making the determination of the GOP set to be decoded more flexible. At the same time, by decoding the GOP set to be decoded in parallel using the target decoder set, the video decoding efficiency is improved.

[0106] Further, the GOP set to be decoded determination module 502 is configured to:

[0107] If the second decoding state of the target GOP set is undecoded, use the target GOP set as the starting GOP set to be decoded, and select a first number of GOP sets to be decoded from at least one set of candidate GOPs;

[0108] Wherein, the first number is less than or equal to the number of decoders in the target decoder set.

[0109] Further, the GOP set to be decoded determination module 502 is further configured to:

[0110] If the first decoding state of the target video frame is undecoded and the second decoding state of the target GOP set is being decoded, determine whether there is a GOP set before the target GOP set among the other GOP sets decoded in parallel with the target GOP set;

[0111] If there is, use this GOP set as the GOP set that can be released;

[0112] Based on the GOP set that can be released, select a second number of candidate decoded GOP sets from at least one set of candidate GOPs corresponding to the video to be played; wherein the second number is less than or equal to the number of GOP sets that can be released;

[0113] Use the candidate decoded GOP sets and the other GOP sets after the target GOP set decoded in parallel with the target GOP set as the GOP sets to be decoded.

[0114] Further, the video to be played determination module 503 is specifically configured to:

[0115] Determine, from the set of target decoders, the target decoder corresponding to the GOP set that can be released as the releaseable decoder;

[0116] Use the releaseable decoder to perform parallel decoding on the candidate decoded GOP set;

[0117] Meanwhile, use the target decoders in the set of target decoders other than the releaseable decoder to continue parallel decoding on the target GOP set and other GOP sets located after the target GOP set that are decoded in parallel with the target GOP set.

[0118] Furthermore, the apparatus further includes:

[0119] A storage module, configured to store video frames to be played into a cache queue for playback display.

[0120] Furthermore, the apparatus further includes:

[0121] An update module, configured to delete and update the video frames to be played in the cache queue according to the memory length of the cache queue, and the cache duration and / or playback status of the video frames to be played.

[0122] Furthermore, the apparatus further includes a candidate GOP set determination module, configured to:

[0123] Before determining the target playback moment of the video to be played, determine at least two candidate GOPs in the video to be played and the dependency relationship between the at least two candidate GOPs;

[0124] Group the at least two candidate GOPs according to the dependency relationship between the at least two candidate GOPs to obtain at least one candidate GOP set.

[0125] Furthermore, the apparatus further includes a target decoder set determination module, configured to:

[0126] Before determining the target video frame corresponding to the target playback moment from the video to be played, determine the number of parallel decodings according to the kernel resources of the local device; where the kernel resources include at least one of the following: the number of CPU cores, the compression ratio, the resolution, and the number of GPU cards;

[0127] Create at least two target decoders according to the number of parallel decodings to obtain a set of target decoders.

[0128] According to an embodiment of the present disclosure, the present disclosure further provides an electronic device, a readable storage medium, and a computer program product.

[0129] Figure 6 It is a block diagram of an electronic device for implementing the video data processing method according to the embodiments of the present disclosure. Figure 6FIG. shows a schematic block diagram of an exemplary electronic device 600 that may be used to implement embodiments of the present disclosure. The electronic device is intended to represent various forms of digital computers, such as, a laptop computer, a desktop computer, a workbench, a personal digital assistant, a server, a blade server, a mainframe computer, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as, a personal digital processor, a cellular phone, a smart phone, a wearable device, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are only examples and are not intended to limit the implementation of the present disclosure described and / or claimed herein.

[0130] As Figure 6 shown, the electronic device 600 includes a computing unit 601, which may perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 602 or a computer program loaded from a storage unit 608 into a random access memory (RAM) 603. In the RAM 603, various programs and data required for the operation of the electronic device 600 may also be stored. The computing unit 601, the ROM 602, and the RAM 603 are connected to each other via a bus 604. An input / output (I / O) interface 605 is also connected to the bus 604.

[0131] A plurality of components in the electronic device 600 are connected to the I / O interface 605, including: an input unit 606, such as a keyboard, a mouse, etc.; an output unit 607, such as various types of displays, speakers, etc.; a storage unit 608, such as a magnetic disk, an optical disk, etc.; and a communication unit 609, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 609 allows the electronic device 600 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.

[0132] The computing unit 601 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 601 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 601 executes the various methods and processes described above, such as the artificial intelligence-based video data processing method. For example, in some embodiments, the artificial intelligence-based video data processing method can be implemented as a computer software program tangibly embodied in a machine-readable medium, such as the storage unit 608. In some embodiments, part or all of the computer program can be loaded and / or installed onto the electronic device 600 via the ROM 602 and / or the communication unit 609. When the computer program is loaded into the RAM 603 and executed by the computing unit 601, one or more steps of the artificial intelligence-based video data processing method described above can be executed. Alternatively, in other embodiments, the computing unit 601 can be configured to execute the artificial intelligence-based video data processing method by any other suitable means (e.g., by means of firmware).

[0133] The various embodiments of the systems and techniques described above in this document can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-chip (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a dedicated or general-purpose programmable processor, receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting the data and instructions to the storage system, the at least one input device, and the at least one output device.

[0134] The program code for implementing the methods of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when the program codes are executed by the processor or controller, the functions / operations specified in the flowcharts and / or block diagrams are implemented. The program codes can be executed entirely on the machine, partially on the machine, as an independent software package partially on the machine and partially on a remote machine, or entirely on a remote machine or server.

[0135] In the context of this disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0136] To provide for interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices can also be used to provide for interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic, speech, or tactile input).

[0137] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer having a graphical user interface or a web browser through which the user can interact with an implementation of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), and the Internet.

[0138] A computer system may include a client and a server. The client and the server are generally far from each other and usually interact through a communication network. The client-server relationship is generated by computer programs running on corresponding computers and having a client-server relationship with each other. The server may be a cloud server, a server of a distributed system, or a server combined with a blockchain.

[0139] Artificial intelligence is a discipline that studies the simulation of certain human thinking processes and intelligent behaviors (such as learning, reasoning, thinking, planning, etc.) by computers, including both hardware-level technologies and software-level technologies. Artificial intelligence hardware technologies generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, and big data processing; artificial intelligence software technologies mainly include several major directions such as computer vision technology, speech recognition technology, natural language processing technology, machine learning / deep learning technology, big data processing technology, and knowledge graph technology.

[0140] Cloud computing refers to a technical system that accesses an elastic and scalable shared physical or virtual resource pool through a network. The resources can include servers, operating systems, networks, software, applications, and storage devices, etc., and the resources can be deployed and managed in a on-demand and self-service manner. Through cloud computing technology, it can provide efficient and powerful data processing capabilities for the application and model training of technologies such as artificial intelligence and blockchain.

[0141] It should be understood that various forms of processes shown above can be used, steps can be reordered, added, or deleted. For example, the steps recorded in this disclosure can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved, and no limitations are imposed herein.

[0142] The above specific implementation manners do not constitute a limitation on the protection scope of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure shall be included within the protection scope of this disclosure.

Claims

1. A video data processing method based on artificial intelligence, comprising: Determining a target video frame corresponding to a target playback moment from a video to be played, taking the candidate picture group GOP to which the target video frame belongs as the target GOP, and taking the candidate GOP set to which the target GOP belongs as the target GOP set; Determining a GOP set to be decoded from at least one candidate GOP set corresponding to the video to be played according to the first decoding state of the target video frame and the second decoding state of the target GOP set, or the second decoding state of the target GOP set; Performing parallel decoding on the GOP set to be decoded based on a target decoder set to obtain video frames to be played; Among them, determining the GOP set to be decoded from at least one candidate GOP set corresponding to the video to be played according to the first decoding state of the target video frame and the second decoding state of the target GOP set includes: If the first decoding state of the target video frame is undecoded and the second decoding state of the target GOP set is being decoded, determining whether there is a GOP set before the target GOP set among other GOP sets decoded in parallel with the target GOP set; If so, taking this GOP set as the GOP set to be released; Based on the GOP set to be released, selecting a second number of candidate decoded GOP sets from at least one candidate GOP set corresponding to the video to be played; where the second number is less than or equal to the number of the GOP sets to be released; Taking the candidate decoded GOP sets and other GOP sets after the target GOP set decoded in parallel with the target GOP set as the GOP set to be decoded.

2. The method according to claim 1, wherein Determining the GOP set to be decoded from at least one candidate GOP set corresponding to the video to be played according to the second decoding state of the target GOP set includes: If the second decoding state of the target GOP set is undecoded, taking the target GOP set as the starting GOP set to be decoded and selecting a first number of GOP sets to be decoded from the at least one candidate GOP set; Wherein, the first number is less than or equal to the number of decoders in the target decoder set.

3. The method according to claim 1, wherein The performing parallel decoding on the GOP set to be decoded based on a target decoder set to obtain the video to be played includes: Determining, from the target decoder set, the target decoder corresponding to the GOP set to be released as the decoder to be released; Using the decoder to be released to perform parallel decoding on the candidate decoded GOP sets; Meanwhile, using the target decoders in the target decoder set other than the decoder to be released to continue performing parallel decoding on the target GOP set and other GOP sets after the target GOP set decoded in parallel with the target GOP set.

4. The method according to claim 1, wherein, It further includes: Storing the video frames to be played into a buffer queue for playback display.

5. The method according to claim 4, wherein It further includes: Delete and update the video frames to be played in the cache queue according to the memory length of the cache queue, and the cache duration and / or playing status of the video frames to be played.

6. The method according to claim 1, wherein Before determining the target playing moment of the video to be played, it further includes: Determine at least two candidate GOPs in the video to be played, and the dependency relationship between the at least two candidate GOPs; Group the at least two candidate GOPs according to the dependency relationship between the at least two candidate GOPs to obtain at least one candidate GOP set.

7. The method according to claim 1, wherein, Before determining the target video frame corresponding to the target playing moment from the video to be played, it further includes: Determine the parallel decoding number according to the kernel resources of the local device; wherein, the kernel resources include at least one of the following: the number of CPU cores, compression ratio, resolution, and the number of GPU cards; Create at least two target decoders according to the parallel decoding number to obtain a target decoder set.

8. A video data processing device based on artificial intelligence, including: A target GOP set determination module, configured to determine a target video frame corresponding to a target playing moment from a video to be played, and use the candidate picture group GOP to which the target video frame belongs as the target GOP, and use the candidate GOP set to which the target GOP belongs as the target GOP set; A GOP set to be decoded determination module, configured to determine a GOP set to be decoded from at least one candidate GOP set corresponding to the video to be played according to the first decoding status of the target video frame and the second decoding status of the target GOP set, or the second decoding status of the target GOP set; A video to be played determination module, configured to perform parallel decoding on the GOP set to be decoded based on the target decoder set to obtain video frames to be played; Wherein, the GOP set to be decoded determination module is further configured to: If the first decoding status of the target video frame is undecoded and the second decoding status of the target GOP set is in decoding, determine whether there is a GOP set before the target GOP set among other GOP sets decoded in parallel with the target GOP set; If so, use this GOP set as the GOP set to be released; Based on the GOP set to be released, select a second number of candidate decoded GOP sets from at least one candidate GOP set corresponding to the video to be played; wherein the second number is less than or equal to the number of the GOP sets to be released; Use the candidate decoded GOP sets and other GOP sets after the target GOP set decoded in parallel with the target GOP set as the GOP sets to be decoded.

9. The device according to claim 8, wherein, The GOP set to be decoded determination module is configured to: If the second decoding status of the target GOP set is undecoded, use the target GOP set as the starting GOP set to be decoded, and select a first number of GOP sets to be decoded from the at least one candidate GOP set; Wherein, the first number is less than or equal to the number of decoders in the target decoder set.

10. The apparatus according to claim 8, wherein, The video to be played determination module is specifically configured to: From the set of target decoders, determine the target decoder corresponding to the set of releasable GOPs as the releasable decoder; Use the releasable decoder to perform parallel decoding on the set of candidate decoded GOPs; Meanwhile, use the target decoders in the set of target decoders other than the releasable decoder to continue to perform parallel decoding on the target GOP set and other GOP sets located after the target GOP set and decoded in parallel with the target GOP set.

11. The apparatus according to claim 8, wherein, Further includes: A storage module for storing the video frames to be played into a cache queue for playback display.

12. The apparatus according to claim 11, wherein Further includes: An update module for deleting and updating the video frames to be played in the cache queue according to the memory length of the cache queue, and the cache duration and / or playback status of the video frames to be played.

13. The apparatus according to claim 8, wherein Further includes a candidate GOP set determination module for: Before determining the target playback moment of the video to be played, determine at least two candidate GOPs in the video to be played and the dependency relationship between the at least two candidate GOPs; Group the at least two candidate GOPs according to the dependency relationship between the at least two candidate GOPs to obtain at least one candidate GOP set.

14. The apparatus according to claim 8, wherein Further includes a target decoder set determination module for: Before determining the target video frame corresponding to the target playback moment from the video to be played, determine the number of parallel decodings according to the kernel resources of the local device; wherein the kernel resources include at least one of the following: the number of CPU cores, the compression ratio, the resolution, and the number of GPU cards; Create at least two target decoders according to the number of parallel decodings to obtain a set of target decoders.

15. An electronic device, comprising: At least one processor; And A memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the video data processing method according to any one of claims 1-7.

16. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are used to cause a computer to execute the video data processing method according to any one of claims 1-7.

17. A computer program product, comprising a computer program which, when executed by a processor, implements the video data processing method according to any one of claims 1-7.

Citation Information

Patent Citations

  • Method and device for video reverse playback

    CN106507204A

  • Video editing method, device and equipment and computer readable storage medium

    CN113015005A