A video data processing method, apparatus, device, and system

By selecting a sleep state decoder in the decoding side device for video data processing, the problem of high resource consumption of multiple decoder is solved, and the saving of memory and decoding resources and the improvement of processing efficiency is achieved.

CN115297331BActive Publication Date: 2025-07-22HANGZHOU HIKVISION DIGITAL TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210910163.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-07-29
Publication Date
2025-07-22
Estimated Expiration
2042-07-29

AI Technical Summary

Technical Problem

In the prior art, the decoding side device needs to support multiple decoders, resulting in excessive consumption of memory resources and decoding resources, and it is impossible to efficiently perform specified video data processing.

Method used

By selecting a decoder in the sleep state in the decoding device, the target code stream is subjected to specified video data processing, and the decoding process of the decoder is restored after the processing is completed, and the memory resources and decoding resources of the N-channel decoder are multiplexed.

Benefits of technology

It effectively saves the consumption of memory resources and decoding resources, improves processing efficiency, and realizes the processing of specified video data without the need for additional resource consumption.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115297331B_ABST
    Figure CN115297331B_ABST
Patent Text Reader

Abstract

The present application provides a video data processing method, apparatus, device and system. The method includes: determining a target bitstream that needs to be subjected to specified video data processing; selecting a target decoder in a dormant state from N decoders; wherein, the target decoder has completed decoding of all frames within the current GOP sequence and the target decoder has not started decoding the next GOP sequence of the current GOP sequence; performing specified video data processing on the target bitstream through the target decoder; after the specified video data processing is completed, decoding the bitstream corresponding to the target decoder through the target decoder, and the bitstream includes the next GOP sequence. Through the technical solution of the present application, the memory resources and decoding resources of the N decoders can be reused, the memory resources and decoding resources can be effectively saved, and the processing efficiency can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of image processing technologies, and in particular, to a video data processing method, apparatus, device, and system. Background Art

[0002] For the purpose of saving space, video images are transmitted after being encoded. A complete video encoding may include processes such as prediction, transformation, quantization, entropy encoding, and filtering. Therefore, after a camera captures a video image, it encodes the video image to obtain an encoded bitstream and sends the encoded bitstream to a decoding-side device. By encoding the video image, data compression of the video image can be achieved, which is convenient for storage and network transmission. The encoding format can be H264, H265, MPEG4, etc.

[0003] After the decoding-side device receives the encoded bitstream, it can decode the bitstream through a decoder to obtain a video image, that is, restore the encoded bitstream to a video image.

[0004] The decoding-side device is connected to N cameras, that is, it decodes the bitstreams corresponding to N cameras simultaneously. Therefore, the decoding-side device needs to support N decoders, and uses the N decoders to decode the bitstreams corresponding to N cameras. In addition, the decoding-side device also needs to reserve M decoders to perform specified video data processing on the bitstream. Obviously, the decoding-side device needs to support at least N + M decoders.

[0005] For each decoder, it needs to occupy independent memory resources and independent decoding resources. Therefore, when the decoding-side device needs to support N + M decoders, it needs to reserve N + M memory resources and N + M decoding resources, thus consuming relatively large memory resources and decoding resources. Summary of the Invention

[0006] This application provides a video data processing method. The decoding-side device includes N decoders for decoding N bitstreams. The method includes:

[0007] Determine a target bitstream that needs to perform specified video data processing;

[0008] Select a target decoder in a dormant state from the N decoders; wherein, the target decoder has completed the decoding of all frames within the current GOP sequence and has not started to decode the next GOP sequence of the current GOP sequence;

[0009] Perform specified video data processing on the target bitstream through the target decoder;

[0010] After the specified video data processing is completed, the target decoder decodes the bitstream corresponding to the target decoder, and the bitstream includes the subsequent GOP sequence.

[0011] This application provides a video data processing device. The decoding-side device includes N decoders for decoding N bitstreams. The device includes:

[0012] A determination module, configured to determine a target bitstream that needs to undergo specified video data processing;

[0013] A selection module, configured to select a target decoder in a dormant state from the N decoders; wherein, the target decoder has completed decoding of all frames within the current GOP sequence and has not started decoding the subsequent GOP sequence of the current GOP sequence;

[0014] A processing module, configured to perform specified video data processing on the target bitstream through the target decoder;

[0015] A decoding module, configured to, after the specified video data processing is completed, decode the bitstream corresponding to the target decoder through the target decoder, and the bitstream includes the subsequent GOP sequence.

[0016] This application provides a decoding-side device, including: a processor and a machine-readable storage medium. The machine-readable storage medium stores machine-executable instructions that can be executed by the processor; the processor is configured to execute the machine-executable instructions to implement the video data processing method in the above example.

[0017] This application provides a video data processing system, including N cameras and a decoding-side device. The N cameras are configured to generate N bitstreams, and the decoding-side device includes N decoders for decoding the N bitstreams, wherein:

[0018] For each camera, the camera is configured to collect the current GOP sequence and the subsequent GOP sequence of the current GOP sequence, generate a bitstream corresponding to the camera, the bitstream including the current GOP sequence and the subsequent GOP sequence, and send the bitstream to the decoder corresponding to the camera; the decoder is configured to decode the bitstream after receiving the bitstream;

[0019] The decoding-side device is configured to determine a target bitstream that needs to undergo specified video data processing; select a target decoder in a dormant state from the N decoders, and send the target bitstream to the target decoder; wherein, the target decoder has completed decoding of all frames within the current GOP sequence and has not started decoding the subsequent GOP sequence of the current GOP sequence;

[0020] The target decoder is used to perform specified video data processing on the target bitstream; after the specified video data processing is completed, continue to decode the bitstream corresponding to the target decoder.

[0021] As can be seen from the above technical solutions, in the embodiments of the present application, the decoding-side device only needs to support N decoders, decode the bitstreams corresponding to N cameras through the N decoders, and perform specified video data processing on the bitstreams through the N decoders, so as to be able to reuse the memory resources and decoding resources of the N decoders, without the need to additionally support M decoders. In this way, the memory resources and decoding resources occupied by the M decoders can be saved, and the memory resources and decoding resources of N+M channels can be reduced to the memory resources and decoding resources of N channels, thereby reducing the consumption of memory resources and decoding resources, achieving the ability to perform specified video data processing without additional resource consumption, effectively saving memory resources and decoding resources, and improving the processing efficiency. Description of the Drawings

[0022] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments of the present application or the prior art. Obviously, the drawings in the following description are only some embodiments recorded in the present application. For those of ordinary skill in the art, other drawings can also be obtained according to these drawings in the embodiments of the present application.

[0023] Figure 1 It is a schematic flowchart of a video data processing method in an embodiment of the present application;

[0024] Figures 2A - 2C It is a schematic structural diagram of a video data processing system in an embodiment of the present application;

[0025] Figure 3 It is a schematic flowchart of the decoding process of the bitstream in an embodiment of the present application;

[0026] Figure 4 It is a schematic diagram of the division of I frames and P frames in an embodiment of the present application;

[0027] Figure 5 It is a schematic flowchart of a video data processing method in an embodiment of the present application;

[0028] Figure 6 It is a schematic flowchart of a video data processing method in an embodiment of the present application;

[0029] Figure 7 It is a schematic structural diagram of a video data processing device in an embodiment of the present application;

[0030] Figure 8 It is a hardware structure diagram of a decoding-side device in an embodiment of the present application. Specific Embodiments

[0031] The terms used in the embodiments of the present application are only for the purpose of describing specific embodiments and do not limit the present application. The singular forms "a", "the", and "said" used in the present application and the claims are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term "and / or" used herein refers to any or all possible combinations of one or more of the associated listed items.

[0032] It should be understood that although the terms first, second, third, etc. may be used in the embodiments of the present application to describe various information, such information should not be limited to these terms. These terms are only used to distinguish the same type of information from each other. For example, without departing from the scope of the present application, the first information may also be referred to as the second information, and similarly, the second information may also be referred to as the first information. Depending on the context, in addition, the word "if" used may be interpreted as "when" or "while" or "in response to a determination".

[0033] In an embodiment of the present application, a video data processing method is proposed, which can be applied to a decoding-side device. The decoding-side device includes N decoders for decoding N streams. The N decoders correspond to the N streams one by one. Refer to Figure 1 As shown, it is a flowchart of the method. The method may include:

[0034] Step 101, determine the target stream that needs to perform specified video data processing.

[0035] Step 102, select a target decoder in a dormant state from the N decoders; wherein, the target decoder has completed the decoding of all frames within the current GOP (Group Of Picture) sequence and has not started to decode the next GOP sequence of the current GOP sequence.

[0036] Exemplarily, it can be determined whether there is a decoder in a dormant state among the N decoders. If so, select a target decoder in a dormant state from the N decoders; if not, after waiting for a preset duration, return to perform the operation of determining whether there is a decoder in a dormant state among the N decoders until there is a decoder in a dormant state, and then select a target decoder in a dormant state from the N decoders.

[0037] In a possible implementation, selecting a target decoder in a dormant state from N decoders may include, but is not limited to: for each decoder, the code stream corresponding to the decoder may include the current GOP sequence and the next GOP sequence after the current GOP sequence. The current GOP sequence may include a first I-frame and multiple first P-frames, and the next GOP sequence may include a second I-frame and multiple second P-frames. If the decoder has completed the decoding of the first I-frame, completed the decoding of multiple first P-frames, and has not started decoding the second I-frame, then the decoder is selected as the target decoder in a dormant state.

[0038] Step 103: Perform specified video data processing on the target code stream through the target decoder.

[0039] In a possible implementation, the specified video data processing may include, but is not limited to, screenshot processing. Based on this, the I-frame in the target code stream can be input to the target decoder; the target decoder decodes the I-frame to obtain a first image in a first image format, converts the first image to a second image in a second image format, and outputs the second image. Exemplarily, the first image format may include, but is not limited to, the YUV format, and the second image format may include, but is not limited to, the JEPG format. Of course, the YUV format and the JEPG format are only examples of the first image format and the second image format, and there is no limitation on this image format.

[0040] Step 104: After the specified video data processing is completed, decode the code stream corresponding to the target decoder through the target decoder. The code stream may include the next GOP sequence after the current GOP sequence.

[0041] Exemplarily, before selecting a target decoder in a dormant state from N decoders, it is also possible to determine whether there are candidate decoders in an idle state among the N decoders; among them, the code stream corresponding to the candidate decoder does not include the GOP sequence to be processed. If so, the candidate decoder can be used to perform specified video data processing on the target code stream; if not, the operation of selecting a target decoder in a dormant state from N decoders is executed, and the target decoder is used to perform specified video data processing on the target code stream.

[0042] As can be seen from the above technical solutions, in the embodiments of the present application, the decoding-side device only needs to support N decoders. The N decoders are used to decode the bitstreams corresponding to N cameras, and the N decoders are used to perform specified video data processing on the bitstreams. Therefore, the memory resources and decoding resources of the N decoders can be reused, and there is no need to additionally support M decoders. In this way, the memory resources and decoding resources occupied by the M decoders can be saved, and the memory resources and decoding resources of N+M are reduced to those of N, thereby reducing the consumption of memory resources and decoding resources. It realizes the specified video data processing without additional resource consumption, can effectively save memory resources and decoding resources, and improve the processing efficiency.

[0043] The technical solutions of the embodiments of the present application will be described below in conjunction with specific application scenarios.

[0044] In a possible implementation manner, as shown in Figure 2A a schematic structural diagram of a video data processing system. The video data processing system may include N cameras and a decoding-side device. The decoding-side device is connected to the N cameras, and the decoding-side device needs to support N decoders. The N decoders are used to decode the bitstreams corresponding to the N cameras. N can be configured according to actual needs. N can be a positive integer, and there is no limitation thereto. For example, N can be 4, 8, 12, 16, 18, etc. In Figure 2A it, N is taken as 4 as an example.

[0045] Exemplarily, the N cameras correspond to the N decoders one by one. The N cameras are used to generate N bitstreams, and the N bitstreams correspond to the N decoders one by one. For example, camera 1 generates bitstream 1, sends bitstream 1 to the decoding-side device, and the decoding-side device sends bitstream 1 to decoder 1, and decoder 1 decodes bitstream 1. Camera 2 generates bitstream 2, sends bitstream 2 to the decoding-side device, and the decoding-side device sends bitstream 2 to decoder 2, and decoder 2 decodes bitstream 2, and so on.

[0046] Exemplarily, for the decoding-side device, it usually has a specified video data processing function (video data processing can be called service processing, that is, a specified service processing function). To implement the specified video data processing function, as shown in Figure 2B the decoding-side device also needs to reserve M decoders, and the M decoders are used to perform specified video data processing on the bitstreams. M can be configured according to actual needs. M can be a positive integer, and there is no limitation thereto. For example, M can be 2, 3, 4, etc. Figure 2B In it, M is taken as 2 as an example.

[0047] As can be seen from the above, the decoding-side device needs to support at least N + M decoders, such as a 4 + 2 decoder. For each decoder, it requires independent memory resources and independent decoding resources. Therefore, when the decoding-side device needs to support N + M decoders, it is necessary to reserve memory resources for N + M channels and decoding resources for N + M channels, thus consuming relatively large memory resources and decoding resources.

[0048] In view of the above findings, in the embodiments of the present application, the decoding-side device only needs to support N decoders. The N decoders are used to decode the code streams corresponding to N cameras, and the N decoders are used to perform specified video data processing, so that the memory resources and decoding resources of the N decoders can be reused, and there is no need to additionally support M decoders, thus saving the memory resources and decoding resources occupied by the M decoders.

[0049] In the embodiments of the present application, a video data processing system is proposed. Refer to Figure 2C As shown in the structure schematic diagram of the video data processing system, the video data processing system may include N cameras and a decoding-side device. The decoding-side device only needs to support N decoders. The N decoders correspond to the N cameras one by one, and the N decoders correspond to the code streams generated by the N cameras one by one. Obviously, the N decoders are used to decode the code streams corresponding to the N cameras, and the N decoders can perform specified video data processing on the code streams.

[0050] In the embodiments of the present application, the decoding-side device may be a backend device (corresponding to front-end devices such as cameras), such as a hard disk recorder. The hard disk recorder may be an XVR (XVideo Recorder, infinite possibility hard disk recorder), an NVR (Network Video Recorder), etc. Hereinafter, the hard disk recorder will be taken as an example for description. After the hard disk recorder is connected to the network, it can support the decoding preview function and the local playback function. After the hard disk recorder is connected to the network, it can also support the specified video data processing function.

[0051] Among them, after the hard disk recorder receives the code stream, it can decode the code stream through the decoder to obtain a video image, that is, restore the encoded code stream to a video image. After obtaining the video image, the decoding preview function and the local playback function can be realized based on the video image.

[0052] After the hard disk recorder receives the code stream, it can also perform specified video data processing on the code stream through the decoder, such as restoring the code stream to a video image and realizing the specified video data processing function based on the video image.

[0053] Exemplarily, specifying video data processing refers to a function with at least one of the following characteristics: a function with low usage frequency (such as being used only once per second), and a function with low usage frequency is also called a low-frequency function; a function that needs to be implemented by a decoder; a function that only processes I-frames, that is, a function that does not process P-frames.

[0054] For example, specifying video data processing can be screenshot processing. Screenshot processing means decoding the I-frame of a certain video stream into a YUV image and encoding the YUV image into a JPEG image. Of course, screenshot processing is only an example and is not limited thereto. Subsequently, screenshot processing will be used as an example for explanation.

[0055] In summary, specifying video data processing can be called a low-frequency function, that is, only the decoder processes I-frames and does not process P-frames, and the processing frequency of I-frames is very low, such as only processing a few I-frames per day, etc. In contrast, the normal decoding function can be called a high-frequency function, that is, the decoder needs to process a continuous GOP sequence, and each GOP sequence includes an I-frame and multiple P-frames, and the decoder needs to continuously decode these video frames.

[0056] See Figure 3 As shown, it is a schematic diagram of the decoding process of the video stream. After the decoding-side device obtains the video stream, it can unpack the video stream to obtain the raw stream. For example, if the video stream obtained by the decoding-side device is an encapsulated video stream, when the video stream is a live stream, the encapsulation format is RTP (Real-time Transport Protocol) format, and when the video stream is a stored stream, the encapsulation format is PS format. Therefore, the decoding-side device can unpack the encapsulated video stream and splice and assemble the unpacked video stream into a raw stream.

[0057] After obtaining the raw stream, a complete frame of the video stream can be parsed from the raw stream, and the complete frame of the video stream can be sent to the decoder, and the decoder decodes the complete frame of the video stream to obtain the video image. Each frame of the video stream can be sent to the decoder, and the decoder decodes each frame of the video stream to obtain each frame of the video image.

[0058] After obtaining the raw stream, subsequent processing such as secondary encapsulation of the raw stream can also be performed, and this process is not limited. For example, the raw stream can be secondarily encapsulated and then stored and transmitted, etc.

[0059] Exemplarily, the decoding-side device supports N decoders, and each decoder implements decoding using the above process. For example, for the video stream 1 corresponding to camera 1, decoder 1 decodes the video stream 1 using the above process, and for the video stream 2 corresponding to camera 2, decoder 2 decodes the video stream 2 using the above process, and so on.

[0060] See Figure 4As shown in the figure, video images can be divided into I-frames and P-frames. An I-frame is a video image that can be independently encoded and decoded during the encoding and decoding process and does not require dependence on other frames for encoding and decoding. A P-frame is a video image that cannot be independently encoded and decoded during the encoding and decoding process and requires dependence on other frames for encoding and decoding, that is, after the frames it depends on are encoded and decoded, the corresponding reference information needs to be obtained to complete the encoding and decoding of this frame of video image.

[0061] Based on the division method of I-frames and P-frames, all video images between two adjacent I-frames can be divided into the same GOP sequence. That is to say, starting from one I-frame to the last P-frame before the next I-frame, they all belong to the same GOP sequence, that is, a GOP sequence includes one I-frame and multiple P-frames. For example, all video frames successively include I-frame 1, P-frame 2, P-frame 3, P-frame 4, I-frame 5, P-frame 6, P-frame 7, P-frame 8, I-frame 9, P-frame 10, …, and so on. GOP sequence 1 includes I-frame 1, P-frame 2, P-frame 3, P-frame 4, GOP sequence 2 includes I-frame 5, P-frame 6, P-frame 7, P-frame 8, and so on.

[0062] Combined with Figure 3 and Figure 4 it can be seen that the decoder needs to decode I-frame 1, P-frame 2, P-frame 3, P-frame 4, I-frame 5, P-frame 6, P-frame 7, P-frame 8, I-frame 9, P-frame 10 in sequence. When the decoder decodes I-frame 1, the decoding process does not need to refer to the decoding information of other video frames. When the decoder decodes P-frame 2, the decoding process needs to refer to the decoding information of I-frame 1. When the decoder decodes P-frame 3, the decoding process needs to refer to the decoding information of I-frame 1. When the decoder decodes P-frame 4, the decoding process needs to refer to the decoding information of I-frame 1. When the decoder decodes I-frame 5, the decoding process does not need to refer to the decoding information of other video frames, and so on.

[0063] In summary, after the decoder decodes I-frame 1, it needs to decode P-frame 2, P-frame 3, P-frame 4 based on the decoding information of I-frame 1, that is, the decoding information of I-frame 1 acts on all P-frames within GOP sequence 1 until the first I-frame 5 of GOP sequence 2. At this time, the decoder no longer needs the decoding information of I-frame 1.

[0064] During the decoding process of GOP sequence 1, since the decoding process needs to use the decoding information of I-frame 1, the decoder cannot be occupied by the specified video data processing process. Otherwise, the decoding information of I-frame 1 will be lost, resulting in the inability to complete the decoding process of GOP sequence 1 until the decoding of P-frame 4 is completed, that is, the decoding of all frames within GOP sequence 1 is completed. Before decoding the first I-frame 5 of GOP sequence 2, since the decoding information of I-frame 1 is no longer used, the decoder can be occupied by the specified video data processing process.

[0065] Based on the above principle, an embodiment of the present application proposes a video data processing method, which is applied to a decoding-side device. The decoding-side device includes N decoders for decoding N streams, and the N decoders correspond to the N streams one by one. Refer to Figure 5 As shown in the flowchart of the method, the method may include:

[0066] Step 501: Determine the target stream that needs to perform specified video data processing.

[0067] Exemplarily, the decoding-side device supports the specified video data processing function, that is, it is necessary to perform specified video data processing on the stream through the decoder. Based on this, in order to implement the specified video data processing function, the stream that needs to perform specified video data processing can be configured for the decoding-side device. For the convenience of distinction, the stream that needs to perform specified video data processing can be called the target stream. For example, the stream 1 corresponding to camera 1 can be used as the target stream, such as a part of the stream corresponding to camera 1 as the target stream, or the stream 2 corresponding to camera 2 can be used as the target stream, and there is no limit to this.

[0068] In summary, the decoding-side device can determine the target stream that needs to perform specified video data processing.

[0069] Step 502: When there is a target stream that needs to perform specified video data processing, determine whether there is a decoder in the N decoders that is in a sleep state. If so, step 503 can be executed. If not, after waiting for a preset duration, continue to determine whether there is a decoder in the N decoders that is in a sleep state, and so on, until there is a decoder in the N decoders that is in a sleep state.

[0070] Exemplarily, a decoder in a sleep state means that the decoder has completed the decoding of all frames within the current GOP sequence and has not started to decode the next GOP sequence of the current GOP sequence.

[0071] Exemplarily, regarding how to know that the decoder has entered the sleep state, the following method can be adopted: Since the GOP sequence includes an I-frame and multiple P-frames, and the first frame of the GOP sequence is an I-frame, when the decoder decodes each frame within the GOP sequence, it can be determined whether the next video frame to be decoded is an I-frame. If it is not an I-frame, it indicates that the decoder has not completed the decoding of all frames within the current GOP sequence, and the decoder has not entered the sleep state. If the next video frame to be decoded is an I-frame, it indicates that the decoder has completed the decoding of all frames within the current GOP sequence but has not started to decode the next GOP sequence after the current GOP sequence, and the decoder has entered the sleep state. Among them, the information indicating whether the video frame within the GOP sequence is an I-frame can be carried in the bitstream. Therefore, it is possible to know whether the video frame within the GOP sequence is an I-frame, and then know whether the next video frame to be decoded is an I-frame.

[0072] In a possible implementation manner, for each decoder, the bitstream corresponding to the decoder may include the current GOP sequence and the next GOP sequence after the current GOP sequence. The current GOP sequence may include a first I-frame and multiple first P-frames, and the next GOP sequence may include a second I-frame and multiple second P-frames. On this basis, if the decoder has completed the decoding of the first I-frame and the decoding of multiple first P-frames, and the decoder has not started to decode the second I-frame, then the decoder is a decoder in the sleep state; otherwise, it can be known that the decoder is not a decoder in the sleep state.

[0073] For example, decoder 1 corresponds to bitstream 1. Bitstream 1 may include I-frame 1, P-frame 2, P-frame 3, P-frame 4, I-frame 5, P-frame 6, P-frame 7, P-frame 8, I-frame 9, P-frame 10, etc. GOP sequence 1 includes I-frame 1, P-frame 2, P-frame 3, P-frame 4, and GOP sequence 2 includes I-frame 5, P-frame 6, P-frame 7, P-frame 8, and so on. On this basis, when the current GOP sequence is GOP sequence 1, the next GOP sequence after the current GOP sequence is GOP sequence 2; when the current GOP sequence is GOP sequence 2, the next GOP sequence after the current GOP sequence is GOP sequence 3, and so on.

[0074] Taking the current GOP sequence as GOP sequence 1 and the next GOP sequence as GOP sequence 2 as an example, then, the first I-frame is I-frame 1, the multiple first P-frames are P-frame 2, P-frame 3, P-frame 4, the second I-frame is I-frame 5, and the multiple second P-frames are P-frame 6, P-frame 7, P-frame 8. On this basis:

[0075] If decoder 1 has completed the decoding of I-frame 1 and has completed the decoding of P-frame 2, P-frame 3, and P-frame 4, but decoder 1 has not started decoding I-frame 5, then decoder 1 is a decoder in the sleep state. Otherwise, if decoder 1 has not completed the decoding of P-frame 4, or decoder 1 has already decoded I-frame 5, etc., it is determined that decoder 1 is not a decoder in the sleep state.

[0076] In summary, it can be known whether decoder 1 is a decoder in the sleep state. Similarly, it can be known whether decoder 2 is a decoder in the sleep state, whether decoder 3 is a decoder in the sleep state, and whether decoder 4 is a decoder in the sleep state.

[0077] In the above embodiment, after the decoder has completed the decoding of all video frames within the current GOP sequence and before the decoder decodes the video frames within the next GOP sequence, it can be considered that this decoder is in the sleep state, and all resources corresponding to this decoder (such as memory resources and decoding resources, etc.) are in a short-term idle state. Therefore, this decoder can be occupied by the specified video data processing.

[0078] In addition, if the decoder has not completed the decoding of all video frames within the current GOP sequence, that is, a complete GOP sequence has not ended, then the resources corresponding to this decoder (such as memory resources and decoding resources, etc.) cannot be released. That is to say, this decoder is not in the sleep state, and the resources corresponding to this decoder are not in a short-term idle state. Therefore, this decoder cannot be occupied by the specified video data processing.

[0079] Step 503: Select a target decoder in the N decoders that is in the sleep state. The target decoder in the sleep state refers to: the target decoder has completed the decoding of all frames within the current GOP sequence and has not started decoding the next GOP sequence of the current GOP sequence.

[0080] Exemplarily, if there is a decoder in the sleep state among the N decoders and there is only one decoder in the sleep state, then this decoder in the sleep state is used as the target decoder. Or, if there are at least two decoders in the sleep state, then one decoder is selected from the at least two decoders in the sleep state as the target decoder. For example, a decoder is randomly selected as the target decoder, or a decoder is selected as the target decoder by adopting a certain strategy. There is no limitation on this selection strategy.

[0081] In summary, a target decoder in a sleep state can be selected from the N-channel decoder. For the target decoder, the bitstream corresponding to the target decoder may include the current GOP sequence and the next GOP sequence of the current GOP sequence. The current GOP sequence may include a first I-frame and multiple first P-frames, and the next GOP sequence may include a second I-frame and multiple second P-frames. The target decoder has completed the decoding of the first I-frame and multiple first P-frames, and the target decoder has not started decoding the second I-frame.

[0082] Step 504: Process the specified video data on the target bitstream through the target decoder.

[0083] In a possible implementation, taking the specified video data processing as screenshot processing as an example, the I-frame in the target bitstream can be input to the target decoder, and the I-frame is decoded through the target decoder to obtain a first image in a first image format, and the first image is converted into a second image in a second image format, and the second image is output. Exemplarily, the first image format may include but is not limited to the YUV format, and the second image format may include but is not limited to the JEPG format, and there is no limitation on this image format.

[0084] For example, the I-frame in the target bitstream can be input to the target decoder, and the I-frame is decoded through the target decoder to obtain a first image in the YUV format. After decoding the I-frame through the target decoder, the decoding resources of the target decoder can be released to resume the decoding process of the target decoder. For the specific process, refer to step 505. In addition, for the first image in the YUV format, the first image in the YUV format can also be encoded into a second image in the JEPG format and the second image is output, and this process will not be elaborated here.

[0085] Exemplarily, by decoding the I-frame through the target decoder to obtain a first image in the YUV format, the processing time of this process is relatively short, within a few milliseconds. Therefore, it will not affect the decoding process of the target decoder and will not cause the decoding process of the target decoder to be abnormal, that is, when implementing the specified video data processing function through the target decoder, the decoding process of the target decoder will not be abnormal.

[0086] Exemplarily, each complete GOP sequence may include video frames of 1 second to several seconds. Therefore, if only the resources of a certain decoder are multiplexed, there will be a long waiting time. In this embodiment, when multiplexing decoder resources is required, it is queried whether there is a decoder in the sleep state among all decoders. If there is no decoder in the sleep state, the query operation continues until a decoder in the sleep state is found, and then the query operation stops. In this way, the waiting time can be saved. When the number of decoders is relatively large, it is easy to find a decoder in the sleep state, that is, the waiting time can be greatly reduced.

[0087] Exemplarily, since the target decoder has completed the decoding of the current GOP sequence and has not started the decoding of the next GOP sequence, the target decoder will have a short sleep time. The screenshot processing function can utilize the short resource idle time to use the decoding resources of the target decoder. Obviously, the target decoder can execute the screenshot processing function during the resource idle time. The start time of the resource idle time is the time when the target decoder has completed the decoding of the current GOP sequence, and the end time of the resource idle time is the time when the target decoder has completed the screenshot processing, that is, the resources are released after the screenshot processing is completed.

[0088] Step 505: After the specified video data is processed, the target decoder decodes the bitstream corresponding to the target decoder, and the bitstream may include the next GOP sequence of the current GOP sequence.

[0089] Exemplarily, after the specified video data is processed, the decoding process of the target decoder can be resumed. That is to say, the target decoder decodes the bitstream corresponding to the target decoder again.

[0090] In summary, it can be seen that assuming the current GOP sequence includes I-frame 1, P-frame 2, P-frame 3, P-frame 4, and the next GOP sequence of the current GOP sequence includes I-frame 5, P-frame 6, P-frame 7, P-frame 8, then after the target decoder completes the decoding of I-frame 1, P-frame 2, P-frame 3, P-frame 4, the target decoder can perform specified video data processing on the target bitstream. After the specified video data processing is completed, the target decoder decodes I-frame 5, P-frame 6, P-frame 7, P-frame 8. The decoding process will not be elaborated here.

[0091] In the embodiment of the present application, another video data processing method is proposed, which can be applied to a decoding-side device. The decoding-side device includes N decoders for decoding N bitstreams, and the N decoders correspond to the N bitstreams one by one. Refer to Figure 6 As shown, it is a schematic flowchart of the method. The method may include:

[0092] Step 601: Determine the target bitstream that needs to be processed for the specified video data.

[0093] Step 602: When there is a target bitstream that needs to be processed for the specified video data, determine whether there is a candidate decoder in the N decoders that is in an idle state; where the bitstream corresponding to the candidate decoder does not include the GOP sequence to be processed. If so, execute Step 603; if not, execute Step 604.

[0094] Exemplarily, for each decoder, if the bitstream corresponding to the decoder does not include the GOP sequence to be processed, that is, the decoder does not have a corresponding bitstream to be processed, that is, the resources corresponding to the decoder (such as memory resources and decoding resources, etc.) are not occupied, then this decoder is a decoder in an idle state, and this decoder can be used as a candidate decoder in an idle state. If the bitstream corresponding to the decoder includes the GOP sequence to be processed, then this decoder is not a decoder in an idle state.

[0095] Step 603: Perform the specified video data processing on the target bitstream through the candidate decoder.

[0096] In a possible implementation manner, taking the specified video data processing as frame capture processing as an example, the I-frame in the target bitstream can be input to the candidate decoder, the I-frame is decoded through the candidate decoder to obtain a first image in a first image format, and the first image is converted into a second image in a second image format and the second image is output. Exemplarily, the first image format may include but is not limited to the YUV format, and the second image format may include but is not limited to the JEPG format, and there is no limitation on this image format.

[0097] Step 604: Determine whether there is a decoder in the N decoders that is in a sleep state. If so, execute Step 605. If not, after waiting for a preset duration, continue to determine whether there is a decoder in the N decoders that is in a sleep state, and so on, until there is a decoder in the N decoders that is in a sleep state. A decoder in a sleep state means that the decoder has completed the decoding of all frames within the current GOP sequence and the decoder has not started to decode the next GOP sequence of the current GOP sequence.

[0098] Step 605: Select a target decoder in the N decoders that is in a sleep state. Among them, the target decoder in a sleep state means that the target decoder has completed the decoding of all frames within the current GOP sequence and the target decoder has not started to decode the next GOP sequence of the current GOP sequence.

[0099] Step 606: Perform the specified video data processing on the target bitstream through the target decoder.

[0100] In a possible implementation manner, taking the specified video data processing as the screenshot processing as an example, the I-frame in the target bitstream can be input to the target decoder, and the target decoder decodes the I-frame to obtain a first image in the first image format, converts the first image into a second image in the second image format, and outputs the second image. Exemplarily, the first image format may include but is not limited to the YUV format, and the second image format may include but is not limited to the JEPG format. There is no limitation on the image format.

[0101] Step 607, after the specified video data processing is completed, the target decoder decodes the bitstream corresponding to the target decoder, and the bitstream may include the next GOP sequence of the current GOP sequence.

[0102] As can be seen from the above technical solutions, in the embodiments of the present application, the decoding-side device only needs to support N decoders, and the N decoders decode the bitstreams corresponding to N cameras, and the N decoders perform specified video data processing on the bitstreams, so as to be able to reuse the memory resources and decoding resources of the N decoders, without the need to additionally support M decoders. In this way, the memory resources and decoding resources occupied by the M decoders can be saved, and the memory resources and decoding resources of N + M channels are reduced to the memory resources and decoding resources of N channels, thereby reducing the consumption of memory resources and decoding resources, and realizing that the specified video data processing can be achieved without additional resource consumption, effectively saving memory resources and decoding resources, and improving the processing efficiency.

[0103] In this embodiment, a resource (such as memory resources and decoding resources) reuse mechanism is proposed, which can reuse the resources of N decoders. For example, when the decoder completes the decoding of all video frames within the GOP sequence, the decoder will have a short idle time, and another screenshot function can utilize the short resource idle time to use the resources of the decoder to complete the screenshot function. Among them, when resource reuse is required, it can be queried whether there are idle resources in the current N decoders, and the resource that can be released fastest is used for processing, so as to be able to reduce the waiting delay and achieve the purpose of resource sharing.

[0104] Based on the same application concept as the above method, a video data processing device is proposed in the embodiments of the present application. The video data processing device is applied to a decoding-side device, and the decoding-side device includes N decoders for decoding N bitstreams. See Figure 7 As shown, it is a schematic structural diagram of the video data processing device. The video data processing device may include:

[0105] A determination module 71, configured to determine a target bitstream that needs to perform specified video data processing;

[0106] A selection module 72 is configured to select a target decoder in a sleep state from the N decoders; wherein, the target decoder has completed decoding of all frames within the current GOP sequence and has not started decoding the next GOP sequence of the current GOP sequence.

[0107] A processing module 73 is configured to perform specified video data processing on the target bitstream through the target decoder.

[0108] A decoding module 74 is configured to, after the specified video data processing is completed, decode the bitstream corresponding to the target decoder through the target decoder, and the bitstream includes the next GOP sequence.

[0109] Exemplarily, when the selection module 72 selects a target decoder in a sleep state from the N decoders, it is specifically configured to: for each decoder, the bitstream corresponding to the decoder includes the current GOP sequence and the next GOP sequence of the current GOP sequence, the current GOP sequence includes a first I-frame and a plurality of first P-frames, and the next GOP sequence includes a second I-frame and a plurality of second P-frames. If the decoder has completed decoding of the first I-frame, completed decoding of the plurality of first P-frames, and has not started decoding the second I-frame, then the decoder is selected as the target decoder in a sleep state.

[0110] Exemplarily, the specified video data processing includes screenshot processing. When the processing module 73 performs specified video data processing on the target bitstream through the target decoder, it is specifically configured to: input the I-frame in the target bitstream to the target decoder; decode the I-frame through the target decoder to obtain a first image in a first image format, convert the first image to a second image in a second image format, and output the second image; the first image format includes the YUV format, and the second image format includes the JEPG format.

[0111] Exemplarily, when the selection module 72 selects a target decoder in a sleep state from the N decoders, it is specifically configured to: determine whether there is a decoder in a sleep state among the N decoders; if so, select a target decoder in a sleep state from the N decoders; if not, after waiting for a preset duration, return to perform the operation of determining whether there is a decoder in a sleep state among the N decoders until there is a decoder in a sleep state.

[0112] Exemplarily, the processing module 73 is further configured to determine whether there is a candidate decoder in an idle state among the N decoders; wherein, the bitstream corresponding to the candidate decoder does not include the GOP sequence to be processed; if so, perform specified video data processing on the target bitstream through the candidate decoder.

[0113] Based on the same application concept as the above method, in an embodiment of the present application, a decoding-side device is proposed. Refer to Figure 8 As shown, the decoding-side device includes: a processor 81 and a machine-readable storage medium 82. The machine-readable storage medium 82 stores machine-executable instructions that can be executed by the processor 81. The processor 81 is configured to execute the machine-executable instructions to implement the video data processing method disclosed in the above examples of the present application.

[0114] Based on the same application concept as the above method, an embodiment of the present application further provides a machine-readable storage medium. A number of computer instructions are stored on the machine-readable storage medium. When the computer instructions are executed by a processor, the video data processing method disclosed in the above examples of the present application can be implemented.

[0115] Among them, the above machine-readable storage medium can be any electronic, magnetic, optical or other physical storage device that can contain or store information, such as executable instructions, data, etc. For example, the machine-readable storage medium can be: RAM (Random Access Memory), volatile memory, non-volatile memory, flash memory, storage drives (such as hard disk drives), solid-state drives, any type of storage disk (such as optical discs, DVDs, etc.), or similar storage media, or a combination thereof.

[0116] Based on the same application concept as the above method, a video data processing system is proposed in an embodiment of the present application. The video data processing system includes N cameras and a decoding-side device. The N cameras are used to generate N bitstreams. The decoding-side device includes N decoders for decoding the N bitstreams. The N decoders correspond to the N bitstreams one by one. Among them:

[0117] For each camera, the camera is configured to collect the current GOP sequence and the next GOP sequence of the current GOP sequence, generate a bitstream corresponding to the camera, the bitstream includes the current GOP sequence and the next GOP sequence, and send the bitstream to the decoder corresponding to the camera. The decoder is configured to decode the bitstream after receiving the bitstream.

[0118] The decoding-side device is configured to determine a target bitstream that needs to be subjected to specified video data processing; select a target decoder in a dormant state from the N decoders, and send the target bitstream to the target decoder. Among them, the target decoder has completed the decoding of all frames within the current GOP sequence, and the target decoder has not started to decode the next GOP sequence of the current GOP sequence.

[0119] The target decoder is used to perform specified video data processing on the target bitstream; after the specified video data processing is completed, decoding of the bitstream corresponding to the target decoder is continued.

[0120] The systems, devices, modules or units illustrated in the above embodiments can be specifically implemented by computer chips or entities, or by products with certain functions. A typical implementation device is a computer, and the specific form of the computer can be a personal computer, a laptop computer, a cellular phone, a camera phone, a smart phone, a personal digital assistant, a media player, a navigation device, an email transceiver device, a game console, a tablet computer, a wearable device, or a combination of any several of these devices.

[0121] For convenience of description, the above devices are described by dividing them into various units according to functions. Of course, when implementing the present application, the functions of each unit can be implemented in one or more software and / or hardware.

[0122] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the embodiments of the present application can adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0123] The present application is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or block in the flowchart and / or block diagram can be implemented by computer program instructions, and the combination of the processes and / or blocks in the flowchart and / or block diagram can also be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate a device for implementing the functions specified in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.

[0124] Moreover, these computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing devices to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured article including an instruction device, and the instruction device implements the functions specified in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.

[0125] These computer program instructions can also be loaded onto a computer or other programmable data processing apparatus, so that a series of operation steps are performed on the computer or other programmable apparatus to generate a computer-implemented process, thereby providing instructions for implementing the steps of the function specified in one process or a plurality of processes and / or blocks Figure 1 one process or a plurality of processes and / or blocks Figure 1 in one block or a plurality of blocks.

[0126] The above are only embodiments of the present application and are not intended to limit the present application. Various changes and modifications can be made to the present application for those skilled in the art. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included within the scope of the claims of the present application.

Claims

1. A video data processing method, characterized in that, The decoding-side device includes N decoders for decoding N bitstreams, and the method includes: Determine a target bitstream that needs to undergo specified video data processing, where the specified video data processing includes capture processing; Select a target decoder in a dormant state from the N decoders; where the target decoder has completed decoding all frames within the current GOP sequence and has not started decoding the next GOP sequence of the current GOP sequence; Perform specified video data processing on the target bitstream through the target decoder; After the specified video data processing is completed, decode the bitstream corresponding to the target decoder through the target decoder, and this bitstream includes the next GOP sequence.

2. The method according to claim 1, wherein: The step of selecting a target decoder in a dormant state from the N decoders includes: For each decoder, the bitstream corresponding to this decoder includes the current GOP sequence and the next GOP sequence of the current GOP sequence. The current GOP sequence includes a first I-frame and multiple first P-frames, and the next GOP sequence includes a second I-frame and multiple second P-frames. If this decoder has completed decoding the first I-frame, completed decoding the multiple first P-frames, and has not started decoding the second I-frame, then select this decoder as the target decoder in a dormant state.

3. The method according to claim 1, wherein: The step of performing specified video data processing on the target bitstream through the target decoder includes: Input the I-frame in the target bitstream to the target decoder; Decode the I-frame through the target decoder to obtain a first image in a first image format, convert the first image to a second image in a second image format, and output the second image.

4. The method according to claim 3, wherein: The first image format includes the YUV format, and the second image format includes the JEPG format.

5. The method according to claim 1 or 2, wherein: The step of selecting a target decoder in a dormant state from the N decoders includes: Judge whether there is a decoder in a dormant state among the N decoders; If so, select a target decoder in a dormant state from the N decoders; If not, wait for a preset duration and then return to perform the operation of judging whether there is a decoder in a dormant state among the N decoders until there is a decoder in a dormant state.

6. The method according to any one of claims 1-4, characterized in that, Before the step of selecting a target decoder in a dormant state from the N decoders, the method further includes: Judge whether there is a candidate decoder in an idle state among the N decoders; where the bitstream corresponding to the candidate decoder does not include the GOP sequence that needs to be processed; If so, perform specified video data processing on the target bitstream through the candidate decoder; If not, perform the operation of selecting a target decoder in a dormant state from the N decoders.

7. A video data processing device, characterized in that, The decoding-side device includes N decoders for decoding N bitstreams, and the device includes: A determination module, configured to determine a target bitstream that needs to perform specified video data processing, where the specified video data processing includes screenshot processing; A selection module, configured to select a target decoder in a dormant state from the N decoders; where the target decoder has completed decoding of all frames within the current GOP sequence, and the target decoder has not started decoding the next GOP sequence of the current GOP sequence; A processing module, configured to perform specified video data processing on the target bitstream through the target decoder; A decoding module, configured to, after the specified video data processing is completed, decode the bitstream corresponding to the target decoder through the target decoder, and the bitstream includes the next GOP sequence.

8. The device according to claim 7, It is characterized in that Wherein When the selection module selects a target decoder in a dormant state from the N decoders, it specifically is configured to: For each decoder, the bitstream corresponding to the decoder includes the current GOP sequence and the next GOP sequence of the current GOP sequence, the current GOP sequence includes a first I-frame and multiple first P-frames, and the next GOP sequence includes a second I-frame and multiple second P-frames. If the decoder has completed decoding of the first I-frame, and completed decoding of the multiple first P-frames, and has not started decoding the second I-frame, then the decoder is selected as the target decoder in a dormant state; Wherein, when the processing module performs specified video data processing on the target bitstream through the target decoder, it specifically is configured to: Input the I-frame in the target bitstream to the target decoder; decode the I-frame through the target decoder to obtain a first image in a first image format, convert the first image to a second image in a second image format, and output the second image; the first image format includes the YUV format, and the second image format includes the JEPG format; Wherein, when the selection module selects a target decoder in a dormant state from the N decoders, it specifically is configured to: Determine whether there is a decoder in a dormant state among the N decoders; if so, select a target decoder in a dormant state from the N decoders; if not, after waiting for a preset duration, return to perform the operation of determining whether there is a decoder in a dormant state among the N decoders until there is a decoder in a dormant state; Wherein, the processing module is further configured to determine whether there is a candidate decoder in an idle state among the N decoders; where the bitstream corresponding to the candidate decoder does not include the GOP sequence that needs to be processed; If so, perform specified video data processing on the target bitstream through the candidate decoder.

9. A decoding-side device, characterized in that, Including: A processor and a machine-readable storage medium, where the machine-readable storage medium stores machine-executable instructions that can be executed by the processor; The processor is configured to execute the machine-executable instructions to implement the method steps described in any one of claims 1-6.

10. A video data processing system, characterized in that, Including N cameras and a decoding-side device, the N cameras are used to generate N bitstreams, and the decoding-side device includes N decoders for decoding the N bitstreams, where: For each camera, the camera is used to collect the current GOP sequence and the next GOP sequence after the current GOP sequence, generate a bitstream corresponding to the camera, the bitstream includes the current GOP sequence and the next GOP sequence, and send the bitstream to the decoder corresponding to the camera; the decoder is used to decode the bitstream after receiving the bitstream; The decoding-side device is used to determine a target bitstream that needs to be subjected to specified video data processing, where the specified video data processing includes frame capture processing; and select a target decoder in a dormant state from the N decoders, and send the target bitstream to the target decoder; where the target decoder has completed decoding all frames within the current GOP sequence and the target decoder has not started decoding the next GOP sequence after the current GOP sequence; The target decoder is used to perform specified video data processing on the target bitstream; after the specified video data processing is completed, continue to decode the bitstream corresponding to the target decoder.

Citation Information

Patent Citations

  • Image decoding method and device based on parallel thread, equipment and storage medium

    CN113395523A

  • Video data processing method and device, equipment and storage medium

    CN114125432A