A video stream decoding method, system, device and medium

By splitting the video stream data, combining hardware decoding and software decoding, combined with image recognition processing and filtering mixing technology, the performance bottleneck problem during high-resolution video stream decoding is solved, and higher decoding accuracy and image rendering effect are achieved.

CN119697384BActive Publication Date: 2025-05-16CHANGSHA LEILEIYUN NETWORK TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510208212.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-25
Publication Date
2025-05-16
Estimated Expiration
2045-02-25

AI Technical Summary

Technical Problem

In the prior art, when decoding high-resolution video streams, hardware decoding and software decoding have performance bottlenecks, resulting in loss of decoding accuracy and poor playback.

Method used

By splitting the video stream data, some of the data are decoded by hardware, some of the data are decoded by software, and combining image recognition processing and filtering mixing technology, the accuracy of decoding and image rendering effect are improved.

Benefits of technology

This method can reduce the decoding accuracy loss while improving the accuracy of video stream data decoding and image rendering effect, and is suitable for video stream data of complex forms and high encoding standards.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119697384B_ABST
    Figure CN119697384B_ABST
Patent Text Reader

Abstract

The present application relates to the field of data processing technology, and in particular to a video stream decoding method, system, device and medium, comprising hardware decoding and software decoding of the video stream data splitting result to obtain a first common texture image and a second common texture image; extracting initial feature information and saving it as identification information; performing image recognition processing on the first common texture image and the second common texture image, and updating the identification information according to the result; obtaining a target texture image according to the identification information, performing a first filtering processing on the first common texture image and the second common texture image to obtain a first filtered texture image; mixing with the first common texture image and the second common texture image to obtain a first mixed image; performing a second filtering processing on the target texture image to obtain a second filtered texture image; mixing with the target texture image to obtain a second mixed image; performing image rendering to improve the accuracy of video stream data decoding and thus improve the image rendering effect.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of data processing technology, and in particular to a video stream decoding method, system, device and medium. Background Art

[0002] With the popularization of the Internet and mobile devices, the production, sharing and consumption of video content have penetrated almost into all walks of life, from social media platforms to online education, from entertainment content to corporate training. Video streaming data has become an important part of people's daily lives. With the improvement of video resolution and the application of multiple encoding standards, the parsing and decoding of video streaming data has become more and more complicated.

[0003] Video stream decoding technology can be divided into two categories: hardware decoding and software decoding. Hardware decoding uses dedicated hardware to handle video stream decoding tasks, such as GPU or dedicated decoding chips. It has higher efficiency and lower power consumption, so it is widely used in high-performance video processing scenarios such as high-definition video playback and video conferencing. However, not all data forms can be decoded in all devices and application scenarios, so software decoding is needed. Software decoding uses the computer's central processing unit to decode video streams. Compared with hardware decoding, it is more flexible and can support a wider range of devices, but it faces a greater computing burden when decoding high-resolution video streams, which may lead to performance bottlenecks and affect playback fluency and efficiency.

[0004] In the current decoding scenario, it is usually determined whether hardware decoding can be performed based on the relevant information and category of the video stream. If it can be performed, the video stream is directly hardware decoded. If it cannot be performed, software decoding is selected, and finally the decoding-related images are output for rendering. This method can only occupy a large amount of system resources for software decoding when hardware decoding is not satisfied. As complex video stream data with high encoding standards becomes more and more popular, the complexity of decoding is also increasing. The use of single hardware decoding or software decoding limits the processing power, resulting in precision loss during the decoding process, affecting the decoding accuracy. The above problems need to be solved. Summary of the invention

[0005] In order to reduce the precision loss in the decoding process, improve the accuracy of video stream data decoding and thus improve the image rendering effect, the present application provides a video stream decoding method, system, device and medium, which adopts the following technical solutions:

[0006] In a first aspect, the present application provides a video stream decoding method, comprising:

[0007] Acquire video stream data, split the video stream data, perform hardware decoding and software decoding on the split results respectively, the hardware decoding obtains a first common texture image, and the software decoding obtains a second common texture image;

[0008] Acquire pre-selected target image information, extract initial feature information of the target image information, and save the initial feature information as identification information;

[0009] Performing image recognition processing on the first common texture image and the second common texture image according to the initial feature information, and updating the identification information according to the image recognition processing result;

[0010] Obtaining a target texture image in the first common texture image and the second common texture image according to the identification information, performing a first filtering process on the first common texture image and the second common texture image after filtering out the target texture image, to obtain a first filtered texture image;

[0011] Mixing the first filtered texture image with the corresponding first common texture image and the corresponding second common texture image according to a preset ratio to obtain a first mixed image;

[0012] Performing a second filtering process on the target texture image to obtain a second filtered texture image;

[0013] Mixing the second filtered texture image and the target texture image according to a preset ratio to obtain a second mixed image;

[0014] Image rendering is performed according to the first mixed image and the second mixed image.

[0015] Preferably, the specific steps of performing image recognition processing on the first common texture image and the second common texture image according to the initial feature information and updating the identification information according to the image recognition processing result are:

[0016] The first common texture image is subjected to image recognition processing according to the initial feature information, and when the first recognition information of the first common texture image is recognized to be the same as or similar to the initial feature information, the identification information is updated according to the first recognition information.

[0017] Preferably, the specific steps of performing image recognition processing on the first common texture image and the second common texture image according to the initial feature information and updating the identification information according to the image recognition processing result are:

[0018] The second common texture image is subjected to image recognition processing according to the initial feature information, and when the second recognition information of the second common texture image is recognized to be the same as or similar to the initial feature information, the identification information is updated according to the second recognition information.

[0019] Preferably, the specific steps of splitting the video stream data are:

[0020] Divide the video stream data into several parts according to the preset data size, and extract the encoding format information, frame type information, resolution information and multi-layer encoding information of each video stream data;

[0021] The coding format information, frame type information, resolution information and multi-layer coding situation information are analyzed according to a pre-built decision model to obtain a first split result suitable for hardware decoding or a second split result suitable for software decoding.

[0022] Preferably, the specific steps of obtaining the first common texture image by hardware decoding are:

[0023] Convert the first split result into OES texture information, and decode the first split result using a hardware decoding framework decoder to obtain an OES extended texture image;

[0024] The OES extended texture image is converted into a normal texture to obtain a first normal texture image.

[0025] Preferably, the specific steps of obtaining the second common texture image by software decoding are:

[0026] The second split result is converted into YUV texture information, and the second split result is decoded by a software decoding framework decoder to obtain a YUV extended texture image;

[0027] The YUV extended texture image is converted into a normal texture to obtain a second normal texture image.

[0028] Preferably, the specific steps of obtaining the pre-selected target image information and extracting the initial feature information of the target image information are:

[0029] Determine the selected range of the target image information, perform edge detection on the selected range, extract the features of the main object, and obtain initial feature information. The main object includes any one of the objects occupying the largest area, the object with the largest color difference from other colors, and the object with the highest clarity.

[0030] Preferably, the specific steps of mixing the first filtered texture image with the corresponding first common texture image and the corresponding second common texture image according to a preset ratio to obtain the first mixed image are:

[0031] Binding the first filtered texture image and the corresponding first common texture image to an OpenGL frame buffer, and mixing the first filtered texture image and the first common texture image in the frame buffer according to a preset ratio to obtain a first mixed image;

[0032] The first filtered texture image and the corresponding second common texture image are bound to the OpenGL frame buffer, and the first filtered texture image and the second common texture image are mixed in the frame buffer according to a preset ratio to obtain a first mixed image.

[0033] Preferably, the specific steps of mixing the second filtered texture image and the target texture image in a preset ratio to obtain the second mixed image are:

[0034] The second filtered texture image and the target texture image are bound to the OpenGL frame buffer, and the second filtered texture image and the target texture image are mixed in the frame buffer according to a preset ratio to obtain a second mixed image.

[0035] In a second aspect, the present application provides a video stream decoding system, comprising:

[0036] A video stream acquisition module is used to acquire video stream data, split the video stream data, and perform hardware decoding and software decoding on the split results, wherein the hardware decoding obtains a first common texture image, and the software decoding obtains a second common texture image;

[0037] A target image acquisition module, used to acquire pre-selected target image information, extract initial feature information of the target image information, and save the initial feature information as identification information;

[0038] An identification updating module, used for performing image recognition processing on the first common texture image and the second common texture image according to the initial feature information, and updating the identification information according to the image recognition processing result;

[0039] A first filtering processing module is used to obtain a target texture image in the first common texture image and the second common texture image according to the identification information, and perform a first filtering process on the first common texture image and the second common texture image after filtering out the target texture image to obtain a first filtered texture image;

[0040] A first image mixing module, used for mixing the first filtered texture image with the corresponding first common texture image and the corresponding second common texture image according to a preset ratio to obtain a first mixed image;

[0041] A second filtering processing module, used for performing a second filtering process on the target texture image to obtain a second filtered texture image;

[0042] A second image mixing module, used for mixing the second filtered texture image and the target texture image according to a preset ratio to obtain a second mixed image;

[0043] The image rendering module is used to perform image rendering according to the first mixed image and the second mixed image.

[0044] Preferably, the identification updating module performs image recognition processing on the first common texture image and the second common texture image according to the initial feature information, and the specific steps of updating the identification information according to the image recognition processing result are:

[0045] The first common texture image is subjected to image recognition processing according to the initial feature information, and when the first recognition information of the first common texture image is recognized to be the same as or similar to the initial feature information, the identification information is updated according to the first recognition information.

[0046] Preferably, the identification updating module performs image recognition processing on the first common texture image and the second common texture image according to the initial feature information, and the specific steps of updating the identification information according to the image recognition processing result are:

[0047] The second common texture image is subjected to image recognition processing according to the initial feature information, and when the second recognition information of the second common texture image is recognized to be the same as or similar to the initial feature information, the identification information is updated according to the second recognition information.

[0048] Preferably, the specific steps of splitting the video stream data by the video stream acquisition module are:

[0049] Divide the video stream data into several parts according to the preset data size, and extract the encoding format information, frame type information, resolution information and multi-layer encoding information of each video stream data;

[0050] The coding format information, frame type information, resolution information and multi-layer coding situation information are analyzed according to a pre-built decision model to obtain a first split result suitable for hardware decoding or a second split result suitable for software decoding.

[0051] Preferably, the specific steps of hardware decoding of the video stream acquisition module to obtain the first common texture image are:

[0052] Convert the first split result into OES texture information, and decode the first split result using a hardware decoding framework decoder to obtain an OES extended texture image;

[0053] The OES extended texture image is converted into a normal texture to obtain a first normal texture image.

[0054] Preferably, the specific steps of decoding the video stream acquisition module software to obtain the second ordinary texture image are:

[0055] The second split result is converted into YUV texture information, and the second split result is decoded by a software decoding framework decoder to obtain a YUV extended texture image;

[0056] The YUV extended texture image is converted into a normal texture to obtain a second normal texture image.

[0057] Preferably, the target image acquisition module acquires the pre-selected target image information, and the specific steps of extracting the initial feature information of the target image information are:

[0058] Determine the selected range of the target image information, perform edge detection on the selected range, extract the features of the main object, and obtain initial feature information. The main object includes any one of the objects occupying the largest area, the object with the largest color difference from other colors, and the object with the highest clarity.

[0059] Preferably, the first image mixing module mixes the first filtered texture image with the corresponding first common texture image and the corresponding second common texture image according to a preset ratio to obtain the first mixed image in the following specific steps:

[0060] Binding the first filtered texture image and the corresponding first common texture image to an OpenGL frame buffer, and mixing the first filtered texture image and the first common texture image in the frame buffer according to a preset ratio to obtain a first mixed image;

[0061] The first filtered texture image and the corresponding second common texture image are bound to the OpenGL frame buffer, and the first filtered texture image and the second common texture image are mixed in the frame buffer according to a preset ratio to obtain a first mixed image.

[0062] Preferably, the second image mixing module mixes the second filtered texture image and the target texture image according to a preset ratio to obtain the second mixed image in the following specific steps:

[0063] The second filtered texture image and the target texture image are bound to the OpenGL frame buffer, and the second filtered texture image and the target texture image are mixed in the frame buffer according to a preset ratio to obtain a second mixed image.

[0064] In a third aspect, the present application provides a video stream decoding device, comprising a memory and a processor, wherein the memory stores a computer program, and the processor is configured to run the computer program to execute the video stream decoding method as described above.

[0065] In a fourth aspect, the present application provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, wherein the computer program is configured to execute the video stream decoding method as described above when run.

[0066] In summary, compared with the prior art, the technical solution provided by this application has at least the following beneficial effects:

[0067] The present application splits the acquired video stream data, performs hardware decoding on part of the split results to obtain a first ordinary texture image, performs software decoding on part of the split results to obtain a second ordinary texture image, uses initial feature information extracted according to target image information after splitting and decoding as reference information, performs image recognition processing on the first ordinary texture image and the second ordinary texture image according to the initial feature information, extracts the recognition result to update identification information, extracts the target texture image according to the identification information, performs a first filtering processing on the first ordinary texture image and the second ordinary texture image after filtering out the target texture image, mixes the first filtered texture image obtained by filtering with the first ordinary texture image and the second ordinary texture image to obtain a first mixed image, renders the target texture image using a higher quality rendering method, that is, performs a second filtering processing on the target texture image to obtain a second filtered texture image, and then mixes the second filtered texture image and the target texture image according to a preset ratio to obtain a second mixed image. The present application can use a more targeted decoding processing method to process different types of data in a video stream data, reduce the precision loss in the decoding process, improve the accuracy of video stream data decoding, and thus improve the image rendering effect. BRIEF DESCRIPTION OF THE DRAWINGS

[0068] Figure 1 It is a flowchart of a video stream decoding method described in an embodiment of the present application.

[0069] Figure 2 It is a module diagram of a video stream decoding system described in an embodiment of the present application.

[0070] Figure 3 It is a module diagram of a video stream decoding device described in an embodiment of the present application.

[0071] Description of reference numerals:

[0072] 1. Video stream acquisition module; 2. Target image acquisition module; 3. Identification update module; 4. First filtering processing module; 5. First image mixing module; 6. Second filtering processing module; 7. Second image mixing module; 8. Image rendering module. DETAILED DESCRIPTION

[0073] The following combination Figure 1-Figure 3 The present application is described in further detail. The terms used in the embodiments of the present application are only for the purpose of describing specific embodiments and are not intended to be limiting.

[0074] Reference Figure 1 , a video stream decoding method involved in this application specifically includes:

[0075] Acquire video stream data, split the video stream data, perform hardware decoding and software decoding on the split results respectively, the hardware decoding obtains a first common texture image, and the software decoding obtains a second common texture image;

[0076] Acquire pre-selected target image information, extract initial feature information of the target image information, and save the initial feature information as identification information;

[0077] Performing image recognition processing on the first common texture image and the second common texture image according to the initial feature information, and updating the identification information according to the image recognition processing result;

[0078] Obtaining a target texture image in the first common texture image and the second common texture image according to the identification information, performing a first filtering process on the first common texture image and the second common texture image after filtering out the target texture image, to obtain a first filtered texture image;

[0079] Mixing the first filtered texture image with the corresponding first common texture image and the corresponding second common texture image according to a preset ratio to obtain a first mixed image;

[0080] Performing a second filtering process on the target texture image to obtain a second filtered texture image;

[0081] Mixing the second filtered texture image and the target texture image according to a preset ratio to obtain a second mixed image;

[0082] Image rendering is performed according to the first mixed image and the second mixed image.

[0083] Specifically, the composition of video stream data is complex. Simply detecting its type and classifying it into hardware decoding or software decoding will result in the video stream information not being fully adapted to the current device in the hardware decoding scenario, resulting in the inability of the hardware-supported format to decode the current video stream data. In the soft decoding scenario of using low-performance devices, the CPU's processing power is limited, resulting in a loss of accuracy during the decoding process. Therefore, a single decoding method has gradually failed to meet the current video stream information processing requirements.

[0084] The present application obtains video stream information, splits the obtained video stream data, performs hardware decoding on a part of the split results that meet the hardware decoding conditions to obtain a first common texture image, and performs software decoding on another part of the split results that meet the software decoding conditions to obtain a second common texture image. By analyzing several types of information of the video stream data, it is determined whether it is suitable for hardware decoding or software decoding after splitting, and a video stream data can be processed in segments and a suitable decoding method can be selected.

[0085] After split decoding, the initial feature information extracted from the target image information is saved as identification information, and the initial feature information is used as reference information for subsequent image recognition processing. Image recognition processing is performed on the first ordinary texture image and the second ordinary texture image according to the initial feature information, and the identification information is updated by extracting the recognition result. The target texture image is extracted according to the identification information, and the first ordinary texture image and the second ordinary texture image that screen out the target texture image are subjected to a first filtering process, and the first filtered texture image obtained by filtering is mixed with the first ordinary texture image and the second ordinary texture image to obtain a first mixed image. The target texture image obtained by screening is rendered using a higher quality rendering method, that is, a second filtering process is performed on the target texture image to obtain a second filtered texture image, and then the second filtered texture image and the target texture image are mixed according to a preset ratio to obtain a second mixed image, and finally the first mixed image and the second mixed image are combined and rendered to the screen. The present application can use a more targeted decoding processing method to process different types of data in a video stream data, reduce the precision loss in the decoding process, improve the accuracy of video stream data decoding, and thus improve the image rendering effect.

[0086] The embodiment of the present application adopts a solution of performing hardware decoding and software decoding simultaneously, which can give full play to the advantages of both. Hardware decoding processes conventional video streams to reduce energy consumption and improve processing efficiency, while software decoding is used to process special video formats or to supplement in scenarios where hardware decoding cannot be processed efficiently. The flexibility of decoding is improved to ensure that a high-quality playback experience can be maintained even in the case of complex video streams.

[0087] The first filtering process and the second filtering process of the present application are the same filtering processes for different target signals. The first and the second here are used to distinguish and process different information. For example, Gaussian filtering is performed on the first ordinary texture image and the second ordinary texture image, and then Gaussian filtering is performed on the target texture image.

[0088] As one implementation method, the specific steps of splitting the video stream data are as follows:

[0089] Divide the video stream data into several parts according to the preset data size, and extract the encoding format information, frame type information, resolution information and multi-layer encoding information of each video stream data;

[0090] The coding format information, frame type information, resolution information and multi-layer coding situation information are analyzed according to a pre-built decision model to obtain a first split result suitable for hardware decoding or a second split result suitable for software decoding.

[0091] Specifically, before decoding the video stream data, the embodiment of the present application needs to split the video stream data to a certain extent. The splitting is based on the preset data size. The video stream data is first divided into several parts, and the pre-divided parts of the video stream data are further allocated according to the encoding format information, frame type information, resolution information and multi-layer encoding information in the video stream data. The data segments that do not use the same type of decoding method are screened out from one video stream data, and then the data segments are divided into adjacent or similar video stream data, specifically, the nearest video stream data with different decoding methods. Through the embodiment of the present application, the video stream data can be evenly divided and then slightly adjusted, so that the split video stream data will not be too large or too small, and the segmented video stream data can be adapted to hardware decoding or software decoding.

[0092] Hardware decoders have good support for certain specific encoding formats and standards, such as H.264, H.265, HEVC, VP9, ​​etc. However, since different hardware supports different formats, the application scope of hardware decoding is limited. When the video stream uses a relatively new encoding format, such as AV1, only specific hardware supports these formats, and other hardware can only use software decoding.

[0093] Considering the frame type, I frame is a key frame in video encoding and is an independent image. The hardware decoder can decode I frame relatively quickly because it does not rely on the data of other frames, while P frame and B frame need to refer to the information of other frames. The hardware decoder needs to have a strong decoding ability to handle the relationship between these frames. When the I frame sequence in the video stream is dense, for example, a higher GOP length, the hardware decoder can decode these video streams relatively easily. If the video stream contains a large number of P frames or B frames, the complexity of the hardware decoder will increase, which is not suitable for running on low-power devices, and it is necessary to choose soft decoding to achieve decoding.

[0094] When the video resolution is very high, such as 4K or higher, hard decoding can usually provide higher performance due to its higher efficiency in processing high-resolution videos. If the video resolution is low or the bit rate is low, the load of soft decoding is usually relatively low, and soft decoding can provide sufficient performance.

[0095] When considering multi-layer encoding and complex encoding configurations, since some video streams may contain multi-layer encoding, for example, multi-level encoding with dynamic adaptive resolution or bit rate, or use complex encoding techniques, such as context adaptive binary arithmetic coding CABAC, they are usually not well implemented in hardware decoding and need to rely on soft decoding to complete, or be further split through the embodiments of the present application and assigned to a suitable hardware decoder for decoding.

[0096] One embodiment is a video stream that uses H.264 encoding and contains content with a resolution of 1080p and a frame rate of 30fps. The video stream also contains the following features: more I frames and fewer P frames and B frames. In this case, the hardware decoder can decode efficiently because the decoding of I frames is independent and the processing power of the hardware decoder is more suitable for decoding these frames. Some videos are encoded with HEVC. HEVC encoding has high hardware requirements, especially on low-end hardware. If the hardware does not support HEVC or does not support certain features in HEVC, this part of the video needs to be processed by soft decoding. AV1 format: if part of the video stream is encoded with AV1, soft decoding is required because the current hardware decoding support for AV1 is relatively limited, especially for older devices.

[0097] In order to better divide the video stream according to information such as encoding format information, frame type information, resolution information, and multi-layer encoding information, a decision model is pre-built. The decision model of the embodiment of the present application calculates the score, and a weight value is assigned to each reference factor. The weight value is adjusted according to the actual situation. For example, the weight of the encoding format is , the weight of the frame type is , the weight of resolution and bit rate is , the weight of the complex encoding configuration is .

[0098] Then determine the score of each parameter of the video stream. For example, if the hard solution supports the encoding format, the score is 1, otherwise, the score is 0. For example, if the frame type is suitable for hard solution, the score is 1; if it is suitable for soft solution, the score is 0. For example, if the resolution is high, the hard solution is better, the score is 1; if the resolution is low, the soft solution is better, the score is 0. For example, if the encoding configuration is simple, the hard solution is better, the score is 1; if the encoding configuration is complex, the soft solution is better, the score is 0. Finally, after multiplying the corresponding weight value and the score value, the final scores of several parameters are added up, the final scores are calculated, and then compared according to the score judgment value to determine whether to choose hard solution or soft solution.

[0099] Based on the decision model, the decision tree method is used to determine whether there are special cases in the segmented data of the video stream data. For example, if the encoding format is not supported by hard decoding, soft decoding is directly selected; if the frame type is mainly I-frame and the resolution is high, hard decoding is selected; if the encoding configuration is complex, such as CABAC and the video resolution is low, soft decoding is selected.

[0100] As one implementation method, the specific steps of hardware decoding to obtain the first common texture image are:

[0101] Convert the first split result into OES texture information, and decode the first split result using a hardware decoding framework decoder to obtain an OES extended texture image;

[0102] The OES extended texture image is converted into a normal texture to obtain a first normal texture image.

[0103] The specific steps of software decoding to obtain the second ordinary texture image are:

[0104] The second split result is converted into YUV texture information, and the second split result is decoded by a software decoding framework decoder to obtain a YUV extended texture image;

[0105] The YUV extended texture image is converted into a normal texture to obtain a second normal texture image.

[0106] Specifically, the embodiment of the present application determines whether to create a hardware decoder MediaCodec framework based on the hardware decoding support attributes of the Android device. If hardware decoding is not supported, it is necessary to create a software decoder, and the ffmpeg decoding framework is usually used for decoding. When creating a MediaCodec hardware decoder, it is necessary to create an additional OpenGL ES extended texture OES, configure it to MediaCodec for decoding, cache the decoded image data, and update the texture image in time to hand it over to OpenGLES for drawing and rendering to display the image. When creating an FFmpeg software decoder, the YUV420P image data is directly decoded and output, and it also needs to be handed over to OpenGL ES for drawing and rendering to display the image.

[0107] Among them, OpenGL is a cross-platform graphics rendering application program interface, which is commonly used in computer graphics, game development, virtual reality, CAD software and other fields. OpenGL provides a set of standard functions for performing various operations in 2D and 3D graphics rendering, and efficiently accelerates graphics through computer hardware such as GPU to generate visual effects. OpenGL ES texture, referred to as 0ES texture, is a graphics rendering technology used on embedded systems. It is used to store image data and is applied to the surface of 3D objects or 2D interfaces during the graphics rendering process. Texture is a key element in graphics for simulating the surface characteristics of objects, providing details such as patterns, colors and gloss for the surface of 3D models. OpenGL ES textures usually work closely with graphics processing units to achieve fast graphics rendering through hardware acceleration. YUV texture is a texture format commonly used in image processing and video encoding. In the YUV color space, the image color is represented by three components: Y brightness, U blue difference, and V red difference. YUV texture is usually used to store image data in video streams, which more effectively represents the human eye's perception of brightness and color. In video encoding and decoding and graphics rendering, YUV texture is widely used, especially in image compression, video playback, and real-time video processing. HEVC encoding, also known as H.265, is a video compression standard that inherits the foundation of H.264 / AVC and provides a more efficient video compression technology. HEVC can compress more efficiently than H.264 / AVC at the same video quality, and can usually achieve a data compression rate of 50%. It improves video compression efficiency by optimizing the encoding process, improving motion compensation, and adding more complex prediction algorithms.

[0108] As one implementation method, the specific steps of obtaining the pre-selected target image information and extracting the initial feature information of the target image information are as follows:

[0109] Determine the selected range of the target image information, perform edge detection on the selected range, extract the features of the main object, and obtain initial feature information. The main object includes any one of the objects occupying the largest area, the object with the largest color difference from other colors, and the object with the highest clarity.

[0110] Specifically, the target image information of the embodiment of the present application can be selected by point selection, frame selection, circle selection, or the corresponding selected area can be selected by the size of pressure, and edge detection can be performed based on the selected range. The main object can be the object with the largest area, the object with the largest color difference from other colors, the object with the highest clarity, or the object with the largest other identification factor. Based on the main object, a corresponding category pointer is identified. The category pointer is specifically one or more of the type of the main object, the color of the main object, the proportion of the area of ​​the main object in the entire frame image, and the area where the main object is located in the entire frame image. The category pointer can be a specific value, a specific category, or an interval range.

[0111] As one implementation method, the first common texture image and the second common texture image are subjected to image recognition processing according to the initial feature information, and the specific steps of updating the identification information according to the image recognition processing result are as follows:

[0112] The first common texture image is subjected to image recognition processing according to the initial feature information, and when the first recognition information of the first common texture image is recognized to be the same as or similar to the initial feature information, the identification information is updated according to the first recognition information.

[0113] The specific steps of performing image recognition processing on the first common texture image and the second common texture image according to the initial feature information and updating the identification information according to the image recognition processing result are as follows:

[0114] The second common texture image is subjected to image recognition processing according to the initial feature information, and when the second recognition information of the second common texture image is recognized to be the same as or similar to the initial feature information, the identification information is updated according to the second recognition information.

[0115] Specifically, after selecting the image recognition result, the identification information is updated, specifically, the corresponding image information is added to the identification information. Through the identification information of the embodiment of the present application, the target texture image in the first ordinary texture image and the second ordinary texture image can be screened out, so as to further process the target texture image.

[0116] As one implementation method, the first filtered texture image and the corresponding first ordinary texture image and the corresponding second ordinary texture image are mixed according to a preset ratio to obtain the first mixed image in the following specific steps:

[0117] Binding the first filtered texture image and the corresponding first common texture image to an OpenGL frame buffer, and mixing the first filtered texture image and the first common texture image in the frame buffer according to a preset ratio to obtain a first mixed image;

[0118] The first filtered texture image and the corresponding second common texture image are bound to the OpenGL frame buffer, and the first filtered texture image and the second common texture image are mixed in the frame buffer according to a preset ratio to obtain a first mixed image.

[0119] Specifically, the first mixed image in this embodiment is formed by combining the result of mixing the first filtered texture image and the first common texture image in a preset ratio and the result of mixing the first filtered texture image and the second common texture image in a preset ratio.

[0120] The present application further processes the first filtered texture image obtained by the first filtering process. The image enhancement process is placed after the decoder is decoded and before the final image is displayed. Specifically, the image data decoded and output by the decoder is first bound to the OpenGL frame buffer, Gaussian filtering and image blending processing are performed in the frame buffer, and the image pixels are efficiently processed concurrently using the GPU. There is a difference between the hardware decoder and the soft decoder when processing images in the frame buffer, which is the issue of the input image format. Since the input image for Gaussian filtering and image blending processing must be in a normal texture format, and the output image of the hardware decoder is in an OES extended texture format, the first step is to convert the OES extended texture into a normal texture in the frame buffer. The output image of the software decoder is a YUV texture, and the first step is to convert the YUV texture format into a normal texture in the frame buffer. After the image textures of the hardware decoder and the software decoder are converted to normal textures, the use of Gaussian filtering and image blending processing is the same process.

[0121] When Gaussian filtering is performed on ordinary texture images, the benefit is that it plays a role in noise suppression. Gaussian filter is a low-pass filter, which is mainly used to smooth images and remove high-frequency noise. Through Gaussian filtering, random noise and small interference in the image can be effectively reduced, so that the overall quality of the image is improved. The second benefit is edge preservation. While smoothing the image, the Gaussian filter can retain larger edge features. Then the smoothed image after Gaussian filtering is mixed with the original image in proportion, usually through linear interpolation. By adjusting the mixing ratio, certain details can be retained while denoising and smoothing. Compared with direct image sharpening, the hybrid method is easier to control noise amplification, better smooth transition and achieve natural visual effects.

[0122] As one implementation manner, the specific steps of mixing the second filtered texture image and the target texture image according to a preset ratio to obtain the second mixed image are:

[0123] The second filtered texture image and the target texture image are bound to the OpenGL frame buffer, and the second filtered texture image and the target texture image are mixed in the frame buffer according to a preset ratio to obtain a second mixed image.

[0124] Specifically, the embodiment of the present application further processes the second filtered texture image obtained by the second filtering process, wherein the second filtering process is specifically a median filtering process, a bilateral filtering process, a convolutional neural network denoising process or a fast Gaussian filtering process. The fast Gaussian filtering process uses an optimization algorithm, such as a separation filter or a FFT fast Fourier transform method, to accelerate the calculation of the original Gaussian filter and reduce its computational complexity. Through optimization technology, the implementation of Gaussian filtering can be accelerated, which is particularly suitable for real-time image processing, thereby reducing the precision loss in the decoding process, improving the accuracy of video stream data decoding, and thus improving image rendering effects.

[0125] The video stream data of the present application is first split, and hardware decoding and software decoding are performed to obtain multiple image frames. During the decoding process of each frame, image feature extraction is performed simultaneously to generate feature information of each frame. After each frame is decoded, the image recognition algorithm is used to match the features of the frame. If the features of the frame match or are similar to the target image selected by the user, the frame is marked as a "target frame". For video frames marked as "target frames", the system uses a more efficient image processing method, such as enhanced filtering or clarity enhancement algorithms for processing. For all frames, regardless of whether they match or not, standard image processing steps such as Gaussian filtering are performed, and the results of hardware decoding and software decoding are mixed to generate the final clear picture. On this basis, users can query the video identifier through retrieval instructions to quickly find frames in the video that are similar to the target image, supporting fast positioning and playback. The enhanced image frames are output with optimized pictures to improve visual quality. By combining image feature extraction and matching mechanisms, video decoding and image processing can achieve more efficient video frame processing. For frames in the video that contain target images, special processing measures can be taken to improve picture quality. With the help of a retrieval mechanism, relevant frames can be quickly located during playback, optimizing the efficiency of video playback and retrieval.

[0126] Each frame of the video stream of this application is decoded and processed, and combined with the image recognition function, after each frame is decoded and processed, such as hardware decoding or software decoding, the image feature extraction and matching of the frame is immediately performed. If a frame is detected to be similar or consistent with the target image selected by the user during the video playback process, the system performs special processing on the frame, such as enhancing the picture clarity or rendering the frame with higher quality.

[0127] When processing each frame of the video stream, this application combines the feature matching mechanism to determine whether the current frame needs to be specially processed. If the current frame matches the target image feature, the system can give priority to using higher-quality decoding and image processing methods. For example, for a specific frame, a higher-quality decoding algorithm is used or the filtering parameters are adjusted to further improve the picture effect.

[0128] The image retrieval mechanism of this application can be used to optimize the decoding and image processing process. For example, when playing a video, when the user triggers a retrieval command, the system can quickly locate the specific position or frame in the video that contains the target image, and then provide more efficient decoding and clarity enhancement for the frame, reducing unnecessary processing and improving overall performance.

[0129] Reference Figure 2 , a video stream decoding system is provided for an embodiment of the present application, the system includes a video stream acquisition module, a target image acquisition module, an identification update module, a first filtering processing module, a first image mixing module, a second filtering processing module, a second image mixing module and an image rendering module.

[0130] A video stream acquisition module is used to acquire video stream data, split the video stream data, and perform hardware decoding and software decoding on the split results, wherein the hardware decoding obtains a first common texture image, and the software decoding obtains a second common texture image;

[0131] A target image acquisition module, used to acquire pre-selected target image information, extract initial feature information of the target image information, and save the initial feature information as identification information;

[0132] An identification updating module, used for performing image recognition processing on the first common texture image and the second common texture image according to the initial feature information, and updating the identification information according to the image recognition processing result;

[0133] A first filtering processing module is used to obtain a target texture image in the first common texture image and the second common texture image according to the identification information, and perform a first filtering process on the first common texture image and the second common texture image after filtering out the target texture image to obtain a first filtered texture image;

[0134] A first image mixing module, used for mixing the first filtered texture image with the corresponding first common texture image and the corresponding second common texture image according to a preset ratio to obtain a first mixed image;

[0135] A second filtering processing module, used for performing a second filtering process on the target texture image to obtain a second filtered texture image;

[0136] A second image mixing module, used for mixing the second filtered texture image and the target texture image according to a preset ratio to obtain a second mixed image;

[0137] The image rendering module is used to perform image rendering according to the first mixed image and the second mixed image.

[0138] As one implementation mode, the identification updating module performs image recognition processing on the first common texture image and the second common texture image according to the initial feature information, and the specific steps of updating the identification information according to the image recognition processing result are as follows:

[0139] The first common texture image is subjected to image recognition processing according to the initial feature information, and when the first recognition information of the first common texture image is recognized to be the same as or similar to the initial feature information, the identification information is updated according to the first recognition information.

[0140] As one implementation mode, the identification updating module performs image recognition processing on the first common texture image and the second common texture image according to the initial feature information, and the specific steps of updating the identification information according to the image recognition processing result are as follows:

[0141] The second common texture image is subjected to image recognition processing according to the initial feature information, and when the second recognition information of the second common texture image is recognized to be the same as or similar to the initial feature information, the identification information is updated according to the second recognition information.

[0142] As one implementation method, the specific steps of the video stream acquisition module splitting the video stream data are as follows:

[0143] Divide the video stream data into several parts according to the preset data size, and extract the encoding format information, frame type information, resolution information and multi-layer encoding information of each video stream data;

[0144] The coding format information, frame type information, resolution information and multi-layer coding situation information are analyzed according to a pre-built decision model to obtain a first split result suitable for hardware decoding or a second split result suitable for software decoding.

[0145] As one implementation method, the specific steps of hardware decoding of the video stream acquisition module to obtain the first common texture image are:

[0146] Convert the first split result into OES texture information, and decode the first split result using a hardware decoding framework decoder to obtain an OES extended texture image;

[0147] The OES extended texture image is converted into a normal texture to obtain a first normal texture image.

[0148] As one implementation method, the specific steps of decoding the video stream acquisition module software to obtain the second ordinary texture image are:

[0149] The second split result is converted into YUV texture information, and the second split result is decoded by a software decoding framework decoder to obtain a YUV extended texture image;

[0150] The YUV extended texture image is converted into a normal texture to obtain a second normal texture image.

[0151] As one implementation method, the target image acquisition module acquires the pre-selected target image information, and the specific steps of extracting the initial feature information of the target image information are as follows:

[0152] Determine the selected range of the target image information, perform edge detection on the selected range, extract the features of the main object, and obtain initial feature information. The main object includes any one of the objects occupying the largest area, the object with the largest color difference from other colors, and the object with the highest clarity.

[0153] As one implementation manner, the first image mixing module mixes the first filtered texture image with the corresponding first ordinary texture image and the corresponding second ordinary texture image according to a preset ratio to obtain the first mixed image in the following specific steps:

[0154] Binding the first filtered texture image and the corresponding first common texture image to an OpenGL frame buffer, and mixing the first filtered texture image and the first common texture image in the frame buffer according to a preset ratio to obtain a first mixed image;

[0155] The first filtered texture image and the corresponding second common texture image are bound to the OpenGL frame buffer, and the first filtered texture image and the second common texture image are mixed in the frame buffer according to a preset ratio to obtain a first mixed image.

[0156] As one implementation manner, the second image mixing module mixes the second filtered texture image and the target texture image according to a preset ratio to obtain the second mixed image in the following specific steps:

[0157] The second filtered texture image and the target texture image are bound to the OpenGL frame buffer, and the second filtered texture image and the target texture image are mixed in the frame buffer according to a preset ratio to obtain a second mixed image.

[0158] Reference Figure 3 An embodiment of the present application provides a video stream decoding device, including a memory and a processor, wherein the memory stores a computer program, and the processor is configured to run the computer program to execute the video stream decoding method as described above.

[0159] An embodiment of the present application provides a computer-readable storage medium, in which a computer program is stored. The computer program is configured to execute the video stream decoding method as described above when running.

[0160] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working process of the above-described device and product can refer to the corresponding process in the aforementioned method embodiment, and will not be repeated here.

[0161] In several embodiments provided in this application, it should be understood that the disclosed methods, systems, devices and program products may be implemented in other ways.

[0162] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of software functional units.

[0163] As described above, the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present application.

Claims

1. A video stream decoding method, characterized in that: include: Acquire video stream data, split the video stream data, perform hardware decoding and software decoding on the split results respectively, the hardware decoding obtains a first common texture image, and the software decoding obtains a second common texture image; Acquire pre-selected target image information, extract initial feature information of the target image information, and save the initial feature information as identification information; Performing image recognition processing on the first common texture image and the second common texture image according to the initial feature information, and updating the identification information according to the image recognition processing result; Obtaining a target texture image in the first common texture image and the second common texture image according to the identification information, performing a first filtering process on the first common texture image and the second common texture image after filtering out the target texture image, to obtain a first filtered texture image; Mixing the first filtered texture image with the corresponding first common texture image and the corresponding second common texture image in a preset ratio to obtain a first mixed image; Performing a second filtering process on the target texture image to obtain a second filtered texture image; Mixing the second filtered texture image and the target texture image according to a preset ratio to obtain a second mixed image; Image rendering is performed according to the first mixed image and the second mixed image.

2. The video stream decoding method according to claim 1, characterized in that: The specific steps of performing image recognition processing on the first common texture image and the second common texture image according to the initial feature information and updating the identification information according to the image recognition processing result are: The first common texture image is subjected to image recognition processing according to the initial feature information, and when the first recognition information of the first common texture image is recognized to be the same as or similar to the initial feature information, the identification information is updated according to the first recognition information.

3. The video stream decoding method according to claim 2, characterized in that: The specific steps of performing image recognition processing on the first common texture image and the second common texture image according to the initial feature information and updating the identification information according to the image recognition processing result are: The second common texture image is subjected to image recognition processing according to the initial feature information, and when the second recognition information of the second common texture image is recognized to be the same as or similar to the initial feature information, the identification information is updated according to the second recognition information.

4. The video stream decoding method according to claim 1, characterized in that: The specific steps of splitting the video stream data are as follows: Divide the video stream data into several parts according to the preset data size, and extract the encoding format information, frame type information, resolution information and multi-layer encoding information of each video stream data; The coding format information, frame type information, resolution information and multi-layer coding situation information are analyzed according to a pre-built decision model to obtain a first split result suitable for hardware decoding or a second split result suitable for software decoding.

5. The video stream decoding method according to claim 4, characterized in that: The specific steps of obtaining the first common texture image by hardware decoding are: Convert the first split result into OES texture information, and decode the first split result using a hardware decoding framework decoder to obtain an OES extended texture image; The OES extended texture image is converted into a normal texture to obtain a first normal texture image.

6. The video stream decoding method according to claim 4, characterized in that: The specific steps of decoding the software to obtain the second ordinary texture image are: The second split result is converted into YUV texture information, and the second split result is decoded by a software decoding framework decoder to obtain a YUV extended texture image; The YUV extended texture image is converted into a normal texture to obtain a second normal texture image.

7. The video stream decoding method according to claim 1, characterized in that: The specific steps of obtaining the pre-selected target image information and extracting the initial feature information of the target image information are as follows: Determine the selected range of the target image information, perform edge detection on the selected range, extract the features of the main object, and obtain initial feature information. The main object includes any one of the objects occupying the largest area, the object with the largest color difference from other colors, and the object with the highest definition.

8. The video stream decoding method according to claim 1, characterized in that: The specific steps of mixing the first filtered texture image with the corresponding first common texture image and the corresponding second common texture image according to a preset ratio to obtain the first mixed image are: Binding the first filtered texture image and the corresponding first common texture image to an OpenGL frame buffer, and mixing the first filtered texture image and the first common texture image in the frame buffer according to a preset ratio to obtain a first mixed image; The first filtered texture image and the corresponding second common texture image are bound to the OpenGL frame buffer, and the first filtered texture image and the second common texture image are mixed in the frame buffer according to a preset ratio to obtain a first mixed image.

9. The video stream decoding method according to claim 1, characterized in that: The specific steps of mixing the second filtered texture image and the target texture image according to a preset ratio to obtain the second mixed image are: The second filtered texture image and the target texture image are bound to the OpenGL frame buffer, and the second filtered texture image and the target texture image are mixed in the frame buffer according to a preset ratio to obtain a second mixed image.

10. A video stream decoding system, characterized in that: include: A video stream acquisition module is used to acquire video stream data, split the video stream data, and perform hardware decoding and software decoding on the split results, wherein the hardware decoding obtains a first common texture image, and the software decoding obtains a second common texture image; A target image acquisition module, used to acquire pre-selected target image information, extract initial feature information of the target image information, and save the initial feature information as identification information; An identification updating module, used for performing image recognition processing on the first common texture image and the second common texture image according to the initial feature information, and updating the identification information according to the image recognition processing result; A first filtering processing module is used to obtain a target texture image in the first common texture image and the second common texture image according to the identification information, and perform a first filtering process on the first common texture image and the second common texture image after filtering out the target texture image to obtain a first filtered texture image; A first image mixing module, used for mixing the first filtered texture image with the corresponding first common texture image and the corresponding second common texture image according to a preset ratio to obtain a first mixed image; A second filtering processing module, used for performing a second filtering process on the target texture image to obtain a second filtered texture image; A second image mixing module, used for mixing the second filtered texture image and the target texture image according to a preset ratio to obtain a second mixed image; The image rendering module is used to perform image rendering according to the first mixed image and the second mixed image.

11. The video stream decoding system according to claim 10, characterized in that: The identification updating module performs image recognition processing on the first common texture image and the second common texture image according to the initial feature information, and the specific steps of updating the identification information according to the image recognition processing result are: The first common texture image is subjected to image recognition processing according to the initial feature information, and when the first recognition information of the first common texture image is recognized to be the same as or similar to the initial feature information, the identification information is updated according to the first recognition information.

12. The video stream decoding system according to claim 11, characterized in that: The identification updating module performs image recognition processing on the first common texture image and the second common texture image according to the initial feature information, and the specific steps of updating the identification information according to the image recognition processing result are: The second common texture image is subjected to image recognition processing according to the initial feature information, and when the second recognition information of the second common texture image is recognized to be the same as or similar to the initial feature information, the identification information is updated according to the second recognition information.

13. The video stream decoding system according to claim 10, characterized in that: The specific steps of the video stream acquisition module for splitting the video stream data are as follows: Divide the video stream data into several parts according to the preset data size, and extract the encoding format information, frame type information, resolution information and multi-layer encoding information of each video stream data; The coding format information, frame type information, resolution information and multi-layer coding situation information are analyzed according to a pre-built decision model to obtain a first split result suitable for hardware decoding or a second split result suitable for software decoding.

14. The video stream decoding system according to claim 13, characterized in that: The specific steps of hardware decoding of the video stream acquisition module to obtain the first common texture image are: Convert the first split result into OES texture information, and decode the first split result using a hardware decoding framework decoder to obtain an OES extended texture image; The OES extended texture image is converted into a normal texture to obtain a first normal texture image.

15. The video stream decoding system according to claim 13, characterized in that: The specific steps of decoding the video stream acquisition module software to obtain the second ordinary texture image are: The second split result is converted into YUV texture information, and the second split result is decoded by a software decoding framework decoder to obtain a YUV extended texture image; The YUV extended texture image is converted into a normal texture to obtain a second normal texture image.

16. The video stream decoding system according to claim 10, characterized in that: The target image acquisition module acquires the pre-selected target image information, and the specific steps of extracting the initial feature information of the target image information are as follows: Determine the selected range of the target image information, perform edge detection on the selected range, extract the features of the main object, and obtain initial feature information. The main object includes any one of the objects occupying the largest area, the object with the largest color difference from other colors, and the object with the highest definition.

17. The video stream decoding system according to claim 10, characterized in that: The first image mixing module mixes the first filtered texture image with the corresponding first common texture image and the corresponding second common texture image according to a preset ratio to obtain the first mixed image in the following specific steps: Binding the first filtered texture image and the corresponding first common texture image to an OpenGL frame buffer, and mixing the first filtered texture image and the first common texture image in the frame buffer according to a preset ratio to obtain a first mixed image; The first filtered texture image and the corresponding second common texture image are bound to the OpenGL frame buffer, and the first filtered texture image and the second common texture image are mixed in the frame buffer according to a preset ratio to obtain a first mixed image.

18. The video stream decoding system according to claim 10, characterized in that: The second image mixing module mixes the second filtered texture image and the target texture image according to a preset ratio to obtain the second mixed image in the following specific steps: The second filtered texture image and the target texture image are bound to the OpenGL frame buffer, and the second filtered texture image and the target texture image are mixed in the frame buffer according to a preset ratio to obtain a second mixed image.

19. A video stream decoding device, characterized in that: The method comprises a memory and a processor, wherein the memory stores a computer program, and the processor is configured to run the computer program to execute the video stream decoding method according to any one of claims 1 to 9.

20. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, wherein the computer program is configured to execute the video stream decoding method according to any one of claims 1 to 9 when running.

Citation Information

Patent Citations

  • Video frame extraction method and device, electronic equipment and computer readable storage medium

    CN111405288A

  • Method for encoding and decoding images, device for encoding and decoding images and corresponding computer programs

    WO2015079179A1