Video encoding method, apparatus, computer device and storage medium

By identifying regions of interest in video frames and employing efficient encoding strategies, the problem of bitrate spikes in high-resolution video transmission is solved, resulting in more stable network transmission and a better user experience.

CN120711176BActive Publication Date: 2026-03-17GUANGZHOU ANYKA MICROELECTRONICS CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202511022368.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-07-24
Publication Date
2026-03-17
Estimated Expiration
2045-07-24

AI Technical Summary

Technical Problem

Traditional video encoders tend to generate superframes containing large amounts of data when processing high-resolution videos, leading to a surge in bitrate during network transmission and impacting network bandwidth and user experience.

Method used

By identifying regions of interest and non-regions of interest in video frames, different encoding strategies are used to encode them, including Intra encoding and Skip/GDR encoding. The encoding algorithm is optimized to smooth data packets, and keyframes are inserted to ensure the integrity of the image.

Benefits of technology

It significantly reduces the risk of jitter and congestion in network transmission, improves the problem of video stuttering in real-time video transmission, and enhances user experience and encoding efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120711176B_ABST
    Figure CN120711176B_ABST
Patent Text Reader

Abstract

The application relates to a video coding method, device, computer equipment and storage medium. The method comprises the following steps: obtaining a region of interest and a non-region of interest in a current video frame of a video stream; using a first coding strategy to code the region of interest, using a second coding strategy to code the non-region of interest, and outputting coding data; wherein the coding quality of the first coding strategy is higher than that of the second coding strategy. The application can effectively control the code rate fluctuation by optimizing the coding algorithm to perform smooth processing on the data packet, thereby reducing the jitter and congestion risk in network transmission, and especially improving the picture freezing problem caused by large I frame transmission delay in the real-time video transmission scene.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of video coding technology, and in particular to a video coding method, apparatus, computer equipment, and storage medium. Background Technology

[0002] With the development of computer technology, video encoding is increasingly widely used in scenarios such as streaming media transmission and remote conferencing. However, when processing high-resolution video, traditional encoders are prone to generating superframes containing large amounts of data, which affects network transmission. Summary of the Invention

[0003] Therefore, it is necessary to provide a video encoding method, apparatus, computer equipment, and storage medium that can optimize network transmission to address the aforementioned technical problems.

[0004] In a first aspect, this application provides a video coding method, the method comprising: obtaining regions of interest and regions of non-interest in the current video frame of a video stream; encoding the regions of interest using a first coding strategy, and encoding the regions of non-interest using a second coding strategy, and outputting encoded data; wherein the coding quality of the first coding strategy is higher than the coding quality of the second coding strategy.

[0005] In one embodiment, the method further includes inserting keyframes into the video stream when it is determined that a global update is needed.

[0006] In one embodiment, the method further includes determining whether a global update is needed based on the differences between video frames in the video stream.

[0007] In one embodiment, obtaining the region of interest and non-region of interest in the current video frame of the video stream includes: performing edge detection processing on the current video frame to divide the region of interest and non-region of interest.

[0008] In one embodiment, obtaining the region of interest and non-region of interest in the current video frame of the video stream includes: inputting the current video frame into the SAM model, segmenting the current video frame using the SAM model, and dividing the region of interest and non-region of interest.

[0009] In one embodiment, the SAM model includes a SAM-SA model and a SAM-SE model; inputting the current video frame into the SAM model includes: determining whether to input the current video frame into the SAM-SA model or the SAM-SE model based on the changes in the video content of consecutive frames in the video stream.

[0010] In one embodiment, the first encoding strategy includes Intra encoding; encoding the region of interest using the first encoding strategy includes: encoding the region of interest into a keyframe using Intra encoding within a preset period.

[0011] In one embodiment, the second encoding strategy includes Skip encoding and GDR encoding; encoding the non-interest region using the second encoding strategy includes: determining whether to encode the non-interest region by Skip encoding or GDR encoding based on the difference between the current video frame and the previous video frame.

[0012] In one embodiment, based on the difference between the current video frame and the previous video frame, it is determined whether to encode the non-interest region by Skip coding or GDR coding, including: if the difference between the current video frame and the previous video frame is small, Skip coding is used; if the difference between the current video frame and the previous video frame is large, GDR coding is used periodically.

[0013] In one embodiment, the difference between the current video frame and the previous video frame includes scene differences.

[0014] Secondly, this application also provides a video encoding apparatus, comprising: a region acquisition module for acquiring regions of interest and regions of non-interest in the current video frame of a video stream; and a region encoding module for encoding the regions of interest using a first encoding strategy and encoding the regions of non-interest using a second encoding strategy, and outputting encoded data; wherein the encoding quality of the first encoding strategy is higher than that of the second encoding strategy.

[0015] Thirdly, this application also provides a computer device, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps of the above-described method.

[0016] Fourthly, this application also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the above-described method.

[0017] Fifthly, this application also provides a computer program product, including a computer program that, when executed by a processor, implements the steps of the above-described method.

[0018] The aforementioned video encoding method, apparatus, computer equipment, and storage medium, after obtaining the region of interest (ROI) and non-ROI regions in the current video frame of the video stream, encode the ROI using a first encoding strategy and encode the non-ROI using a second encoding strategy, outputting encoded data; wherein, the encoding quality of the first encoding strategy is higher than that of the second encoding strategy; this application optimizes the encoding algorithm to smooth data packets, effectively controlling bit rate fluctuations, thereby reducing jitter and congestion risks in network transmission, especially improving the screen stuttering problem caused by large I-frame transmission delay in real-time video transmission scenarios, significantly enhancing the user experience. Attached Figure Description

[0019] To more clearly illustrate the technical solutions in the embodiments of this application or related technologies, the drawings used in the description of the embodiments of this application or related technologies will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.

[0020] Figure 1 This is a flowchart illustrating a video encoding method in one embodiment;

[0021] Figure 2 This is a flowchart illustrating a video encoding method in another embodiment;

[0022] Figure 3 This is a structural block diagram of a video encoding device in one embodiment;

[0023] Figure 4 This is an internal structural diagram of a computer device in one embodiment;

[0024] Figure 5 This is a diagram of the internal structure of a computer device in another embodiment. Detailed Implementation

[0025] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.

[0026] It should be noted that the terms "first," "second," etc., used in this application can be used to describe various elements, but these elements are not limited by these terms. These terms are only used to distinguish the first element from the second element. The terms "comprising" and "having," and any variations thereof, used in this application, are intended to cover non-exclusive inclusion. The term "multiple" used in this application refers to two or more. The term "and / or" used in this application refers to one of the embodiments, or any combination of multiple embodiments.

[0027] Existing video coding standards handle high-resolution superframes using the following methods: ① Slice Processing: This method divides a large frame into multiple smaller blocks (called "slices"). Each slice can be encoded and decoded independently, but the encoding efficiency is low. Boundaries between slices can also reduce efficiency because pixels at the boundaries cannot utilize information from adjacent slices for prediction. The bitrate increases slightly as each slice requires additional header information. ② Tile Technology: This method divides the video frame into multiple rectangular regions (called "tiles"). Each tile can be encoded and decoded independently, but the encoding efficiency is low, and boundaries between tiles can also reduce efficiency, similar to slicing. The complexity increases as tile technology requires more complex encoder and decoder designs, increasing implementation complexity. ③ Hierarchical Coding: This method divides the video frame into multiple layers, each containing image information at different resolutions. For example, a low-resolution version is encoded first, then progressively refined to a high-resolution version. However, the complexity increases as hierarchical coding requires more complex encoder and decoder designs, increasing implementation complexity. Increased latency may occur due to the need for progressive refinement, potentially increasing decoding delay. ④ Progressive Coding: This method divides the video frame encoding process into multiple stages, progressively refining image information. However, it increases complexity, requiring more complex encoder and decoder designs, thus increasing implementation complexity. Increased latency may also occur due to the need for progressive refinement. ⑤ Key Frame Optimization: This method optimizes the encoding strategy of key frames (I-frames) to reduce the amount of data in each key frame. For example, using more efficient encoding modes or reducing redundant information. However, it also increases complexity, requiring more complex encoder designs, thus increasing implementation complexity. Furthermore, the optimization process may increase encoding time, reducing encoding speed.

[0028] In summary, traditional encoders currently have significant drawbacks when processing high-resolution videos: when encoding ultra-high-definition content, they are prone to generating superframes containing large amounts of data, and such sudden large data packets can cause a surge in the instantaneous bitrate.

[0029] To address the issue of super I-frames generated by high-resolution video transmission, reduce bandwidth pressure, further decrease bitrate, and reduce storage space, this application proposes an optimized encoding scheme. By optimizing the encoding algorithm to smooth data packets, bitrate fluctuations can be effectively controlled, thereby reducing jitter and congestion risks in network transmission. The embodiments of this application particularly improve the stuttering problem caused by the transmission delay of large I-frames in real-time video transmission scenarios, significantly enhancing the user experience.

[0030] It should be noted that the beneficial effects or technical problems solved by the embodiments of this application are not limited to this one, but may also be other implicit or related problems. For details, please refer to the description of the embodiments below.

[0031] Before introducing the specific embodiments of this application, the technical terms involved in this application will be explained:

[0032] I-frame: Encoded keyframe.

[0033] QP: Quantization Parameter.

[0034] ROI: Region of Interest.

[0035] P-frame: Non-keyframe.

[0036] GDR: Gradual Decoder Refresh.

[0037] Intra: Intra-frame encoding.

[0038] The technical solution of this application and how the technical solution of this application solves the above-mentioned technical problems are described in detail below with specific embodiments. These specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described again in some embodiments. The embodiments of this application will now be described with reference to the accompanying drawings.

[0039] In one exemplary embodiment, such as Figure 1 As shown, a video encoding method is provided. Taking an application environment with limited network bandwidth or high real-time requirements as an example, the method includes steps 102 to 104. Wherein:

[0040] Step 102: Obtain the region of interest and non-region of interest in the current video frame of the video stream.

[0041] In this context, a video stream refers to a video stream that requires encoding processing. A video stream includes multiple video frames, and the current video frame can refer to the latest video frame in the video stream.

[0042] Specifically, during the video encoding process, the region of interest (ROI) and non-ROI in the current video frame of the video stream can be obtained. For example, by identifying and segmenting different regions, adaptive ROI settings can be achieved. For instance, the division of ROI and non-ROI can be dynamically adjusted according to changes in video content (such as motion detection, scene switching, etc.).

[0043] It is understandable that the sensing area can be called the Region of Interest (ROI), and the non-ROI region can be called the non-ROI region. Regarding the methods for obtaining the ROI and non-ROI regions, for example, based on the feature information extracted by edge detection, the system can initially divide the ROI and non-ROI regions; alternatively, the Segment Anything Model (SAM) algorithms, specifically Segment Anything Mode (SAM-SA) and Segment Everything Mode (SAM-SE), can be used to replace edge detection, thereby extracting segmentation features more efficiently for adaptive ROI setting and encoding optimization.

[0044] It should be noted that the above-mentioned methods for obtaining regions of interest and regions of non-interest can also take other forms, and are not limited to the forms already mentioned in the above embodiments, as long as they can achieve the function of identifying and segmenting different regions.

[0045] Step 104: Encode the region of interest using a first coding strategy and encode the region of non-interest using a second coding strategy, and output the encoded data; wherein the coding quality of the first coding strategy is higher than that of the second coding strategy.

[0046] Specifically, given the regions of interest and non-regions of interest, a first encoding strategy can be used to encode the regions of interest, and a second encoding strategy can be used to encode the non-regions of interest, outputting encoded data. In this embodiment, the encoding quality of the first encoding strategy is higher than that of the second encoding strategy.

[0047] For example, the Region of Interest (ROI) typically contains important information in the video and requires higher-quality encoding, while the Region of Non-Interest (Non-ROI) can be encoded with lower-quality encoding. This application encodes the ROI using a first encoding strategy with higher encoding quality than the second encoding strategy to ensure the integrity and high quality of the region, maintaining high quality and random access capability while avoiding the large amount of data redundancy found in traditional I-frames. Furthermore, by employing the second encoding strategy to encode the Non-ROI, the encoding complexity and data volume are further reduced. The data volume in the Non-ROI region can be significantly reduced, while maintaining the overall coherence of the image and avoiding the large amount of data redundancy found in traditional I-frames.

[0048] This application embodiment performs high-quality encoding of the ROI (Region of Interest) area, effectively replacing redundant data in traditional ultra-large frames (such as I-frames). This not only ensures high-quality transmission of critical information but also reduces the data volume of key frames, further optimizing encoding efficiency.

[0049] It should be noted that the first and second encoding strategies can adopt corresponding encoding methods, such as Intra encoding, Skip encoding, etc., or other forms, not limited to those mentioned in the above embodiments, as long as they can achieve the function of encoding the ROI region and the non-ROI region with different qualities respectively.

[0050] The aforementioned video encoding method significantly reduces the encoding bitrate and improves video transmission efficiency while maintaining video quality. It is particularly suitable for video applications with limited network bandwidth or high real-time requirements. Furthermore, it eliminates the need for additional chip design and extra encoding header information, avoiding the development costs and technical difficulties associated with complex encoder and decoder designs. Specifically, by eliminating the need for additional hardware design, optimizing ROI region encoding, and employing appropriate encoding strategies, it achieves a dual improvement in both high-efficiency encoding and visual quality. This not only reduces encoding complexity and cost but also significantly improves encoding efficiency, making it particularly suitable for scenarios with high requirements for visual quality and bitrate control, such as video conferencing, real-time monitoring, and environments with limited network bandwidth.

[0051] In one embodiment, the method further includes inserting keyframes into the video stream when it is determined that a global update is needed. Specifically, when it is determined that a global update is needed, embodiments of this application propose inserting keyframes into the video stream. By inserting keyframes when needed, the information of the entire frame can be quickly recovered, ensuring the integrity and continuity of the video stream.

[0052] For example, scene transition detection can be used to determine whether a global update is needed. For instance, if a global content change is detected, a global update is deemed necessary. Optionally, regarding keyframe insertion, a complete I-frame can be inserted when a scene transition or global content change is detected, performing a global update on the entire screen. Furthermore, regarding the insertion position of the I-frame, it can be done before encoding, generating a keyframe whenever a change is detected.

[0053] The video encoding method described above will flexibly insert keyframes when it detects that the entire screen needs to be updated, so as to achieve a global refresh of the screen. This will allow for the rapid restoration of complete screen information when dynamic scene switching or global content changes, ensuring the smoothness and integrity of the video stream.

[0054] In one embodiment, the method may further include determining whether a global update is needed based on differences between video frames in the video stream. Specifically, global content changes can be detected by analyzing differences between video frames (such as histogram differences, segmentation mask changes, etc.).

[0055] It is understood that the above determination of whether a global update is needed can also take other forms, not limited to those mentioned in the above embodiments, as long as they can achieve the function of scene switching detection.

[0056] In some embodiments, obtaining the region of interest (ROI) and non-ROI regions in the current video frame of the video stream includes: performing edge detection processing on the current video frame to divide the ROI and non-ROI regions. Specifically, during video encoding, edge detection feature extraction can be introduced to accurately identify and segment different regions, thereby achieving adaptive ROI setting. For example, based on the feature information extracted by edge detection, the system can initially divide the ROI region and non-ROI region.

[0057] In one embodiment, obtaining the region of interest (ROI) and non-ROI regions in the current video frame of a video stream includes: inputting the current video frame into a SAM model, segmenting the current video frame using the SAM model, and identifying the ROI and non-ROI regions. Specifically, the SAM algorithm can be used for segmentation feature extraction. For example, inputting the current video frame into a SAM model, segmenting the current video frame using the SAM model, and identifying the ROI and non-ROI regions.

[0058] The aforementioned video encoding method, by accurately identifying and optimizing regions of interest (ROIs), significantly reduces the overall bitrate while maintaining clarity and detail in important areas. This adaptive encoding strategy fully utilizes the characteristics of video content, achieving a dual improvement in encoding efficiency and visual quality.

[0059] In one embodiment, the SAM model includes a SAM-SA model and a SAM-SE model; inputting the current video frame into the SAM model includes: determining whether to input the current video frame into the SAM-SA model or the SAM-SE model based on the changes in the video content of consecutive frames in the video stream.

[0060] Specifically, in the video encoding process, SAM-SA and SAM-SE algorithms can be used to replace traditional edge detection, thereby extracting segmentation features more efficiently for adaptive region of interest (ROI) setting and encoding optimization.

[0061] The SAM-SA model can be understood as a Segment Anything Mode algorithm that combines SAM. SAM-SA is a flexible segmentation mode that can generate a corresponding segmentation mask based on input cues (such as points, bounding boxes, and masks). This mode is suitable for scenarios requiring precise segmentation based on specific cues. In practical applications, in video encoding, each frame can be input into the SAM-SA model, and by specifying foreground and background cues, a precise segmentation mask is generated, thereby distinguishing between ROI and non-ROI regions.

[0062] Furthermore, SAM-SE can be understood as a Segment Everything Mode algorithm that combines SAM. SAM-SE is a fully automatic segmentation mode capable of thoroughly exploring and segmenting all potential objects in the input image, outputting a large number of mask proposals. This mode is suitable for scenarios requiring automatic identification and segmentation of all objects in an image. In practical applications, in video coding, SAM-SE can automatically identify all objects in each frame and generate corresponding segmentation masks. These masks can be directly used to divide ROI (Real Area of ​​Interest) and non-ROI regions.

[0063] This application enables adaptive ROI region setting. For example, based on changes in the video content of consecutive frames in a video stream, it determines whether to input the current video frame into the SAM-SA model or the SAM-SE model. Exemplarily, in SAM-based ROI partitioning, the segmentation mask generated by SAM-SA or SAM-SE can accurately divide the ROI and non-ROI regions in each frame. ROI regions typically contain important information in the video and require higher-quality encoding, while non-ROI regions can use lower-quality encoding. This allows for dynamic adjustment: the partitioning of ROI and non-ROI regions can be dynamically adjusted based on changes in video content (such as motion detection, scene transitions, etc.). For example, when a new object is detected entering the frame, the segmentation mask is regenerated using SAM-SE, updating the ROI region.

[0064] In summary, by combining the Segment Anything Mode and SegmentEverything Mode algorithms of SAM, this application's embodiments can achieve more accurate segmentation feature extraction and adaptive ROI region setting. This strategy not only significantly reduces the encoding bitrate but also improves video transmission efficiency, making it particularly suitable for video application environments with limited network bandwidth or high real-time requirements.

[0065] Regarding the first and second encoding strategies in the embodiments of this application, in one embodiment, the first encoding strategy includes Intra encoding; encoding the region of interest using the first encoding strategy includes: encoding the region of interest into a keyframe through Intra encoding within a preset period.

[0066] Specifically, the Region of Interest (ROI) can be encoded using the first encoding strategy, whereby the ROI can be encoded into keyframes via Intra encoding within a preset period. For the defined ROI, a periodic Intra encoding strategy is adopted. This strategy aims to replace the traditional GOP (Group of Pictures) refresh cycle for keyframes by using P-frames instead of I-frames, effectively reducing the amount of data per frame and thus significantly reducing encoding redundancy while ensuring the integrity of key information.

[0067] For example, encoding the Region of Interest (ROI) into keyframes via Intra encoding within a preset period can be understood as periodic Intra encoding of the ROI region; that is, a periodic Intra encoding strategy is adopted for the defined ROI region. The preset period can refer to a set period, and within the set period (such as every 10 or 20 frames), the ROI region is encoded into I-frames to ensure the integrity and high quality of the region. In this embodiment, through periodic Intra encoding, the ROI region can maintain high quality and random access capability, while avoiding the large amount of data redundancy inherent in traditional I-frames.

[0068] The embodiments of this application propose periodic Intra encoding and ultra-large frame optimization. By adopting the periodic Intra encoding strategy, this application performs high-quality encoding of the ROI region, effectively replacing the redundant data of traditional ultra-large frames (such as I-frames). This not only ensures high-quality transmission of key information, but also reduces the data volume of key frames, further optimizing encoding efficiency.

[0069] The aforementioned video encoding method achieves a dual improvement in both high-efficiency encoding and visual quality by eliminating the need for additional hardware design, optimizing ROI region encoding, and employing a periodic intra encoding strategy. This not only reduces encoding complexity and cost but also significantly improves encoding efficiency, making it particularly suitable for scenarios with high requirements for visual quality and bitrate control, such as video conferencing, real-time monitoring, and environments with limited network bandwidth.

[0070] In an exemplary embodiment, the second encoding strategy includes Skip encoding and GDR encoding; encoding the non-interest region using the second encoding strategy includes: determining whether to encode the non-interest region by Skip encoding or GDR encoding based on the difference between the current video frame and the previous video frame.

[0071] Specifically, embodiments of this application can implement Skip encoding and GDR encoding for non-ROI regions. Skip encoding refers to performing Skip encoding in non-ROI regions, skipping the encoding process and directly referencing data from the previous frame. GDR encoding refers to performing GDR encoding on non-ROI regions to update the data in those regions. GDR encoding avoids the significant data redundancy of traditional I-frames by dividing the I-frame into multiple parts and distributing them across multiple P-frames. Optionally, GDR encoding can be periodic GDR encoding.

[0072] For example, based on the difference between the current video frame and the previous video frame, it can be determined whether to encode the non-ROI region using Skip encoding or GDR encoding. In this embodiment, by using Skip encoding and GDR encoding, the amount of data in the non-ROI region can be significantly reduced while maintaining the overall continuity of the image.

[0073] In one embodiment, based on the difference between the current video frame and the previous video frame, it is determined whether to encode the non-interest region by Skip coding or GDR coding, including: if the difference between the current video frame and the previous video frame is small, Skip coding is used; if the difference between the current video frame and the previous video frame is large, GDR coding is used periodically.

[0074] Specifically, Skip encoding is used only when the difference between the current frame (i.e., the current video frame) and the previous frame (i.e., the previous video frame) is small. If the difference between the current frame and the previous frame is large, Skip encoding is not used, and only GDR encoding is used periodically.

[0075] For non-ROI areas, the aforementioned video encoding method can employ Skip encoding to further reduce encoding complexity and data volume. Furthermore, periodic GDR (Geometric Rendering) is implemented to ensure the periodic updating of non-ROI area data, guaranteeing the stability and consistency of overall image quality.

[0076] In one embodiment, the difference between the current video frame and the previous video frame includes scene differences. Specifically, the difference between the current video frame and the previous video frame may include scene differences. For example, this application can determine scene differences using a video scene dynamic difference grading standard.

[0077] Regarding the video scene dynamic difference grading standard, optionally, high-difference scenes include dynamic content with significant pixel-level changes, specifically characterized by: abrupt changes in detail (such as rapid object displacement), dynamic changes in illumination intensity (brightness fluctuations exceeding 20%), and high-frequency motion textures (such as swaying leaves, water ripples, etc.). Low-difference scenes may include, but are not limited to, static objects (buildings, fixed devices, etc.), still images under stable lighting conditions, and pseudo-static scenes with motion amplitude below the pixel threshold.

[0078] For example, inter-frame difference detection techniques, such as visual metrics, can also be used, including but not limited to: Sum of Squared Differences (SSD): calculating the Euclidean distance between pixels (L2); Mean Squared Error (MSE): establishing an energy function model for inter-frame differences; and Structural Similarity (SSIM): simulating the perceptual evaluation of the human visual system. Furthermore, spatiotemporal feature analysis methods, such as 3D Convolutional Neural Networks (3D-CNN): extracting joint spatiotemporal features; optical flow field estimation: constructing pixel-level motion vector fields; and motion saliency detection: identifying key regions based on attention mechanisms.

[0079] This application's embodiments utilize optimized encoding algorithms to smooth data packets, effectively controlling bitrate fluctuations and thus reducing jitter and congestion risks during network transmission. In particular, this application improves the video stuttering problem caused by large I-frame transmission delays in real-time video transmission scenarios, significantly enhancing the user experience.

[0080] To further illustrate the scheme of this application, a specific example is provided below, such as... Figure 2 As shown, edge detection is used to obtain the region of interest (ROI) and non-ROI regions in the current video frame of the video stream. This application embodiment comprehensively utilizes edge detection feature extraction, adaptive ROI setting, periodic intracoding, skip encoding, and GDR encoding strategies. This significantly reduces the encoding bitrate and improves video transmission efficiency while ensuring video quality, making it particularly suitable for video application environments with limited network bandwidth or high real-time requirements.

[0081] It should be understood that although the steps in the flowcharts of the embodiments described above are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the embodiments described above may include multiple steps or multiple stages. These steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the steps or stages in other steps. It is understood that the steps in different embodiments can be freely combined as needed, and all non-contradictory solutions formed by such combinations are within the scope of protection of this application.

[0082] Based on the same inventive concept, this application also provides a video encoding apparatus for implementing the video encoding method described above. The solution provided by this apparatus is similar to the implementation described in the above method; therefore, the specific limitations in one or more video encoding apparatus embodiments provided below can be found in the limitations of the video encoding method described above, and will not be repeated here.

[0083] In one exemplary embodiment, such as Figure 3 As shown, a video encoding apparatus is provided, the apparatus comprising:

[0084] The region acquisition module 901 is used to acquire the region of interest and non-region of interest in the current video frame of the video stream.

[0085] The region coding module 902 is used to encode the region of interest using a first coding strategy and to encode the region of non-interest using a second coding strategy, and to output coded data; wherein the coding quality of the first coding strategy is higher than that of the second coding strategy.

[0086] In one embodiment, the apparatus further includes a frame insertion module for inserting keyframes into the video stream when it is determined that a global update is needed.

[0087] In one embodiment, the apparatus further includes a global update module for determining whether a global update is needed based on the differences between video frames in the video stream.

[0088] In one embodiment, the region acquisition module 901 is used to perform edge detection processing on the current video frame to divide it into regions of interest and regions of non-interest.

[0089] In one embodiment, the region acquisition module 901 is used to input the current video frame into the SAM model, and to segment the current video frame by the SAM model to divide it into regions of interest and regions of non-interest.

[0090] In one embodiment, the SAM model includes a SAM-SA model and a SAM-SE model; the region acquisition module 901 is used to determine whether to input the current video frame into the SAM-SA model or the SAM-SE model based on the changes in the video content of consecutive frames in the video stream.

[0091] In one embodiment, the first encoding strategy includes Intra encoding; the region encoding module 902 is used to encode the region of interest into a keyframe through Intra encoding within a preset period.

[0092] In one embodiment, the second encoding strategy includes Skip encoding and GDR encoding; the region encoding module 902 is used to determine whether to encode the non-interest region by Skip encoding or GDR encoding based on the difference between the current video frame and the previous video frame.

[0093] In one embodiment, the region coding module 902 is configured to use Skip coding if the difference between the current video frame and the previous video frame is small, and to periodically use GDR coding if the difference between the current video frame and the previous video frame is large.

[0094] In one embodiment, the difference between the current video frame and the previous video frame includes scene differences.

[0095] Each module in the aforementioned video encoding device can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in hardware within or independently of the processor in a computer device, or stored in software within the memory of the computer device, so that the processor can call and execute the operations corresponding to each module.

[0096] In one exemplary embodiment, a computer device is provided, which may be a server, and its internal structure diagram may be as follows: Figure 4As shown, this computer device includes a processor, memory, input / output interfaces (I / O), and a communication interface. The processor, memory, and I / O interfaces are connected via a system bus, and the communication interface is also connected to the system bus via the I / O interfaces. The processor provides computational and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system, computer programs, and a database. The internal memory provides the environment for the operation of the operating system and computer programs stored in the non-volatile storage media. The database stores model data. The I / O interfaces are used for exchanging information between the processor and external devices. The communication interface is used for communication with external terminals via a network connection. When the computer program is executed by the processor, it implements a video encoding method.

[0097] In one exemplary embodiment, a computer device is provided, which may be a terminal, and its internal structure diagram may be as follows: Figure 5 As shown, the computer device includes a processor, memory, input / output interfaces, a communication interface, a display unit, and an input device. The processor, memory, and input / output interfaces are connected via a system bus, and the communication interface, display unit, and input device are also connected to the system bus via the input / output interfaces. The processor provides computing and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs stored in the non-volatile storage media. The input / output interfaces are used for exchanging information between the processor and external devices. The communication interface is used for wired or wireless communication with external terminals; wireless communication can be achieved through Wi-Fi, mobile cellular networks, Near Field Communication (NFC), or other technologies. When the computer program is executed by the processor, it implements a video encoding method. The display unit is used to form a visually visible image and can be a display screen, a projection device, or a virtual reality imaging device. The display screen can be an LCD screen or an e-ink screen. The input device of the computer device can be a touch layer covering the display screen, or buttons, trackballs, or touchpads set on the casing of the computer device, or external keyboards, touchpads, or mice, etc.

[0098] Those skilled in the art will understand that Figure 4 , Figure 5The structure shown is only a block diagram of a part of the structure related to the present application and does not constitute a limitation on the computer device on which the present application is applied. The specific computer device may include more or fewer components than shown in the figure, or combine certain components, or have different component arrangements.

[0099] In one embodiment, a computer device is also provided, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps in the above method embodiments.

[0100] In one embodiment, a computer-readable storage medium is provided having a computer program stored thereon that, when executed by a processor, implements the steps in the above method embodiments.

[0101] In one embodiment, a computer program product is provided, including a computer program that, when executed by a processor, implements the steps in the above method embodiments.

[0102] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. Any references to memory, databases, or other media used in the embodiments provided in this application can include at least one of non-volatile memory and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take many forms, such as Static Random Access Memory (SRAM) or Dynamic Random Access Memory (DRAM). The databases involved in the embodiments provided in this application may include at least one type of relational database and non-relational database. Non-relational databases may include, but are not limited to, blockchain-based distributed databases. The processors involved in the embodiments provided in this application may be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, quantum computing-based data processing logic devices, artificial intelligence (AI) processors, etc., and are not limited to these.

[0103] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this application.

[0104] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of this patent application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this application should be determined by the appended claims.

Claims

1. A method of video coding, the method comprising: The method comprises: obtaining a region of interest and a non-region of interest in a current video frame of a video stream; encoding the region of interest using a first encoding strategy and encoding the non-region of interest using a second encoding strategy, and outputting encoded data; wherein the encoding quality of the first encoding strategy is higher than that of the second encoding strategy; the second encoding strategy comprises Skip encoding and GDR encoding; encoding the non-region of interest using the second encoding strategy comprises: determining whether to encode the non-region of interest using the Skip encoding or the GDR encoding according to the difference between the current video frame and a previous video frame of the current video frame.

2. The method of claim 1, wherein, Before outputting the encoded data, the method further comprises: when it is determined that global updating is needed, inserting a key frame in the video stream.

3. The method of claim 2, wherein, The method further comprises: determining whether global updating is needed according to the difference between video frames in the video stream.

4. The method of claim 1, wherein, Obtaining a region of interest and a non-region of interest in a current video frame of a video stream comprises: performing edge detection processing on the current video frame to divide the region of interest and the non-region of interest.

5. The method of claim 1, wherein, Obtaining a region of interest and a non-region of interest in a current video frame of a video stream comprises: inputting the current video frame into a SAM model, and dividing the region of interest and the non-region of interest by using the SAM model to segment the current video frame.

6. The method of claim 5, wherein, The SAM model comprises a SAM-SA model and a SAM-SE model; inputting the current video frame into the SAM model comprises: determining whether to input the current video frame into the SAM-SA model or the SAM-SE model according to the change of the content of the continuous frames in the video stream.

7. The method according to any one of claims 1 to 6, characterized in that, The first encoding strategy comprises Intra encoding; encoding the region of interest using the first encoding strategy comprises: encoding the region of interest into a key frame by using the Intra encoding within a preset period.

8. The method according to any one of claims 1 to 6, characterized in that, Determining whether to encode the non-region of interest using the Skip encoding or the GDR encoding according to the difference between the current video frame and a previous video frame of the current video frame comprises: if the difference between the current video frame and the previous video frame is small, using the Skip encoding; if the difference between the current video frame and the previous video frame is large, using the GDR encoding periodically.

9. The method according to any one of claims 1 to 6, characterized in that, The difference between the current video frame and the previous video frame comprises a scene difference.

10. A video encoding apparatus, comprising: The device comprises: a region obtaining module configured to obtain a region of interest and a non-region of interest in a current video frame of a video stream; a region encoding module configured to encode the region of interest using a first encoding strategy and encode the non-region of interest using a second encoding strategy, and output encoded data; wherein the encoding quality of the first encoding strategy is higher than that of the second encoding strategy; The second encoding strategy includes Skip encoding and GDR encoding; and the region encoding module is configured to determine to encode the non-interest region by the Skip encoding or the GDR encoding according to a difference between the current video frame and a previous video frame of the current video frame. 11.A computer device, comprising a memory and a processor, wherein the memory stores a computer program, and the computer device is configured to perform the method according to any one of claims 1-10 when the computer program is executed by the processor. The computer program is executed by the processor to implement the steps of the method of any one of claims 1 to 9.

12. A computer readable storage medium having stored thereon a computer program, characterized in that, The computer program is executed by the processor to implement the steps of the method of any one of claims 1 to 9.

13. A computer program product comprising a computer program, characterized in that, The computer program is executed by the processor to implement the steps of the method of any one of claims 1 to 9. The computer program is executed by the processor to implement the steps of the method of any one of claims 1 to 9.

Citation Information

Patent Citations

  • Adaptive frame skipping intelligent encoding method

    CN108833915A

  • Method and system for encoding and decoding video data in conjunction with performing search

    CN116033171A

  • Coding method, real-time communication method and device, equipment and storage medium

    CN116567228A