Method for determining Lagrange multiplier of coding frame in video stream and coding method

By receiving the initial video stream during video encoding, determining the target global quantization parameters and loss ratio, and dynamically updating the Lagrange multipliers, the problem of Lagrange multiplier settings not adapting to changes in video content in existing technologies is solved, thus improving encoding efficiency and adaptability.

CN121644819APending Publication Date: 2026-03-10RONG MING MICROELECTRONICS (JINAN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-01
Publication Date
2026-03-10

AI Technical Summary

Technical Problem

In existing video coding technologies, the Lagrange multiplier parameter settings lack adaptability and are difficult to adjust dynamically, resulting in low video content adaptability and coding efficiency, as well as high computational complexity.

Method used

By receiving the initial video stream, the target global quantization parameters are determined, precoding is performed to obtain coding statistics, the target loss ratio is calculated, and the Lagrange multipliers are updated to adapt to changes in video content. The optimal Lagrange multipliers are dynamically matched using a statistical coding data-driven approach.

Benefits of technology

It achieves dynamic adjustment of Lagrange multipliers without increasing computational complexity, thereby improving the coding efficiency and video content adaptability of intra-coded frames, and optimizing video coding quality and bitrate.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121644819A_ABST
    Figure CN121644819A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a method for determining a Lagrange multiplier of a coding frame in a video stream and a coding method. The method comprises the following steps: receiving an initial video stream; determining target global quantization parameters respectively corresponding to the plurality of initial frames; controlling to use the corresponding target global quantization parameter to pre-code the corresponding initial frame to obtain coding statistical data; determining a target loss ratio according to the loss values corresponding to the plurality of reconstructed frames; and according to the target loss ratio, updating a predetermined Lagrange multiplier to obtain a target Lagrange multiplier corresponding to the intra-frame coding frame. Particularly, the method is especially suitable for the video stream with the fixed intra-frame coding frame interval, the reference function of the frame in the scene is more stable, the updated Lagrange multiplier can be more accurately matched with the coding requirement of the intra-frame coding frame under the fixed interval, and the problems that in the related technology, the Lagrange multiplier is difficult to dynamically adjust, and the coding efficiency is poor are solved. And the video content adaptability and the coding efficiency are difficult to balance.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] Embodiments of the present application relate to the technical field of video coding, and particularly relate to a method for determining a Lagrange multiplier of a coded frame in a video stream and a coding method. BACKGROUND

[0002] With the explosive growth of digital video applications, video coding technology has become the core technology of modern multimedia communication. From the early H.261, MPEG-1 standard to the current H.266 / VVC (Versatile Video Coding), video coding technology is constantly optimized in terms of compression efficiency, coding quality and computational complexity. Among them, rate-distortion optimization (RDO) as the core technology of modern video encoder, through the best balance between the coding rate (Rate) and the distortion degree (Distortion), the maximization of coding efficiency is realized, so it is necessary to reasonably set the Lagrange multiplier (lambda) parameter.

[0003] The existing scheme such as hierarchical Lambda setting strategy, usually according to the importance level of frame (such as I frame, P frame, B frame), sets different Lambda parameters for different types of frames, and adjusts the basic Lambda value according to the frame type. However, the adjustment strategy is relatively fixed, and lacks adaptability, and is poor in adaptability to changes in video content. Or usually need double encoding, resulting in increased computational complexity, there is a problem of low coding efficiency of coded frame. It can be seen that in the prior art, it is difficult to dynamically adjust the Lagrange multiplier, and there is a technical problem that it is difficult to balance the adaptability of video content and coding efficiency. SUMMARY

[0004] Embodiments of the present application provide a method for determining a Lagrange multiplier of a coded frame in a video stream and a coding method, to solve the technical problem that in related technologies, it is difficult to dynamically adjust the Lagrange multiplier, so as to balance the adaptability of video content and coding efficiency.

[0005] In a first aspect, an embodiment of the present application provides a method for determining a Lagrange multiplier of a coded frame in a video stream, comprising: receiving an initial video stream, wherein the initial video stream comprises a plurality of initial frames, and the plurality of initial frames comprises an intra-coded frame and other frames; determining target global quantization parameters corresponding to the plurality of initial frames respectively; controlling pre-encoding of the corresponding initial frames using the corresponding target global quantization parameters to obtain coding statistics, wherein the coding statistics comprise loss values corresponding to the plurality of reconstructed frames respectively; determining a target loss ratio value according to the loss values corresponding to the plurality of reconstructed frames respectively, wherein the target loss ratio value is a ratio of a first loss value and a second loss value, the first loss value is the loss value corresponding to the intra-coded frame, and the second loss value is the loss value corresponding to the other frames; updating a predetermined Lagrange multiplier according to the target loss ratio value to obtain a target Lagrange multiplier corresponding to the intra-coded frame, wherein the predetermined Lagrange multiplier comprises one of the following: a historical Lagrange multiplier, a reference Lagrange multiplier.

[0006] In a second aspect, an embodiment of the present application provides a method for coding a coded frame in a video stream, comprising: receiving an initial video stream, wherein the initial video stream comprises a plurality of initial frames, and the plurality of initial frames comprises an intra-coded frame and other frames; determining target global quantization parameters corresponding to the plurality of initial frames respectively; controlling pre-encoding of the corresponding initial frames using the corresponding target global quantization parameters to obtain coding statistics, wherein the coding statistics comprise loss values corresponding to the plurality of reconstructed frames respectively; determining a target loss ratio value according to the loss values corresponding to the plurality of reconstructed frames respectively, wherein the target loss ratio value is a ratio of a first loss value and a second loss value, the first loss value is the loss value corresponding to the intra-coded frame, and the second loss value is the loss value corresponding to the other frames; updating a predetermined Lagrange multiplier according to the target loss ratio value to obtain a target Lagrange multiplier corresponding to the intra-coded frame, wherein the predetermined Lagrange multiplier comprises one of the following: a historical Lagrange multiplier, a reference Lagrange multiplier; determining a target coding mode corresponding to the intra-coded frame according to the target Lagrange multiplier corresponding to the intra-coded frame; and coding the intra-coded frame according to the target coding mode corresponding to the intra-coded frame to output a single-frame coded stream corresponding to the intra-coded frame.

[0007] In a third aspect, an embodiment of the present application provides a device for determining a Lagrange multiplier of a coded frame in a video stream, comprising: a first receiving module configured to receive an initial video stream, wherein the initial video stream comprises a plurality of initial frames, and the plurality of initial frames comprise an intra-coded frame and other frames; a first determining module configured to determine target global quantization parameters corresponding to the plurality of initial frames respectively; a first control module configured to control pre-encoding of the corresponding initial frames using the corresponding target global quantization parameters to obtain coding statistics, wherein the coding statistics comprise loss values corresponding to a plurality of reconstructed frames respectively; a second determining module configured to determine a target loss ratio value according to the loss values corresponding to the plurality of reconstructed frames respectively, wherein the target loss ratio value is a ratio of a first loss value and a second loss value, the first loss value is the loss value corresponding to the intra-coded frame, and the second loss value is the loss value corresponding to the other frames; and a first updating module configured to update a predetermined Lagrange multiplier according to the target loss ratio value to obtain a target Lagrange multiplier corresponding to the intra-coded frame, wherein the predetermined Lagrange multiplier comprises one of a historical Lagrange multiplier and a reference Lagrange multiplier.

[0008] In a fourth aspect, an embodiment of the present application provides a device for coding a coded frame in a video stream, comprising: a second receiving module configured to receive an initial video stream, wherein the initial video stream comprises a plurality of initial frames, and the plurality of initial frames comprise an intra-coded frame and other frames; a third determining module configured to determine target global quantization parameters corresponding to the plurality of initial frames respectively; a second control module configured to control pre-encoding of the corresponding initial frames using the corresponding target global quantization parameters to obtain coding statistics, wherein the coding statistics comprise loss values corresponding to a plurality of reconstructed frames respectively; a fourth determining module configured to determine a target loss ratio value according to the loss values corresponding to the plurality of reconstructed frames respectively, wherein the target loss ratio value is a ratio of a first loss value and a second loss value, the first loss value is the loss value corresponding to the intra-coded frame, and the second loss value is the loss value corresponding to the other frames; a second updating module configured to update a predetermined Lagrange multiplier according to the target loss ratio value to obtain a target Lagrange multiplier corresponding to the intra-coded frame, wherein the predetermined Lagrange multiplier comprises one of a historical Lagrange multiplier and a reference Lagrange multiplier; a fifth determining module configured to determine a target coding mode corresponding to the intra-coded frame according to the target Lagrange multiplier corresponding to the intra-coded frame; and a coding module configured to code the intra-coded frame according to the target coding mode corresponding to the intra-coded frame to output a single-frame coded stream corresponding to the intra-coded frame.

[0009] Fifthly, embodiments of this application provide a computing device, including a processing component and a storage component; the storage component stores one or more computer instructions; the one or more computer instructions are to be invoked and executed by the processing component to implement the method as described in any of the foregoing descriptions.

[0010] Sixthly, embodiments of this application provide a computer-readable storage medium having a computer program stored thereon, which, when executed by a processing component, implements the method as described in any of the foregoing descriptions.

[0011] In a seventh aspect, embodiments of this application provide a computer program product, including a computer program / instructions, which, when executed by a processing component, implement the method as described above.

[0012] In this embodiment, an initial video stream is received, comprising multiple initial frames, including intra-coded frames and other frames; target global quantization parameters corresponding to the multiple initial frames are determined; the corresponding initial frames are pre-coded using the target global quantization parameters to obtain coding statistics, wherein the coding statistics include loss values ​​corresponding to the multiple reconstructed frames; a target loss ratio is determined based on the loss values ​​corresponding to the multiple reconstructed frames, wherein the target loss ratio is the ratio of a first loss value and a second loss value, the first loss value being the loss value corresponding to the intra-coded frames and the second loss value being the loss value corresponding to other frames; and a predetermined Lagrange multiplier is updated based on the target loss ratio to obtain a target Lagrange multiplier corresponding to the intra-coded frames, wherein the predetermined Lagrange multiplier includes one of the following: a historical Lagrange multiplier and a reference Lagrange multiplier. By employing a statistical coding data-driven approach, this method obtains multi-frame loss values ​​through precoding and calculates the target loss ratio between intra-coded frames and other frames. This achieves the goal of dynamically matching the optimal Lagrange multiplier (Lambda) for intra-coded frames, thereby improving the coding efficiency of intra-coded frames. Since the target loss ratio directly reflects the distortion correlation characteristics between intra-coded frames and other frames under different video content, the update of the Lagrange multiplier always fits the actual content characteristics of the current video without relying on fixed empirical parameters. At the same time, lightweight precoding focuses on statistical data collection, avoiding redundant calculations caused by dual coding. It achieves accurate adaptation of parameters and content without increasing complexity, thus solving the technical problem in related technologies where it is difficult to dynamically adjust the Lagrange multiplier, making it difficult to balance video content adaptability and coding efficiency.

[0013] It should be noted that, in particular, this method, while applicable to various video streams, is especially suitable for video streams with fixed intra-coded frame intervals, such as live streams, achieving a better balance between video content adaptability and coding efficiency. This is because the reference base of intra-coded frames is more stable in such scenarios, and dynamically updated Lagrange multipliers can more accurately match the coding requirements of intra-coded frames at fixed intervals. When subsequent frames also need coding, it can better leverage the reference gain of optimized intra-coded frames for subsequent frames, thus more efficiently balancing video content adaptability and coding efficiency. In other words, it more effectively and efficiently solves the technical problem in related technologies where it is difficult to dynamically adjust Lagrange multipliers, leading to a difficulty in balancing video content adaptability and coding efficiency.

[0014] These or other aspects of this application will become more apparent in the following description of the embodiments. Attached Figure Description

[0015] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0016] Figure 1 A flowchart of a method for determining Lagrange multipliers of encoded frames in a video stream, provided in this application, is shown. Figure 2 A flowchart illustrating an encoding method for encoded frames in a video stream provided in this application is shown; Figure 3 This is a flowchart of the Lambda adjustment method provided in the optional implementation of this application; Figure 4 This is a schematic diagram illustrating the relationship between the IP frame loss ratio and rate-distortion optimization parameters provided in the optional implementation of this application; Figure 5 This is a schematic diagram illustrating the relationship between the IB frame loss ratio and rate-distortion optimization parameters provided in the optional implementation of this application; Figure 6 A schematic diagram of the structure of a device for determining the Lagrange multipliers of encoded frames in a video stream, as provided in this application, is shown. Figure 7 A schematic diagram of the structure of an encoding device for encoding frames in a video stream, as provided in this application, is shown. Detailed Implementation

[0017] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings.

[0018] In some of the processes described in the specification, claims, and accompanying drawings of this application, multiple operations appearing in a specific order are included. However, it should be clearly understood that these operations may not be executed in the order they appear herein, or may be executed in parallel. The operation numbers, such as 101, 102, etc., are merely used to distinguish different operations and do not themselves represent any execution order. Furthermore, these processes may include more or fewer operations, and these operations may be executed sequentially or in parallel. It should be noted that the descriptions such as "first," "second," etc., in this document are used to distinguish different messages, devices, modules, etc., and do not represent a chronological order, nor do they limit "first" and "second" to different types.

[0019] For ease of reference, some terms used in this description are defined as follows. The terms presented and their respective definitions are not strictly limited to these definitions—a term may be further defined by its use in this disclosure. The term “example” as used herein means used as an example, instance, or illustration. Any aspect or design described herein as “exemplary” should not necessarily be construed as superior to other aspects or designs. Rather, the term “exemplary” is used to present the concept in a concrete manner. In this application and the appended claims, the term “or” is intended to mean an inclusive “or” rather than an exclusive “or.” That is, unless otherwise stated or clearly apparent from the context, “X uses A or B” is intended to mean any natural inclusive arrangement. That is, “X uses A or B” is satisfied if X employs A, X employs B, or X employs both A and B. As used herein, at least one of A or B means at least one of A, or at least one of B, or at least one of A and B. In other words, this phrase is disjunctive. The articles “a” and “an” used in this application and in the appended claims should generally be interpreted as “one or more”, unless otherwise stated or clearly apparent from the context that they are in the singular form.

[0020] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties. Furthermore, the collection, use and processing of the relevant data must comply with the relevant laws, regulations and standards of the relevant countries and regions, and corresponding operation entry points are provided for users to choose to authorize or refuse.

[0021] Figure 1A flowchart of a method for determining Lagrange multipliers in coded frames of a video stream, as provided in this application, is shown. Figure 1 As shown, the method may include the following steps: 101. Receive the initial video stream, wherein the initial video stream includes multiple initial frames, and the multiple initial frames include intra-coded frames and other frames; The initial video stream is the raw video data stream to be encoded, which is the input data source for the encoding process and contains a continuous sequence of image frames.

[0022] The initial frame is the raw image frame in the initial video stream that has not been encoded and is the basic unit of the encoding operation.

[0023] Intra-coded frames, also known as I-frames or IDR frames (Instant Decoding Refresh Frames), can be decoded independently without relying on other frames.

[0024] Other frames refer to frames of other types that are different from intra-coded frames, such as P-frames (inter-predictive frames) or B-frames (bidirectional predictive frames).

[0025] By receiving the initial video stream, the core components of the encoding input were identified, providing a complete data source for subsequent pre-analysis and encoding parameter adjustment.

[0026] 102. Determine the target global quantization parameters corresponding to the multiple initial frames respectively; The target global quantization parameter (target global QP) is a global quantization coefficient set separately for each initial frame. It is used to control the bit rate allocation and distortion level during the encoding process. Optionally, the target global quantization parameter can be obtained by adjusting the baseline global quantization parameter determined by the constant quality factor CRF.

[0027] By determining the target global quantization parameters corresponding to multiple initial frames, the method breaks through the one-size-fits-all QP setting mode in the existing technology. The target global QP is allocated according to the frame, which can adapt to the coding characteristics of different frames. For example, IDR frames require a higher quality benchmark, while P / B frames can be flexibly adjusted.

[0028] 103. Control the use of the corresponding target global quantization parameters to pre-encode the corresponding initial frame to obtain encoding statistics, wherein the encoding statistics include the loss values ​​corresponding to multiple reconstructed frames respectively; Among them, the precoding is a lightweight statistically driven coding, the core purpose of which is to collect coding statistics rather than to generate the final coding stream, and it does not execute complete optimization logic.

[0029] Among them, the encoding statistics are the core information collected during the precoding process. They can include various types of data, such as the loss values ​​of the reconstructed frame and the original frame corresponding to the I-frame, the loss value corresponding to the P-frame, the loss value corresponding to the B-frame, the intra-frame percentage, and the skipped kip percentage. There are no restrictions here, and they can be set adaptively according to the actual application and scenario.

[0030] Among them, the reconstructed frame is the image frame obtained after precoding the initial frame, and it is the key comparison object for calculating the degree of distortion.

[0031] The loss value, representing the loss, can be calculated using the mean squared error (MSE).

[0032] By employing lightweight precoding instead of full encoding, the drawback of the increased complexity of dual encoding in existing technologies is addressed, thereby reducing computational overhead and controlling encoding latency. Furthermore, by accurately obtaining the core statistical data of the loss value, a direct basis is provided for the loss ratio calculation in step 104 and the Lambda adjustment in step 105, avoiding the blindness of existing technologies where the Lambda parameter relies on fixed formulas and lacks data support.

[0033] 104. Determine the target loss ratio based on the loss values ​​corresponding to the multiple reconstructed frames respectively, wherein the target loss ratio is the ratio of the first loss value and the second loss value, the first loss value is the loss value corresponding to the intra-coded frame, and the second loss value is the loss value corresponding to the other frames; The target loss ratio is the ratio of the first loss value of the intra-coded frame (I-frame) to the second loss value of other frames (P-frame / B-frame). Different types of frames have corresponding optimal ranges for the ratio. For example, the optimal range for the I / P loss ratio can be set to 0.7-0.8, and the optimal range for the I / B loss ratio is 0.58-0.67.

[0034] Among them, the first loss value corresponding to the intra-coded frame is the MSE value of the original frame and the reconstructed frame after the I-frame is precoded, which is the core metric of the quality benchmark.

[0035] Among them, the second loss value corresponding to other frames, such as the MSE value of P-frames / B-frames after precoding, can optionally be calculated by averaging the previous and subsequent frames and weighting it in combination with the GOP size for P-frames, and by averaging the previous and subsequent frames for B-frames.

[0036] By establishing a quantitative correlation between the loss value and the coding parameters, the problem of the Lambda parameter being disconnected from content characteristics, which is common in existing technologies, is avoided. The loss ratio can accurately reflect the differences in inter-frame coding characteristics, improving the accuracy of subsequent Lambda adjustments.

[0037] 105. Based on the target loss ratio, update the predetermined Lagrange multiplier to obtain the target Lagrange multiplier corresponding to the intra-coded frame. The predetermined Lagrange multiplier includes one of the following: historical Lagrange multiplier or reference Lagrange multiplier.

[0038] The pre-defined Lagrange multipliers are the Lambda parameters before the update, used to control the rate-distortion tradeoff.

[0039] Among them, the historical Lagrange multiplier is the historical Lagrange multiplier. For example, when the video stream is divided into multiple image group structures (GOPs), it can be set to the final Lambda value corresponding to the intra-coded frame of the previous GOP, and used for iterative inheritance of the Lambda parameter corresponding to the current GOP.

[0040] Among them, the benchmark Lagrange multiplier is the initial default Lambda value, which can be set for parameter initialization of the first GOP based on industry experience, CRF benchmark values ​​or video resolution / frame rate derivation.

[0041] The target Lagrange multiplier is the optimal Lambda value adapted to the current intra-coded frame after updating the target loss ratio, and will be used for the recoding of I-frames.

[0042] By using the loss ratio to iteratively update historical or baseline Lambda values, the optimal Lambda parameters for the current IDR frame are obtained. This enables dynamic adaptive adjustment of Lambda parameters, overcoming the shortcomings of existing technologies that rely on fixed mapping or simple hierarchical strategies for Lambda. It allows parameters to be optimized in real-time according to the characteristics of the video content. Furthermore, it supports dual modes of historical parameter inheritance and baseline parameter initialization, ensuring parameter consistency and avoiding quality fluctuations while simultaneously initiating the first encoding process. Ultimately, this improves I-frame encoding efficiency and overall Bjorngard bitrate gain (BD-rate) performance.

[0043] It should be noted that since other frames (P-frames / B-frames) rely on the I-frame for inter-frame prediction, the optimized I-frame possesses high-quality, low-redundancy reference characteristics. When subsequent frames are predicted based on this I-frame, the complexity and volume of the residual data are significantly reduced, thereby lowering the computational overhead and bitrate consumption of encoding other frames. This achieves a cascading gain, improving the prediction accuracy from I-frame optimization to full-frame prediction, and ultimately enhancing the coding efficiency of all frames.

[0044] Through the above steps, an initial video stream is received, comprising multiple initial frames, including intra-coded frames and other frames; target global quantization parameters corresponding to each of the multiple initial frames are determined; the corresponding initial frames are pre-coded using the target global quantization parameters to obtain coding statistics, which include loss values ​​corresponding to each of the multiple reconstructed frames; a target loss ratio is determined based on the loss values ​​corresponding to each of the multiple reconstructed frames, wherein the target loss ratio is the ratio of a first loss value and a second loss value, the first loss value being the loss value corresponding to the intra-coded frame and the second loss value being the loss value corresponding to other frames; and a predetermined Lagrange multiplier is updated based on the target loss ratio to obtain a target Lagrange multiplier corresponding to the intra-coded frame, wherein the predetermined Lagrange multiplier includes one of the following: a historical Lagrange multiplier or a reference Lagrange multiplier. By employing a statistical coding data-driven approach, this method obtains multi-frame loss values ​​through precoding and calculates the target loss ratio between intra-coded frames and other frames. This achieves the goal of dynamically matching the optimal Lagrange multiplier (Lambda) for intra-coded frames, thereby improving the coding efficiency of intra-coded frames. Since the target loss ratio directly reflects the distortion correlation characteristics between intra-coded frames and other frames under different video content, the update of the Lagrange multiplier always fits the actual content characteristics of the current video without relying on fixed empirical parameters. At the same time, lightweight precoding focuses on statistical data collection, avoiding redundant calculations caused by dual coding. It achieves accurate adaptation of parameters and content without increasing complexity, thus solving the technical problem in related technologies where it is difficult to dynamically adjust the Lagrange multiplier, making it difficult to balance video content adaptability and coding efficiency.

[0045] As an optional embodiment, updating the predetermined Lagrange multiplier according to the target loss ratio to obtain the target Lagrange multiplier corresponding to the intra-coded frame includes: if the coding statistics also include the target intra-frame proportion and the target skip proportion, determining the picture complexity based on the target intra-frame proportion and the target skip proportion; determining the Lagrange multiplier limit range based on the picture complexity; and determining the target Lagrange multiplier based on the predetermined Lagrange multiplier and the Lagrange multiplier limit range according to the target loss ratio.

[0046] In this step, the target Lagrange multiplier is the optimal Lambda value that fits the current intra-coded frame after being calculated based on the image complexity limit and loss ratio.

[0047] Among them, the target intra-frame percentage, also known as the target intra-frame percentage, is the percentage of intra-frame prediction mode usage. It is the percentage of intra-frame coding mode application in each frame. A higher percentage usually indicates that the scene changes drastically and is difficult to predict.

[0048] Among them, the target skip ratio can be called the target skip ratio, which is the proportion of skip mode selection. It is the proportion of application of skip encoding mode in each frame. A higher ratio usually indicates that the scene changes smoothly and the temporal correlation is strong.

[0049] Among them, the image complexity is a quantitative indicator calculated based on the target frame ratio and the target skip ratio. It reflects the degree of change and encoding difficulty of the current video image. It can be set into multiple intervals, such as simple, normal and complex intervals.

[0050] The Lagrange multiplier limit is defined by the upper and lower limits of the Lambda ratio based on the complexity of the image. It is used to constrain the updated Lambda parameter and prevent it from deviating from a reasonable range.

[0051] This embodiment illustrates a more refined Lagrange multiplier (Lambda) update scheme. When the coding statistics simultaneously include the target intra-frame proportion, the target skip proportion, and the loss value, the image complexity is first calculated using these two proportions. Then, a reasonable limit range for the Lambda parameter is determined based on the image complexity. Finally, the optimal target Lambda value for the current intra-coded frame is derived by combining the target loss ratio and the predetermined Lambda value.

[0052] This approach improves the accuracy of Lambda parameter adjustment, avoiding the problem of technical parameters being out of sync with content characteristics caused by existing technologies that rely solely on loss values ​​or fixed frame types to adjust Lambda in layers without considering the changing characteristics of the image itself. This solution quantifies image complexity by using intra-frame percentage and skip percentage, ensuring that Lambda adjustment aligns with both inter-frame loss relationships and the actual encoding difficulty of the image. For example, increasing Lambda when the image is complex (high intra-frame percentage) prioritizes bitrate, while decreasing Lambda when the image is smooth (high skip percentage) prioritizes quality, thus resolving the disconnect between existing technical parameters and content characteristics. Furthermore, defining the Lambda limit range by image complexity prevents Lambda parameters from deviating from a reasonable range due to extreme scenarios (such as sudden violent movement or prolonged stillness), ensuring smooth and stable encoding quality and avoiding over- or under-voltage issues caused by parameter mutations in existing technologies.

[0053] As an optional embodiment, obtaining encoding statistics includes: when other frames include the current other frames and historical other frames, determining a second loss value corresponding to other frames based on the loss value corresponding to the current other frames and the loss value corresponding to historical other frames.

[0054] Among them, the other frames currently being encoded are the P-frames or B-frames being processed, and are the immediate objects of loss value statistics.

[0055] Among them, other historical frames are P-frames or B-frames of the same type that have been encoded before the current other frames, and their loss value is historical statistical data.

[0056] Among them, the loss value corresponding to other frames is the MSE value obtained after precoding of the P-frame or B-frame currently being encoded.

[0057] Among them, the loss value corresponding to other historical frames is the MSE value obtained after precoding of the previously encoded P-frames or B-frames of the same type.

[0058] This embodiment illustrates the detailed logic for calculating the target loss ratio. When other frames include the currently encoded P / B frame (current other frames) and previously encoded P / B frames of the same type (historical other frames), the loss value of a single frame is not used directly. Instead, the loss values ​​of the current other frames and historical other frames are fused through a specific calculation method to obtain the comprehensive loss value of the other frames. The target loss ratio is then calculated based on this comprehensive loss value and the loss value of the intra-coded frame.

[0059] This approach improves the accuracy and representativeness of the loss value. Existing technologies often use only the single loss value of the current frame to calculate the ratio, making them susceptible to occasional distortions in a single frame (such as sudden noise or local abnormal textures), causing the ratio to deviate from the true inter-frame characteristics. This embodiment, however, integrates the loss values ​​of the current frame and other historical frames, fully utilizing the temporal correlation between frames. This allows the combined loss value of other frames to better reflect the overall coding distortion characteristics of a class of frames, providing more reliable data support for the target loss ratio. Furthermore, the ratio obtained by combining the current and historical loss values ​​has less fluctuation, avoiding drastic fluctuations in the Lambda parameter caused by sudden changes in loss in a single frame. This eliminates quality fluctuations caused by parameter fluctuations, ensuring smooth and stable coding quality and solving the over- or under-voltage problems caused by parameter mutations in existing technologies.

[0060] As an optional embodiment, determining a second loss value corresponding to other frames based on the loss value corresponding to the current other frames and the loss values ​​corresponding to the historical other frames includes: determining the image group size corresponding to other frames; and determining the second loss value corresponding to other frames based on the loss value corresponding to the current other frames, the loss values ​​corresponding to the historical other frames, and the image group size.

[0061] The group size (GOP) represents the total number of frames contained in a group of images. Optionally, it can be understood as the total number of frames between two adjacent I-frames. For example, a GOP size of 12 means that there are 12 frames between the current I-frame and the next I-frame, including other frames (P / B frames). It is a core parameter in video coding that defines the inter-frame prediction range.

[0062] In this embodiment, the specific calculation logic for the loss values ​​corresponding to other frames is clarified. First, the image group structure to which the current other frames belong is determined. Then, combined with the size of the image group, the instantaneous loss value of the current other frames is fused with the historical loss values ​​of other frames to finally obtain the comprehensive loss value of other frames, providing accurate data support for the subsequent calculation of the target loss ratio.

[0063] It should be noted that this step can be performed adaptively for P-frames. For B-frames, the loss value can be accurately determined without a GOP structure. This is because P-frames are the forward reference core within a GOP, and distortion accumulates and propagates along the reference chain; B-frames are bidirectional auxiliary references and do not assume the primary reference responsibility within a GOP. Therefore, P-frames require weighted statistics based on the GOP structure, while B-frames only require simple averaging.

[0064] Specifically, for P-frames, since they are forward-predicted frames, they only refer to the previous frame (I-frame or P-frame) and are the primary reference node for subsequent P / B frames. Therefore, the GOP size, including information about the interval between two P-frames, directly determines the reference weight of the current P-frame. The larger the GOP, the more times the current P-frame is referenced by subsequent frames, and the more significant its distortion's impact on overall coding quality. The GOP size, as a weighting coefficient, ensures that the distortion of the current P-frame accounts for a higher proportion in the statistics, reflecting both the reference priority of the current frame and smoothing out accumulated errors through historical P-frame distortion, thus preventing sudden distortion changes from affecting subsequent frames.

[0065] For B-frames, since they are bidirectional prediction frames, referencing both preceding and following frames (I / P frames), but not being referenced by any subsequent frames (because B-frames are typically not used as reference frames in coding standards), their distortion only affects themselves and is not propagated to other frames. Therefore, B-frame distortion statistics do not need to consider the GOP structure; simple averaging can smooth out distortion fluctuations between adjacent B-frames, satisfying the statistical requirements of Lambda adjustment without introducing an additional GOP to increase computational complexity.

[0066] This approach allows the loss value calculation to adapt to the coding structure logic. Existing technologies often ignore the impact of image group size on loss statistics, simply summing or averaging current and historical loss values, leading to a disconnect between data and the actual coding dependencies. This embodiment combines the image group structure of other frames; for example, the GOP size is introduced as a weighting coefficient in P-frame loss calculation, making the comprehensive loss value more closely reflect the inter-frame coding dependency characteristics, improving data correlation and rationality. This significantly improves the accuracy and relevance of the loss value. Because the image group size determines the inter-frame reference strength (e.g., the larger the GOP, the stronger the reference correlation between P-frames and historical P-frames), calculating the loss value based on this size avoids statistical bias caused by size differences. This makes the comprehensive loss value more reflective of the true coding distortion characteristics, providing a reliable basis for subsequent Lambda parameter adjustments.

[0067] As an optional embodiment, the control uses the corresponding target global quantization parameters to pre-encode the corresponding initial frame to obtain coding statistics corresponding to multiple reconstructed frames, including: when other frames include multiple types of coded frames, controlling the control uses the corresponding target global quantization parameters to pre-encode the corresponding initial frame to obtain the initial intra-frame percentage and initial skip percentage corresponding to the multiple types of coded frames respectively; determining the adjustment weights corresponding to the multiple types of coded frames respectively; determining the target intra-frame percentage based on the initial intra-frame percentage and the corresponding adjustment weights corresponding to the multiple types of coded frames respectively, and determining the target skip percentage based on the initial skip percentage and the corresponding adjustment weights corresponding to the multiple types of coded frames respectively.

[0068] Among them, multiple types of encoded frames are different types of encoded frames contained in other frames, such as P frames and B frames, which differ in their encoding dependencies and reference logic.

[0069] Among them, the initial intra-frame proportion is the proportion of intra-frame prediction mode used after precoding of multiple types of coded frames (P-frames, B-frames), which is a basic indicator reflecting the coding characteristics of a single type of frame.

[0070] The initial skip ratio, which is the skip mode selection ratio obtained by precoding multiple types of coded frames (P-frames, B-frames), is a basic indicator reflecting the characteristics of image changes in a single type of frame.

[0071] The adjustment weights are weighting coefficients set based on the coding importance and reference value of multiple types of coded frames, used to balance the impact of statistical data from different types of frames on the final result. Optionally, the adjustment weights can be determined based on the video scene type and / or inter-frame dependencies, where inter-frame dependencies can be dependencies between multiple types of coded or intra-coded frames.

[0072] Among them, the target intra-frame proportion is the comprehensive intra-frame prediction mode usage ratio obtained after fusing the initial intra-frame proportions of multiple types of coded frames and their corresponding adjustment weights. It is the core indicator for quantifying the overall image coding characteristics.

[0073] Among them, the target skip ratio is the comprehensive skip mode selection ratio obtained by integrating the initial skip ratios of multiple types of coded frames and their corresponding adjustment weights. It is the core indicator for quantifying the overall image change characteristics.

[0074] This embodiment illustrates the specific logic for obtaining the target intra-frame percentage and target skip percentage in the encoded statistics. When other frames contain multiple types of encoded frames such as P-frames and B-frames, pre-coding is first performed on the corresponding initial frames using the target global quantization parameters to obtain the initial intra-frame percentage and initial skip percentage for each type of encoded frame. Then, adjustment weights are set according to the encoding importance of different types of encoded frames. Finally, through weighted calculation, the initial intra-frame percentage and initial skip percentage of each type of encoded frame are fused to obtain the target intra-frame percentage and target skip percentage that reflect the overall picture characteristics.

[0075] This approach enhances the comprehensiveness and representativeness of the proportion statistics, avoiding the pitfalls of existing technologies that rely solely on proportion data from a single frame type or simple averaging, neglecting the unique characteristics of different types of coded frames (e.g., P-frames have higher reference value than B-frames). This embodiment differentiates between multiple types of coded frames and sets adjustment weights, fully integrating statistical information from various frame types. This ensures that the proportion within the target frame and the target skip proportion comprehensively reflect the overall image's coding characteristics, avoiding statistical biases caused by data from a single frame type. Furthermore, this weight allocation is deeply matched to the inter-frame coding dependencies, making the weighted target proportion more closely aligned with actual coding logic. This provides more accurate input data for subsequent image complexity calculations, resolving the disconnect between proportion statistics and coding dependencies in existing technologies.

[0076] It should be noted that the weight settings can be adapted to different scenarios. For example, the encoding characteristics of P-frames and B-frames vary significantly in different video scenarios (e.g., in motion scenarios, P-frames have a high intra-frame ratio, while in static scenarios, B-frames have a high skipped ratio). This embodiment balances the influence of various frames by adjusting the weights, enabling the target ratio calculation logic to adapt to different scenarios without manual parameter adjustment. This overcomes the limitations of existing technologies where the ratio statistics method is fixed and difficult to adapt to diverse scenarios, thus improving the versatility and robustness of the solution.

[0077] As an optional embodiment, determining the target global quantization parameters corresponding to the multiple initial frames includes: determining the region of interest mapping map corresponding to the multiple initial frames; and adjusting the local quantization parameters included in the corresponding reference global quantization parameters according to the region of interest mapping map corresponding to the multiple initial frames to obtain the target global quantization parameters corresponding to the multiple initial frames.

[0078] Among them, the region of interest map, or ROImap, is a map that marks the importance distribution of different regions in an image (such as the face of a person and the core subject as high importance regions, and the background as low importance regions). Optionally, it can be generated by Lookahead and CTU tree analysis.

[0079] Among them, the benchmark global quantization parameters are basic quantization parameters established based on the preset CRF (constant quality factor), which include a globally unified benchmark value and subdivided local quantization parameters, and serve as the initial basis for subsequent adjustments.

[0080] Among them, the local quantization parameter is the quantization coefficient for a local region of the image in the baseline global quantization parameter. It can be adjusted individually according to the importance of the region and is a key sub-parameter for achieving fine quality control.

[0081] This embodiment illustrates the specific logic for determining the target global quantization parameters. First, a corresponding region of interest mapping map is generated for each initial frame. Then, based on the importance of the regions marked in the mapping map, the local quantization parameters in the baseline global quantization parameters are differentially adjusted to obtain the target global quantization parameters adapted to the regional characteristics of each initial frame.

[0082] This approach enables fine-grained bitrate allocation, improving coding efficiency. Existing technologies often use globally uniform quantization parameters, which can lead to distortion in important areas due to insufficient bitrate, and wasted resources in less important areas due to redundant bitrate. This embodiment combines the region of interest (ROI) mapping with local quantization parameters. Local QP is lowered for highly important areas to ensure quality, while local QP is increased for less important areas to save bitrate, achieving higher overall coding efficiency under the same bitrate constraint. Furthermore, by explicitly marking the core content of the image in the ROI mapping, such as the subject and key details, targeted adjustment of local quantization parameters prioritizes the reconstruction quality of these visually sensitive areas, avoiding the core area distortion problem caused by the traditional one-size-fits-all parameter approach. This makes the encoded output more consistent with human visual perception characteristics. Moreover, the scene content and regional importance distribution differ for each initial frame; for example, some frames focus on dynamic subjects while others focus on static backgrounds. A separate ROI mapping is generated for each initial frame and parameters are adapted accordingly. This overcomes the limitation of existing technologies where uniform parameters cannot adapt to the differences in regional characteristics between frames, improving the targeting of parameter settings.

[0083] As an optional embodiment, updating the predetermined Lagrange multiplier according to the target loss ratio to obtain the target Lagrange multiplier corresponding to the intra-coded frame includes: determining the spatiotemporal coding features corresponding to multiple initial frames respectively; updating the predetermined Lagrange multiplier according to the target loss ratio and the spatiotemporal coding features corresponding to multiple initial frames respectively, and determining the target Lagrange multiplier corresponding to the intra-coded frame.

[0084] Among them, the spatiotemporal domain coding features are a set of coding features that integrate the temporal and spatial characteristics of video frames. Temporal features include inter-frame differences, motion intensity, scene switching trends, etc., while spatial features include texture complexity, edge density, brightness and chromaticity distribution, CTU texture mode, etc., which are the core dimensions reflecting the coding characteristics of video content.

[0085] This embodiment illustrates a Lagrange multiplier (Lambda) update scheme that integrates multi-dimensional features. First, the spatiotemporal coding features of each initial frame are extracted to comprehensively capture the temporal motion characteristics and spatial texture characteristics of the video content. Then, these features are combined with the target loss ratio to update and adjust the predetermined Lagrange multiplier, ultimately determining the target Lagrange multiplier that is suitable for the current intra-coded frame.

[0086] This approach improves the accuracy and comprehensiveness of Lambda parameter adjustment. It avoids the problem that existing technologies often rely solely on loss ratios or single-type features to adjust Lambda, failing to fully reflect the coding characteristics of video content. This embodiment introduces spatiotemporal coding features. Temporal features capture inter-frame motion relationships, while spatial features characterize the complexity of single-frame content, complementing the target loss ratio. This ensures that parameter adjustment not only aligns with inter-frame distortion relationships but also adapts to the spatiotemporal characteristics of the content itself, solving the problem of existing technologies having a single dimension of parameter adjustment and being disconnected from content characteristics.

[0087] Specifically, the linkage between pre-analysis and encoding decisions can be strengthened. The encoding process can be clearly divided into two stages: pre-analysis and actual encoding. In the pre-analysis stage, spatiotemporal encoding features (such as temporal inter-frame differences and spatial texture complexity) can be obtained through Lookahead and CTU tree analysis. This embodiment reuses pre-analysis results to guide Lambda updates, deeply binding encoding decisions with pre-analysis features. This avoids a disconnect between parameter settings and content analysis, leveraging the advantages of the pre-analysis and encoding linkage architecture and improving the overall coherence of the encoding logic.

[0088] Furthermore, the spatiotemporal characteristics of different video scenes differ significantly (e.g., sports events have intense temporal motion, while static documentaries have complex spatial textures). By perceiving scene characteristics through spatiotemporal coding features, the Lambda parameter adjustment can be more targeted. For example, in scenes with intense motion (significant temporal features), the Lambda can be appropriately increased to emphasize bitrate, while in high-texture scenes (significant spatial features), the Lambda can be appropriately decreased to emphasize quality. This allows for adaptation to different scenes without manual intervention, overcoming the limitation of existing technologies where fixed adjustment logic is difficult to adapt to diverse scenarios.

[0089] As an optional embodiment, determining the spatiotemporal coding features corresponding to multiple initial frames includes: splitting the target frame to obtain multiple coding tree units, wherein the target frame is any one of the multiple initial frames; and performing temporal and spatial dual-layer analysis processing on the multiple coding tree units based on a buffer queue and sliding window mechanism to obtain the spatiotemporal coding features corresponding to the target frame.

[0090] The target frame is any one of the multiple initial frames. It is the specific object of the current spatiotemporal coding feature extraction. It can be used as the target frame for feature extraction in turn for each initial frame. This also means that any initial frame can be used to perform this step.

[0091] Among them, the splitting process, which is the operation of dividing the target frame into multiple independent coding units of a fixed size, is a fundamental step in realizing refined feature analysis.

[0092] The coding tree unit (CTU) is the basic coding unit obtained after splitting the target frame. It is the smallest unit for temporal and spatial analysis, and its characteristics directly reflect the coding characteristics of local regions within the frame.

[0093] Among them, the cache queue is a fixed-length data queue, such as a fixed-length data queue of 40-60 frames built in the pre-analysis stage, used to store the frame data to be analyzed and the analyzed frames, providing inter-frame correlation information support for time domain analysis.

[0094] Among them, the sliding window mechanism is a dynamic management mechanism for maintaining the frame order relationship in the buffer queue. It realizes the updating and iteration of frame data by sliding the window, ensuring that the temporal domain analysis can capture the latest inter-frame temporal correlation.

[0095] Among them, the temporal and spatial dual-layer analysis processing is a two-layer feature analysis performed simultaneously on the coding tree unit. The temporal analysis focuses on temporal characteristics such as inter-frame motion correlation and difference, while the spatial analysis focuses on single-frame regional characteristics such as texture complexity and edge density.

[0096] This embodiment illustrates the specific extraction logic of spatiotemporal coding features. First, any one of the multiple initial frames is taken as the target frame and split into multiple coding tree units. Then, relying on the buffer queue and sliding window mechanism built in the pre-analysis stage, temporal analysis (capturing inter-frame temporal correlation) and spatial analysis (characterizing single-frame regional characteristics) are simultaneously performed on these coding tree units, ultimately obtaining spatiotemporal coding features that can comprehensively reflect the coding characteristics of the target frame.

[0097] This approach enables refined and comprehensive feature extraction. Existing technologies often perform single-dimensional feature analysis based on the entire frame, making it difficult to capture local regional characteristics and inter-frame correlations. This embodiment splits the target frame into coding tree units and simultaneously acquires temporal and spatial features through dual-layer analysis. It covers both local texture and edge details within a single frame and temporal characteristics such as inter-frame motion and differences, solving the problems of single-dimensional and coarse-grained traditional feature extraction. Furthermore, it ensures the timeliness and relevance of features. Because the cache queue stores a sufficient amount of frame data, the sliding window mechanism ensures real-time updates to frame order relationships, enabling temporal analysis to accurately capture the correlation characteristics between adjacent frames, historical frames, and the current target frame, avoiding feature distortion caused by a lack of temporal data.

[0098] As an optional embodiment, updating the predetermined Lagrange multiplier according to the target loss ratio to obtain the target Lagrange multiplier corresponding to the intra-coded frame includes: determining cost information corresponding to multiple initial frames respectively; updating the predetermined Lagrange multiplier according to the cost information and the target loss ratio to determine the target Lagrange multiplier corresponding to the intra-coded frame.

[0099] Among them, cost information, which can be called cost information, reflects the resource consumption and difficulty characteristics in the coding process. It can be the information included in the core coding data obtained through Look ahead and CTU tree analysis in the pre-analysis stage. Optionally, it can include CTU prediction cost, residual energy distribution and coding complexity score.

[0100] This embodiment illustrates a Lagrange multiplier (Lambda) update scheme that incorporates cost information. First, cost information for each initial frame is extracted to capture the prediction difficulty, residual characteristics, and complexity level during the encoding process. Then, this cost information is combined with the target loss ratio to update and adjust the predetermined Lagrange multiplier, ultimately determining the target Lagrange multiplier that is suitable for the current intra-frame coded frame.

[0101] This approach improves the accuracy and relevance of Lambda parameter adjustments. It avoids the problem of existing technologies relying solely on loss ratios to adjust Lambda, neglecting the actual cost characteristics during the encoding process (such as prediction difficulty and residual energy). The cost information introduced in this embodiment directly reflects the encoding difficulty; high CTU prediction cost indicates high regional encoding difficulty, residual energy distribution reflects distortion distribution, and encoding complexity score quantifies the overall encoding overhead, complementing the target loss ratio. This ensures that parameter adjustments align with inter-frame distortion relationships and adapt to the actual difficulty of the encoding process, solving the problem of disconnect between parameter adjustments and encoding execution in existing technologies. Furthermore, it optimizes the practical implementation of rate-distortion balancing. The core of the rate-distortion cost function is balancing encoding cost (R) and distortion (D), and cost information is directly related to the actual distribution of encoding cost. By adjusting Lambda in conjunction with cost information, the rate-distortion tradeoff can be made more closely aligned with actual coding performance. For example, in scenarios with high coding complexity (large cost information values), Lambda can be appropriately increased to control bitrate overhead, while in scenarios with low coding complexity, Lambda can be appropriately decreased to improve quality. This ensures that the rate-distortion balance not only remains at the theoretical level but also adapts to the resource consumption characteristics of actual coding, thereby improving overall BD-rate performance.

[0102] As an optional embodiment, after updating the predetermined Lagrange multiplier according to the target loss ratio to obtain the target Lagrange multiplier corresponding to the intra-coded frame, the method further includes: determining the target coding mode corresponding to the intra-coded frame based on the target Lagrange multiplier corresponding to the intra-coded frame.

[0103] Among them, the target coding mode is the optimal intra-frame coding configuration selected in the rate-distortion optimization (RDO) calculation based on the target Lagrange multiplier, and is the specific execution scheme to achieve efficient intra-frame coding.

[0104] This embodiment illustrates the subsequent execution steps after dynamic Lambda adjustment. After obtaining the target Lagrange multiplier corresponding to the intra-coded frame through target loss ratio update, the optimized Lambda parameter is used as the core input and substituted into the rate-distortion optimization (RDO) cost function (J=D+λ×R). By calculating the total cost under different coding modes, the optimal combination of coding modes is selected, and finally the target coding mode that is suitable for the current intra-coded frame is determined.

[0105] This approach improves the accuracy of encoding mode selection, avoiding the problem in existing technologies where encoding mode selection relies on fixed Lambda parameters, leading to a disconnect between mode selection and video content characteristics. In this embodiment, the target Lagrange multiplier is dynamically optimized based on content statistics. Selecting the target encoding mode based on this ensures a deep match between the encoding mode and the content characteristics of the intra-coded frames. For example, complex scenes correspond to more refined CTU partitioning modes, while simple scenes correspond to simpler encoding strategies, solving the problem of blind mode selection in traditional methods. This ultimately forms a complete closed loop from parameter optimization to mode adaptation.

[0106] As an optional embodiment, determining the target coding mode corresponding to the intra-coded frame based on the target Lagrange multiplier corresponding to the intra-coded frame includes: determining coding reference information corresponding to multiple initial frames respectively; excluding invalid coding modes based on the corresponding coding reference information to obtain candidate coding modes corresponding to multiple initial frames respectively; and determining the target coding mode corresponding to the intra-coded frame from the corresponding candidate coding modes based on the target Lagrange multiplier corresponding to the intra-coded frame respectively.

[0107] The encoding reference information is used to determine the effectiveness of the encoding pattern. It can be the core guiding data obtained through Lookahead and CTU tree analysis in the pre-analysis stage, and optionally may include the predicted optimal segmentation pattern, motion trends, reference frame selection suggestions, CTU texture patterns, and encoding difficulty assessment.

[0108] Invalid coding modes are those that do not match the coding characteristics of the current frame (such as texture complexity and motion trends), or have extremely low coding efficiency and excessive risk of distortion, and therefore do not have practical coding application value.

[0109] Among them, candidate coding modes are the set of remaining coding modes that are compatible with the coding characteristics of the current frame and are feasible after excluding invalid coding modes. They form the basis for subsequent selection of target coding modes. Optionally, invalid coding modes can be coding modes that do not match the texture complexity of the current frame (e.g., are below a predetermined complexity range) and / or whose coding efficiency is below a threshold.

[0110] The target coding mode is the optimal coding configuration selected from the candidate coding modes. Optionally, it can be the optimal coding mode verified by RDO calculation, and can include CTU partitioning mode, intra-frame prediction mode, residual coding strategy, etc. It is the execution scheme for the actual coding of the intra-frame coded frame.

[0111] This embodiment illustrates the refined selection logic for determining the target coding mode of an intra-coded frame. First, corresponding coding reference information is determined for each initial frame to clarify its coding characteristics and adaptation direction. Then, based on this coding reference information, invalid coding modes that do not match the frame characteristics are eliminated, resulting in feasible candidate coding modes. Finally, using the target Lagrange multiplier of the intra-coded frame as the core criterion, the optimal solution selected is the target coding mode that adapts to the current intra-coded frame.

[0112] This approach reduces the computational complexity of encoding. Existing technologies often involve traversing all encoding modes for RDO calculations, resulting in massive computational loads and increased encoding latency. This embodiment eliminates invalid encoding modes in advance using encoding reference information, significantly reducing the number of candidate encoding modes and the number of subsequent calculation traversals, thus solving the efficiency problem caused by redundant encoding mode selection in traditional schemes. Furthermore, it improves the accuracy of encoding mode selection. The encoding reference information directly reflects the encoding characteristics of the frame, and candidate encoding modes selected based on this information naturally fit the actual encoding requirements of the frame. Combined with dynamically optimized target Lagrange multipliers, the target encoding mode is both adapted to frame characteristics and achieves optimal rate-distortion balance, avoiding the blind reliance on fixed experience and disconnect between encoding mode selection and frame characteristics in existing technologies.

[0113] As an optional embodiment, after determining the target coding mode corresponding to the intra-coded frame based on the target Lagrange multiplier corresponding to the intra-coded frame, the method further includes: when there are multiple intra-coded frames, determining the target intra-coded frame corresponding to other frames from the multiple intra-coded frames based on the image group size corresponding to other frames; and obtaining the coding mode corresponding to other frames based on the target coding mode corresponding to the target intra-coded frame and the coding modes corresponding to other frames in pre-coding.

[0114] Among them, multiple intra-coded frames are a set of multiple I-frames contained in the initial video stream. Each intra-coded frame has independent decoding and random access functions and can be used as the starting frame and quality benchmark for different GOPs.

[0115] Among them, the target intra-coded frame is the IDR frame selected from multiple intra-coded frames and directly associated with the GOP structure to which other frames belong. It is the core reference benchmark when encoding other frames.

[0116] Among them, the encoding modes corresponding to other frames in the precoding are the encoding modes that were initially determined in the precoding stage of other frames. They are feasible but have not been adapted and optimized with the encoded frames in the target frame, and are the basis for subsequent adjustments.

[0117] Among them, the encoding mode corresponding to other frames is the optimal encoding mode that is finally determined by combining the target encoding mode of the target intra-coded frame and the encoding mode of other pre-coded frames, and is adapted to the current GOP reference logic. It is the execution scheme for the actual encoding of other frames.

[0118] This embodiment illustrates the logic for determining the encoding modes of other frames in a multi-frame intra-coded scenario. When the video stream contains multiple intra-coded frames, firstly, based on the Group of Pictures (GOP) structure to which the other frames belong, a target intra-coded frame directly associated with that GOP is found among the multiple intra-coded frames; then, using the target encoding mode of the target intra-coded frame as a reference, combined with the encoding modes determined by the other frames during the pre-coding stage, the final encoding modes of the other frames are obtained after comprehensive adaptation.

[0119] This approach not only precisely optimizes the reference characteristics of the I-frame itself but also drives a global improvement in the coding performance of other frames (P-frames / B-frames), enhancing the prediction accuracy of their coding. Existing technologies often use randomly matched intra-frame coded frames as references or directly employ fixed coding patterns, leading to a mismatch in coding logic between other frames and the reference frame, resulting in significant prediction distortion. This embodiment precisely matches the target intra-frame coded frame based on the image group structure. Furthermore, this target I-frame is a product optimized through dynamic Lambda adjustment, possessing high-quality, low-redundancy reference characteristics. This ensures a deep fit between the coding patterns of other frames and the coding configuration of the reference frame, strengthening the correlation between inter-frame predictions and reducing prediction distortion. Since other frames rely on the I-frame for inter-frame prediction, the complexity and volume of residual data are significantly reduced when subsequent frames make predictions based on this optimized I-frame, simultaneously reducing the computational overhead and bitrate consumption of other frame coding. Simultaneously, this scheme considers both the feasibility and optimizability of the coding pattern. Since the coding patterns corresponding to the pre-coding of other frames already possess basic feasibility, the target coding pattern is the optimal optimized solution. The combination of these two approaches avoids redundant calculations in designing coding schemes from scratch and achieves optimization and upgrades through the reference of the target coding scheme. This resolves the contradiction in existing technologies where coding schemes either lack optimization or are impractical, thus balancing coding efficiency and optimization effects. Ultimately, by leveraging the characteristic of the I-frame as a reference base for other frames, the optimization benefits of the I-frame are passed on to all subsequent frames, achieving a cascading gain from I-frame optimization to improved prediction accuracy across the entire frame. This leads to enhanced coding efficiency across all frames, comprehensively improving overall coding performance from the reference base to the entire frame stream.

[0120] As an optional embodiment, after obtaining the encoding modes corresponding to other frames based on the target encoding mode corresponding to the target intra-coded frame and the encoding modes corresponding to other frames in pre-coding, the method further includes: encoding the intra-coded frame according to the target encoding mode corresponding to the intra-coded frame to output a first single-frame encoded stream, and encoding other frames according to the encoding modes corresponding to other frames to obtain a second single-frame encoded stream; and integrating the first single-frame encoded stream and the second single-frame encoded stream to obtain the target video frame.

[0121] The first single-frame encoded stream is the encoded data stream corresponding to a single frame after the intra-frame encoded frame is encoded using the target encoding mode.

[0122] The second single-frame encoded stream is the encoded data stream corresponding to a single frame after other frames are encoded using the corresponding encoding mode.

[0123] The target video frame is a complete encoded output unit obtained by integrating the first single-frame encoded stream and the second single-frame encoded stream according to the GOP frame order relationship, and it is the basic component of the final video stream.

[0124] This embodiment illustrates the execution logic of the final encoded output. First, the intra-coded frame is encoded according to the target coding mode optimized for intra-coded frames, generating a first single-frame encoded stream. Simultaneously, other frames are encoded according to the coding mode adapted for other frames, generating a second single-frame encoded stream. Finally, the two types of single-frame encoded streams are integrated according to the frame order association relationship in the GOP structure to obtain the target video frame that conforms to the coding standard.

[0125] This approach ensures a precise balance between coding efficiency and quality. The target coding mode of the intra-coded frame is the optimal scheme after dynamic Lambda adjustment and RDO calculation optimization. The coding modes of other frames are adapted to the GOP reference logic. The generation of both types of coded streams is based on refined decision-making, avoiding the efficiency waste or quality distortion caused by fixed coding modes in existing technologies. It can improve reconstruction quality at the same bitrate and reduce the bitrate at the same quality. The integration process can be based on the frame order relationship of the GOP structure, ensuring that the reference role of the intra-coded frame is effectively connected with the prediction dependency logic of other frames. This avoids subsequent decoding anomalies or prediction distortion caused by coded stream fragmentation, ensures the consistency of coding logic within the GOP, and solves the problem of inter-frame coding disconnect in existing technologies. Both types of encoded streams are generated based on encoding modes that adapt to frame characteristics. The integrated target video frame not only ensures the quality benchmark of the intra-frame encoded frame, but also takes into account the encoding efficiency of other frames, avoiding quality fluctuations. At the same time, it is integrated according to the standard frame order, which conforms to the video encoding protocol specifications and ensures the decoding compatibility of the output video stream, thus solving the limitations of insufficient encoding output stability and poor compatibility in the existing technology.

[0126] As an optional embodiment, after receiving the initial video stream, the method further includes: dividing the initial video stream to obtain multiple image group structures; performing precoding processing on the multiple image group structures sequentially according to the frame order to obtain corresponding coding statistics; updating the target Lagrange multiplier corresponding to the current image group structure based on the corresponding coding statistics; and using the target Lagrange multiplier corresponding to the current image group structure as the benchmark distortion trade-off index for the next set of image group structures.

[0127] The multiple picture group structure consists of multiple GOP (picture group) sets obtained after the initial video stream is divided. Each picture group starts with an intra-coded frame (IDR frame) and contains P frames and B frames, and has independent inter-frame reference relationships and coding logic.

[0128] The frame order is a rule that arranges image groups and frames within groups according to the chronological order of the original video stream, ensuring that the encoding and playback timing are consistent.

[0129] The encoding process is a complete encoding flow performed on each image group structure, including the precoding, statistical information collection, parameter adjustment, and formal encoding steps mentioned above. The core is to generate the encoded stream and collect encoded statistical data.

[0130] The current image group structure, which is the image group undergoing encoding processing and parameter updates, is the computational object of the target Lagrange multiplier.

[0131] The next image group structure is the next image group to be encoded arranged in frame order after the current image group, and the target Lagrange multipliers of the previous group should be used as initial parameters.

[0132] Among them, the baseline distortion trade-off index is the Lambda value used to initialize the coding parameters of the next set of images. It is directly inherited from the target Lagrange multiplier of the current image set and is the basis for the dynamic update of the next set of parameters.

[0133] This embodiment illustrates the iterative update logic of the Lagrange multiplier (Lambda) in a multi-image group scenario. After receiving the initial video stream, it is first divided into multiple independent image group structures; then, according to the original frame order, encoding processing is performed on each image group sequentially, and encoding statistics for that group are collected during the process; using these statistics, the target Lagrange multiplier adapted to the current image group is dynamically updated; finally, this target index is directly used as the benchmark distortion trade-off index for the next image group, providing an initial basis for parameter optimization of the next group, forming a cyclic iterative parameter update mechanism.

[0134] This method enables cross-group iterative optimization of Lambda parameters, improving overall coding coherence. In existing technologies, Lambda parameters for different image groups are often initialized independently, lacking correlation and leading to quality fluctuations between groups. This embodiment utilizes an inheritance mechanism from the current target index to the next set of benchmark indices, ensuring temporal correlation in parameter updates and a smooth transition in coding quality between adjacent image groups, thus resolving the problem of parameter disconnect between groups. Since coding statistics reflect the content characteristics of the current image group (such as image complexity and motion trends), the updated target Lagrange multipliers are already adapted to the characteristics of that group. Using this as the benchmark index for the next group ensures that the initialization of the next set of parameters aligns with the coding characteristics of adjacent scenes, avoiding the blind calculations and redundant computations caused by initialization from zero, and improving the accuracy and efficiency of overall parameter adjustment.

[0135] Furthermore, it can adapt to long video scenarios with multiple image groups, enhancing the versatility of the solution. Long videos typically contain a large number of image groups, and existing technologies struggle to achieve adaptive parameter transitions between groups. The iterative update mechanism in this embodiment requires no manual intervention and can automatically adapt to the encoding requirements of different image groups in long videos. Whether it's the alternation between moving and static scenes or the gradual change in texture complexity, smooth adaptation can be achieved through parameter inheritance and dynamic updates, overcoming the limitation of existing technologies where fixed initialization logic is difficult to adapt to long videos. Moreover, the I-frame quality of each image group is determined by the target Lagrange multiplier, and cross-group parameter inheritance ensures the consistency and coherence of I-frame quality across image groups, avoiding the impact of inter-group I-frame quality fluctuations on the prediction accuracy of subsequent P / B frames. This consistency makes the rate-distortion balance of the entire video stream more stable, ultimately improving overall BD-rate performance.

[0136] Figure 2 A flowchart illustrating an encoding method for encoded frames in a video stream provided in this application is shown, as follows: Figure 2 As shown, the method may include the following steps: 201, Receive initial video stream, wherein the initial video stream includes multiple initial frames, including intra-coded frames and other frames; 202, Determine the target global quantization parameters corresponding to the multiple initial frames respectively; 203. Control the use of the corresponding target global quantization parameters to pre-encode the corresponding initial frame to obtain encoding statistics, wherein the encoding statistics include the loss values ​​corresponding to multiple reconstructed frames respectively; 204. Determine the target loss ratio based on the loss values ​​corresponding to the multiple reconstructed frames respectively, wherein the target loss ratio is the ratio of the first loss value and the second loss value, the first loss value is the loss value corresponding to the intra-coded frame, and the second loss value is the loss value corresponding to the other frames; 205. Based on the target loss ratio, update the predetermined Lagrange multiplier to obtain the target Lagrange multiplier corresponding to the intra-coded frame. The predetermined Lagrange multiplier includes one of the following: historical Lagrange multiplier or reference Lagrange multiplier. 206. Determine the target coding mode corresponding to the intra-coded frame based on the target Lagrange multiplier corresponding to the intra-coded frame; 207. Encode the intra-coded frame according to the target coding mode corresponding to the intra-coded frame, and output the single-frame coded stream corresponding to the intra-coded frame.

[0137] Through the above steps, an initial video stream is received, comprising multiple initial frames, including intra-coded frames and other frames; target global quantization parameters corresponding to each of the initial frames are determined; the corresponding initial frames are pre-coded using the target global quantization parameters to obtain coding statistics, which include loss values ​​corresponding to each of the reconstructed frames; a target loss ratio is determined based on the loss values ​​corresponding to each of the reconstructed frames, wherein the target loss ratio is the ratio of a first loss value and a second loss value, where the first loss value is the loss value corresponding to the intra-coded frame and the second loss value is the loss value corresponding to other frames; a predetermined Lagrange multiplier is updated based on the target loss ratio to obtain a target Lagrange multiplier corresponding to the intra-coded frame, wherein the predetermined Lagrange multiplier includes one of the following: a historical Lagrange multiplier and a reference Lagrange multiplier; a target coding mode corresponding to the intra-coded frame is determined based on the target Lagrange multiplier corresponding to the intra-coded frame; the intra-coded frame is encoded according to the target coding mode corresponding to the intra-coded frame, and a single-frame coded stream corresponding to the intra-coded frame is output. By employing a statistical coding data-driven approach, this method acquires multi-frame loss values ​​through precoding and calculates the target loss ratio between intra-coded frames and other frames. This achieves the goal of dynamically matching the optimal Lagrange multiplier (Lambda) for intra-coded frames, thereby improving the coding efficiency of intra-coded frames. Since the target loss ratio directly reflects the distortion correlation characteristics between intra-coded frames and other frames under different video content, the update of the Lagrange multiplier always closely matches the actual content characteristics of the current video, without relying on fixed empirical parameters. At the same time, lightweight precoding focuses on statistical data collection, avoiding redundant calculations caused by dual coding. It achieves accurate adaptation of parameters and content without increasing complexity, ultimately enabling the encoding of intra-coded frames and balancing video content adaptability and coding efficiency. This solves the technical problem in related technologies where it is difficult to dynamically adjust the Lagrange multiplier, thus making it difficult to balance video content adaptability and coding efficiency.

[0138] Based on the above embodiments and optional embodiments, an optional implementation method is provided, which is described in detail below.

[0139] Among related technologies, with the explosive growth of digital video applications, video coding technology has become a core technology of modern multimedia communication. From the early H.261 and MPEG-1 standards to the current H.266 / VVC (Versatile Video Coding), video coding technology has been continuously optimized in terms of compression efficiency, coding quality, and computational complexity. Among these, Rate-Distortion Optimization (RDO), as a core technology of modern video encoders, maximizes coding efficiency by seeking the optimal balance between coding rate and distortion.

[0140] Rate-distortion optimization (RDO) is based on rate-distortion theory in information theory. Its core idea is to select the optimal coding scheme given a rate-distortion cost function. A typical RDO cost function is expressed as: J = D + λ × R; Where J is the total cost, D is the distortion between the reconstructed signal and the original signal (usually the mean square error MSE), R is the number of bits required for encoding, and λ is the Lagrange multiplier (Lambda parameter), used to control the rate-distortion tradeoff.

[0141] IDR frames (Instantaneous Decoder Refresh Frames) are a key frame type in video coding, with the following characteristics: 1) Independent decoding: IDR frames do not depend on any reference frames and can be decoded independently; 2) Random access point: They provide random access functionality to the video stream, supporting fast location and error recovery; 3) High bitrate overhead: Due to the inability to utilize temporal prediction, IDR frames typically occupy a high bitrate; 4) Quality baseline: The quality of IDR frames directly affects the coding efficiency of subsequent frames and the overall visual quality.

[0142] The choice of the Lambda parameter directly affects the rate-distortion performance of the encoder.

[0143] Traditional lambda settings are typically based on the empirical formula: λ = α × 2^((QP - 12) / 3), where α is an empirical coefficient selected based on different temporalIds, and QP is the quantization parameter obtained through rate control. The lambda parameter settings are mainly based on a fixed mapping relationship of the quantization parameter (QP). While this method is simple and effective, it does not consider the encoding characteristics of different video scenarios.

[0144] The traditional one-size-fits-all parameter setting method of Lambda settings leads to a number of problems, such as: 1) the rate-distortion performance of IDR frames is not optimal; 2) there is room for improvement in overall coding efficiency; 3) under the same bitrate constraint, the video quality is not ideal.

[0145] Based on this, existing solutions have developed hierarchical lambda setting strategies, which set different lambda parameters for different types of frames according to their importance level (e.g., I-frames, P-frames, B-frames), and adjust the base lambda value according to the frame type. However, the adjustment strategy is still relatively fixed, lacks adaptability, and is poorly adapted to changes in video content. Alternatively, existing solutions require the following steps: video input, I / P / B frame encoding, loss statistics, ratio calculation, lambda adjustment, I-frame re-encoding, and final output. It is evident that this involves double encoding, increasing computational complexity. Not only does it require pre-encoding the current I-frame to obtain loss information, but it also requires performing two complete encoding processes for each I-frame, increasing encoding time and introducing additional encoding latency.

[0146] In view of this, the optional implementation of this application proposes a method for determining the Lagrange multipliers of encoded frames in a video stream, which can also be called a technical solution that can dynamically adjust the Lambda parameter in the RDO process of IDR frames. It can improve the coding efficiency of IDR frames, thereby optimizing the overall BD-rate performance and solving the technical problems of poor coding efficiency of IDR frames and lack of adaptability of parameter settings in the prior art.

[0147] Figure 3 This is a flowchart of the Lambda adjustment method provided in the optional implementation of this application, such as... Figure 3 As shown, it will be introduced below: S1, divide the initial video stream to obtain multiple image group structures; S2, according to the frame order, pre-encode multiple image group structures sequentially to obtain the corresponding encoding statistics; S3, based on the corresponding encoded statistical data, update the target Lagrange multiplier corresponding to the current image group structure; S4, use the target Lagrange multiplier corresponding to the current image group structure as the benchmark distortion tradeoff index for the next image group structure.

[0148] For steps S2-S3, the following is also included: (a) Preprocessing stage The current frame is split to obtain multiple coding tree units; Based on the buffer queue and sliding window mechanism, multiple coding tree units are subjected to temporal and spatial dual-layer analysis to obtain the spatiotemporal coding features corresponding to the current frame. Determine the cost information corresponding to the current frame; Determine the encoding reference information corresponding to the current frame; Determine the region of interest mapping map corresponding to the current frame.

[0149] Specifically, let's take an example: For the preprocessing stage of dynamic Lambda adjustment: System inputs include: Receive the input video stream.

[0150] In this phase, we build the Lookahead and CTU tree analysis system. First, we establish a fixed-length buffer queue of 40-60 frames and use a sliding window mechanism to maintain the frame order relationship and GOP structure.

[0151] A two-layer analysis is performed on the cached content to obtain the analysis results: during the analysis, the inter-frame difference is calculated in the temporal domain, the intensity of motion and scene switching points are identified, and the texture complexity, edge density and brightness and chromaticity distribution are evaluated in the spatial domain.

[0152] Based on the analysis results, GOP optimization and IDR frame location determination are performed: each frame is divided into fixed-size CTU units and a CTU tree is constructed. The texture complexity and motion features of each CTU are calculated, the encoding difficulty is evaluated, the partitioning pattern is determined, and the hierarchical structure and CTU relationships are established. CTU-level features, including motion vectors, texture patterns, variance gradients, and other information, are extracted. CTU-level parallel processing and feature caching mechanisms are employed to improve efficiency and ensure the accuracy and real-time performance of the pre-analysis results.

[0153] System output includes: (1) Cost information: CTU prediction cost, residual energy distribution and coding complexity score; (2) ROImap: Marks the distribution of important regions in an image; (3) Encoding reference information: the predicted optimal segmentation pattern, motion trend and reference frame selection suggestions.

[0154] (II) Pre-analysis stage Based on the region of interest mapping maps corresponding to multiple initial frames, the local quantization parameters included in the corresponding baseline global quantization parameters are adjusted to obtain the target global quantization parameters corresponding to the multiple initial frames.

[0155] The control uses the corresponding target global quantization parameters to pre-encode the corresponding initial frame to obtain encoding statistics. The encoding statistics include the loss value corresponding to multiple reconstructed frames, the proportion within the target frame, and the proportion of target skips. The target loss ratio is determined based on the loss values ​​corresponding to multiple reconstructed frames. The target loss ratio is the ratio of the first loss value and the second loss value. The first loss value is the loss value corresponding to the intra-coded frame, and the second loss value is the loss value corresponding to other frames. Based on the target loss ratio, the spatiotemporal coding features and cost information corresponding to multiple initial frames, the predetermined Lagrange multipliers are updated to determine the target Lagrange multipliers corresponding to the intra-coded frames.

[0156] Specifically, let's take an example: For the pre-analysis phase of dynamic Lambda adjustment: Encoding occurs after the preprocessing stage. The system uses a constant quality factor (CRF) as the basic bitrate control strategy, and combines it with the ROImap obtained from the CTU tree analysis in the preprocessing stage for fine-grained control.

[0157] First, a basic quality control baseline is established based on a preset baseline CRF value. Then, local adjustments are made according to the importance of the regions marked in the ROImap: for regions marked as high importance in the ROImap, the local quantization parameter (QP) value is lowered to ensure that these regions receive higher bitrate allocation and better reconstruction quality; for regions marked as low importance in the ROImap, the local quantization parameter (QP) is appropriately increased to save bitrate (similar to the above, adjusting the local quantization parameters included in the corresponding baseline global quantization parameters according to the region of interest mapping maps corresponding to multiple initial frames to obtain the target global quantization parameters corresponding to multiple initial frames).

[0158] At the encoder execution level, these CRF values ​​are adjusted by modifying the local quantization parameter (QP) to ensure that the encoder can accurately execute the quality control strategy.

[0159] Through this CRF-based adaptive quality control mechanism, combined with the importance guidance of ROImap, the system can achieve better bitrate allocation while ensuring subjective quality, thus significantly improving coding efficiency.

[0160] Then, the corresponding initial frame is pre-encoded using the corresponding target global quantization parameters to obtain encoding statistics. These statistics include the loss values ​​for multiple reconstructed frames, the proportion within the target frame, and the proportion of target skips. In other words, statistical information is collected and analyzed after P / B frame encoding.

[0161] It should be noted that dynamically adjusting lambda requires collecting and recording encoding information for all frames, including the loss magnitude, intra (intra) percentage, and skip (skip) percentage. This allows the optimal lambda value for the current frame to be derived based on the available information when encoding a new I-frame.

[0162] First, at the 8x8 block level, the distribution of coding modes is statistically analyzed: the usage percentage of the intra prediction mode in each frame is calculated, and the selection percentage of the skip mode is recorded. For each coded frame, the system calculates the mean squared error (MSE) as a distortion measure: first, the squared difference between the original frame and the reconstructed frame is calculated at the pixel level, and then the average value is calculated at the CTU level and the frame level respectively to obtain the local MSE and global MSE indices. Specifically, the MSE calculation is divided into three channels: luma component (Y) and chroma components (U, V). A weighted average is used to obtain the comprehensive distortion index, and the loss value corresponding to multiple reconstructed frames can be determined by the MSE. At the same time, the system establishes an information statistics table, including statistical data on frame type (I / P / B), to evaluate the stability of coding quality.

[0163] In particular, the system uses intra percentage and skip percentage as important reference indicators: a higher intra percentage usually indicates that the current frame contains more unpredictable content and the overall picture changes more drastically; a higher skip percentage indicates that the overall picture changes more gradually.

[0164] Specifically, the methods for obtaining data in coded statistical data can be as follows: For B-frames: Among the loss values ​​corresponding to multiple reconstructed frames, the loss value corresponding to frame B is: For the statistics of B-frame loss, if the current B-frame is the first B-frame, it is recorded directly. If not, it is calculated using the following formula (similar to the above, in the case of other frames including the current other frames and historical other frames, the second loss value corresponding to other frames is determined based on the loss value corresponding to the current other frames and the loss value corresponding to the historical other frames).

[0165] The intra percentage and skip percentage corresponding to B-frame: For the intra and skip percentages of B-frames, if the current frame is the first B-frame, they are recorded directly; otherwise, they are calculated using the formula below.

[0166] b_last represents the information of the previous B-frame, and b_cur represents the information of the current B-frame. This information will be continuously iterated as the encoding progresses.

[0167] For P-frames: Among the loss values ​​corresponding to multiple reconstructed frames, the loss value corresponding to frame P is: For the statistics of P-frame loss, other parameters of GOP also need to be considered, such as the GOP size, which represents the interval between two P-frames. If the current P-frame is the first P-frame, it is recorded directly; otherwise, it is calculated using the following formula (same as above for determining the image group size corresponding to other frames; based on the current loss value corresponding to other frames, the historical loss values ​​corresponding to other frames, and the image group size, determine the second loss value corresponding to other frames).

[0168] The intra percentage and skip percentage corresponding to P-frames: For the intra and skip ratios of P-frames, if the current frame is the first P-frame, they are recorded directly; otherwise, they are calculated using the formula below.

[0169] p_last represents the information of the previous P-frame, and p_cur represents the information of the current P-frame. This information will be continuously iterated as encoding progresses.

[0170] After pre-analyzing the I-frame, the loss corresponding to the current lambda is obtained.

[0171] Figure 4 This is a schematic diagram illustrating the relationship between the IP frame loss ratio and rate-distortion optimization parameters provided in the optional implementation of this application, such as... Figure 4 As shown in the figure, experimental tests revealed that the loss ratio between I-frames and P-frames is within a relatively optimal range of 0.7-0.8. The optimal lambda was determined by testing the BD-rate performance of the current frame under different lambda values.

[0172] For the loss ratio between I-frames and P-frames: The method for updating the predetermined Lagrange multipliers based on the target loss ratio to obtain the target Lagrange multipliers corresponding to the intra-coded frame can be as follows: The ratio between the new lambda and the previous lambda is related to the loss between I-frames and P-frames as follows: Figure 5 This is a schematic diagram illustrating the correlation between the IB frame loss ratio and rate-distortion optimization parameters provided in the optional implementation of this application, as shown below. Figure 5 As shown in the figure, after experimental testing, it was found that the I-frame and B-frame are in a relatively optimal range of 0.58-0.67.

[0173] For the loss ratio between I-frames and B-frames: The method for updating the predetermined Lagrange multipliers based on the target loss ratio to obtain the target Lagrange multipliers corresponding to the intra-coded frame can be as follows: The ratio between the new lambda and the previous lambda is related to the loss between I-frames and B-frames as follows: When both loss ratios exist, the final target loss ratio can be obtained by averaging the values ​​of lambdaratio (the coefficient after lambdaold) calculated from the MSE of P-frames and B-frames, respectively, and then multiplying it by lambdaold.

[0174] After obtaining the new lambda using the above formula, the range of the new lambda can be limited by the intra ratio and skip ratio (similar to the above case where the target intra-frame ratio and target skip ratio are also included in the encoding statistics, the image complexity is determined based on the target intra-frame ratio and target skip ratio; the Lagrange multiplier limit range is determined based on the image complexity; and the target Lagrange multiplier is determined based on the target loss ratio and the Lagrange multiplier limit range).

[0175] It's important to note that a larger `intra` percentage indicates more drastic changes in the image. For the current I-frame, subsequent frames provide less reference, so the loss of the current I-frame can be appropriately increased, reducing image quality. Correspondingly, in RDO calculation, more emphasis is placed on bitrate, and the lambda value is increased. Conversely, a larger `skip` percentage indicates smoother changes in the image. For the current I-frame, subsequent frames provide more reference, so the loss of the current I-frame can be appropriately reduced, improving image quality. Correspondingly, in RDO calculation, more emphasis is placed on loss, and the lambda value is decreased.

[0176] The method for calculating the integrated intraRatio and skipRatio can be as follows (similar to the above, when other frames include multiple types of coded frames, control the precoding of the corresponding initial frame using the corresponding target global quantization parameters to obtain the initial intra-frame proportion and initial skip proportion corresponding to each of the multiple types of coded frames; determine the adjustment weights corresponding to each of the multiple types of coded frames; determine the target intra-frame proportion based on the initial intra-frame proportion and the corresponding adjustment weights corresponding to each of the multiple types of coded frames, and determine the target skip proportion based on the initial skip proportion and the corresponding adjustment weights corresponding to each of the multiple types of coded frames): The current screen complexity is calculated using intraRatio and skipRatio: When Complexity∈[0,0.35), it is considered a simple picture, and the minimum value of lambdaatio is limited to 0.2, and the maximum value is 1.2; When Complexity ∈ [0.35, 0.7), it is considered a normal image. The minimum value of lambdaatio is limited to 0.5, and the maximum value is 1.5. When Complexity ∈ [0.7, 1], it is considered a complex image. The minimum value of lambdaatio is limited to 0.8, and the maximum value is 1.8. After the above steps, and by adjusting the spatiotemporal domain coding features and cost information, the final lambda can be obtained.

[0177] (III) Coding Stage Based on the corresponding coding reference information, invalid coding modes are eliminated to obtain candidate coding modes corresponding to multiple initial frames respectively; Based on the target Lagrange multipliers corresponding to the intra-coded frames, the target coding mode corresponding to the intra-coded frames is determined from the corresponding candidate coding modes.

[0178] For example: Through the above steps (i) and (ii), the final lambda value is obtained. This value is then used to calculate the RDO of the I-frame and the I-frame is encoded (similar to the above method of determining the target coding mode corresponding to the intra-coded frame based on the target Lagrange multiplier corresponding to the intra-coded frame, in order to encode the I-frame), thus realizing the actual encoding process.

[0179] The following beneficial effects can be achieved through the above optional implementation methods: 1) The optional implementation of this application aims to solve the problem of Lambda parameter optimization in video coding. Specifically, in the prior art, the setting of Lambda parameters mainly relies on fixed mapping relationships or simple adaptive strategies, which cannot fully consider the coding characteristics of different types of video content. The optional implementation of this application establishes an adaptive Lambda adjustment mechanism by collecting intra proportion, skip proportion, and MSE loss information of encoded frames. This allows the encoder to dynamically adjust parameters according to the actual content characteristics, rather than relying on fixed empirical formulas.

[0180] 2) In this optional implementation, the encoding process is divided into two stages: pre-analysis and actual encoding. Content features are obtained through Lookahead and CTU tree analysis to guide subsequent encoding decisions. This architecture ensures encoding efficiency while avoiding the problem of fixed mappings that cannot be dynamically adjusted in traditional methods. Moreover, by adding more statistical features for in-depth analysis, it supports multi-frame joint optimization strategies and can expand the optimization objectives according to specific application requirements.

[0181] 3) The method proposed in this application's optional implementation has controllable computational complexity and is easy to implement in actual hardware. Real-time processing capabilities are ensured through optimized parallel processing and memory access strategies. Accurate distortion control is maintained during encoding, and the predictability of encoded output quality is ensured through comprehensive evaluation of MSE and quality, resulting in an efficient hardware implementation scheme.

[0182] 4) The optional implementation of this application can automatically adjust the encoding strategy according to different types of video content, and achieve dynamic optimization of parameters through real-time analysis of statistical information to adapt to the encoding needs of different scenarios.

[0183] 5) The method provided by the optional implementation of this application can preserve the original coding quality characteristics. By adaptive adjustment based on statistical information, it can maintain the original characteristics of video content, avoid the quality distortion problem caused by fixed parameters in traditional methods, and ensure the visual consistency between the encoded output and the source content.

[0184] 6) The method provided in the optional implementation of this application can eliminate quality fluctuations caused by parameter fluctuations, use global statistical information to guide parameter adjustment, and ensure smooth changes in coding parameters. Through temporal correlation analysis, it avoids abrupt changes between adjacent frames, establishes a stable parameter update mechanism, and prevents drastic fluctuations in coding quality.

[0185] 7) The method provided by the optional implementation of this application can perform a fully automatic parameter optimization mechanism, a statistically driven adjustment strategy without manual intervention, adaptive threshold determination and parameter update, an automatic correction mechanism based on encoding feedback, and a universal solution that supports multiple encoding scenarios.

[0186] According to an embodiment of this application, a device for determining the Lagrange multipliers of encoded frames in a video stream is provided. Figure 6 This application provides a schematic diagram of the structure of a device for determining the Lagrange multipliers of encoded frames in a video stream. Figure 6 As shown, the device includes: a first receiving module 601, a first determining module 602, a first control module 603, a second determining module 604, and a first updating module 605. The device will be described below.

[0187] The first receiving module 601 is used to receive an initial video stream, wherein the initial video stream includes multiple initial frames, and the multiple initial frames include intra-coded frames and other frames; The first determining module 602, connected to the first receiving module 601, is used to determine the target global quantization parameters corresponding to the plurality of initial frames respectively. The first control module 603, connected to the first determining module 602, is used to control the precoding of the corresponding initial frame using the corresponding target global quantization parameters to obtain encoding statistics, wherein the encoding statistics include loss values ​​corresponding to multiple reconstructed frames respectively. The second determining module 604, connected to the first control module 603, is used to determine a target loss ratio based on the loss values ​​corresponding to the plurality of reconstructed frames respectively, wherein the target loss ratio is the ratio of a first loss value and a second loss value, the first loss value is the loss value corresponding to the intra-coded frame, and the second loss value is the loss value corresponding to the other frames; The first update module 605, connected to the second determination module 604, is used to update the predetermined Lagrange multiplier according to the target loss ratio to obtain the target Lagrange multiplier corresponding to the intra-coded frame. The predetermined Lagrange multiplier includes one of the following: historical Lagrange multiplier and reference Lagrange multiplier.

[0188] It should be noted that the first receiving module 601, the first determining module 602, the first control module 603, the second determining module 604, and the first updating module 605 mentioned above correspond to steps S101 to S105 in the embodiments. Multiple modules and their corresponding steps implement the same instances and application scenarios, and their implementation principles and technical effects will not be elaborated further. The specific methods by which each module and unit in the device of the above embodiments performs operations have been described in detail in the embodiments related to the method, and will not be elaborated upon here.

[0189] According to embodiments of this application, an encoding apparatus for encoded frames in a video stream is provided. Figure 7 This application provides a schematic diagram of the structure of an encoding device for encoding frames in a video stream. Figure 7 As shown, the device includes: a second receiving module 701, a third determining module 702, a second control module 703, a fourth determining module 704, a second updating module 705, a fifth determining module 706, and an encoding module 707. The device will be described below.

[0190] The second receiving module 701 is used to receive an initial video stream, wherein the initial video stream includes multiple initial frames, and the multiple initial frames include intra-coded frames and other frames; The third determining module 702, connected to the second receiving module 701, is used to determine the target global quantization parameters corresponding to the plurality of initial frames respectively. The second control module 703, connected to the third determination module 702, is used to control the precoding of the corresponding initial frame using the corresponding target global quantization parameters to obtain encoding statistics, wherein the encoding statistics include loss values ​​corresponding to multiple reconstructed frames respectively. The fourth determining module 704, connected to the second control module 703, is used to determine a target loss ratio based on the loss values ​​corresponding to the plurality of reconstructed frames, wherein the target loss ratio is the ratio of a first loss value and a second loss value, the first loss value is the loss value corresponding to the intra-coded frame, and the second loss value is the loss value corresponding to the other frames. The second update module 705, connected to the fourth determination module 704, is used to update the predetermined Lagrange multiplier according to the target loss ratio to obtain the target Lagrange multiplier corresponding to the intra-coded frame. The predetermined Lagrange multiplier includes one of the following: historical Lagrange multiplier and reference Lagrange multiplier. The fifth determining module 706, connected to the second updating module 705, is used to determine the target coding mode corresponding to the intra-coded frame based on the target Lagrange multiplier corresponding to the intra-coded frame. The encoding module 707, connected to the fifth determining module 706, is used to encode the intra-coded frame according to the target encoding mode corresponding to the intra-coded frame, and output a single-frame encoded stream corresponding to the intra-coded frame.

[0191] It should be noted that the second receiving module 701, the third determining module 702, the second control module 703, the fourth determining module 704, the second updating module 705, the fifth determining module 706, and the encoding module 707 mentioned above correspond to steps S201 to S207 in the embodiments. Multiple modules and their corresponding steps implement the same instances and application scenarios, and their implementation principles and technical effects will not be elaborated further. The specific methods by which each module and unit in the device of the above embodiments perform operations have been described in detail in the embodiments related to the method, and will not be elaborated upon here.

[0192] This application also provides a computer device, which may include a storage component and a processing component; The storage component has one or more computer instructions, wherein the one or more computer instructions are invoked and executed by the processing component to implement the method described in any one of the above.

[0193] Of course, computer equipment may also include other components, such as input / output interfaces, display components, communication components, etc.

[0194] Input / output interfaces provide interfaces between processing components and peripheral interface modules, which can be output devices, input devices, etc. Communication components are configured to facilitate wired or wireless communication between computing devices and other devices.

[0195] The processing component may include one or more processors to execute computer instructions to complete all or part of the steps in the above-described method. Alternatively, the processing component may be implemented as one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field-programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components to perform the above-described method.

[0196] Storage components are configured to store various types of data to support operations on the terminal. Storage components can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk.

[0197] The display component can be an electroluminescent (EL) element, a liquid crystal display or a microdisplay with a similar structure, or a retina-direct display or a similar laser scanning display.

[0198] It should be noted that the aforementioned computing device implementation method or processing method can be a physical device or an elastic computing host provided by a cloud computing platform. It can be implemented as a distributed cluster composed of multiple servers or terminal devices, or as a single server or a single terminal device.

[0199] When the above-mentioned computing device implements the above method, it can be specifically implemented as an electronic device. An electronic device can refer to a device used by a user that has the computing, Internet access, communication and other functions required by the user, such as a mobile phone, tablet computer, personal computer, wearable device, etc.

[0200] It should be noted that the aforementioned computing devices can be physical devices or elastic computing hosts provided by cloud computing platforms. They can be implemented as a distributed cluster of multiple servers or terminal devices, or as a single server or a single terminal device.

[0201] This application also provides a computer-readable storage medium storing a computer program that, when executed by a computer, can implement the above-described method. This computer-readable medium may be included in the electronic device described in the above embodiments; alternatively, it may exist independently and not be assembled into the electronic device.

[0202] This application also provides a computer program product comprising a computer program carried on a computer-readable storage medium, which, when executed by a computer, can implement the methods described above. In such an embodiment, the computer program may be downloaded and installed from a network, and / or installed from a removable medium. When the computer program is executed by a processor, it performs the various functions defined in the system of this application.

[0203] It should be noted that the embodiments of this application may involve the use of user data. In practical applications, user-specific personal data may be used in the scheme described herein within the scope permitted by applicable laws and regulations, provided that it complies with the applicable laws and regulations of the country (e.g., with the user's explicit consent, with the user being properly notified, etc.).

[0204] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.

[0205] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.

[0206] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.

[0207] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application.

Claims

1. A method for determining a Lagrange multiplier of a coded frame in a video stream, characterized in that, The method comprises: receiving an initial video stream, wherein the initial video stream comprises a plurality of initial frames, and the plurality of initial frames comprise an intra-coded frame and other frames; determining a target global quantization parameter corresponding to each of the plurality of initial frames; controlling pre-encoding of the corresponding initial frame using the corresponding target global quantization parameter to obtain encoding statistics, wherein the encoding statistics comprise loss values corresponding to a plurality of reconstructed frames respectively; determining a target loss ratio value according to the loss values corresponding to the plurality of reconstructed frames respectively, wherein the target loss ratio value is a ratio of a first loss value and a second loss value, the first loss value is a loss value corresponding to the intra-coded frame, and the second loss value is a loss value corresponding to the other frames; updating a predetermined Lagrange multiplier according to the target loss ratio value to obtain a target Lagrange multiplier corresponding to the intra-coded frame, wherein the predetermined Lagrange multiplier comprises one of a historical Lagrange multiplier and a reference Lagrange multiplier.

2. The method of claim 1, wherein, updating a predetermined Lagrange multiplier according to the target loss ratio value to obtain a target Lagrange multiplier corresponding to the intra-coded frame, comprising: in a case where the encoding statistics further comprise a target intra-coded frame ratio and a target skip ratio, determining a picture complexity according to the target intra-coded frame ratio and the target skip ratio; determining a Lagrange multiplier limit range according to the picture complexity; determining the target Lagrange multiplier according to the target loss ratio value, the predetermined Lagrange multiplier and the Lagrange multiplier limit range.

3. The method of claim 1, wherein, obtaining the encoding statistics, comprising: in a case where the other frames comprise a current other frame and a historical other frame, determining the second loss value corresponding to the other frames according to a loss value corresponding to the current other frame and a loss value corresponding to the historical other frame.

4. The method of claim 3, wherein, determining the second loss value corresponding to the other frames according to a loss value corresponding to the current other frame and a loss value corresponding to the historical other frame, comprising: determining a group of pictures size corresponding to the other frames; determining the second loss value corresponding to the other frames according to the loss value corresponding to the current other frame, the loss value corresponding to the historical other frame and the group of pictures size.

5. The method of claim 2, wherein, controlling pre-encoding of the corresponding initial frame using the corresponding target global quantization parameter to obtain the encoding statistics corresponding to the plurality of reconstructed frames, comprising: in a case where the other frames comprise a plurality of type coding frames, controlling pre-encoding of the corresponding initial frame using the corresponding target global quantization parameter to obtain an initial intra-coded frame ratio and an initial skip ratio corresponding to the plurality of type coding frames respectively; determining an adjustment weight corresponding to each of the plurality of type coding frames; determining the target intra-coded frame ratio according to the initial intra-coded frame ratio corresponding to each of the plurality of type coding frames and the corresponding adjustment weight, and determining the target skip ratio according to the initial skip ratio corresponding to each of the plurality of type coding frames and the corresponding adjustment weight.

6. The method of claim 1, wherein, determining the target global quantization parameter corresponding to each of the plurality of initial frames, comprising: determining a region of interest mapping corresponding to each of the plurality of initial frames; According to the region-of-interest mapping diagram corresponding to the plurality of initial frames respectively, local quantization parameters included in corresponding reference global quantization parameters are adjusted respectively to obtain target global quantization parameters corresponding to the plurality of initial frames respectively.

7. The method of claim 1, wherein, According to the target loss ratio, a predetermined Lagrange multiplier is updated to obtain a target Lagrange multiplier corresponding to the intra-frame coded frame, including: Determine the spatio-temporal domain coding features corresponding to the plurality of initial frames respectively; According to the target loss ratio and the spatio-temporal domain coding features corresponding to the plurality of initial frames respectively, the predetermined Lagrange multiplier is updated to determine the target Lagrange multiplier corresponding to the intra-frame coded frame.

8. The method of claim 7, wherein, Determine the spatio-temporal domain coding features corresponding to the plurality of initial frames respectively, including: Split the target frame to obtain a plurality of coding tree units, wherein the target frame is any one of the plurality of initial frames; Based on the buffer queue and the sliding window mechanism, the plurality of coding tree units are analyzed in time domain and space domain to obtain the spatio-temporal domain coding features corresponding to the target frame.

9. The method of claim 1, wherein, According to the target loss ratio, a predetermined Lagrange multiplier is updated to obtain a target Lagrange multiplier corresponding to the intra-frame coded frame, including: Determine the cost information corresponding to the plurality of initial frames respectively; According to the cost information and the target loss ratio, the predetermined Lagrange multiplier is updated to determine the target Lagrange multiplier corresponding to the intra-frame coded frame.

10. The method of claim 1, wherein, According to the target loss ratio, a predetermined Lagrange multiplier is updated to obtain a target Lagrange multiplier corresponding to the intra-frame coded frame, including: According to the target Lagrange multiplier corresponding to the intra-frame coded frame, determine the target coding mode corresponding to the intra-frame coded frame.

11. The method of claim 10, wherein, According to the target Lagrange multiplier corresponding to the intra-frame coded frame, determine the target coding mode corresponding to the intra-frame coded frame, including: Determine the coding reference information corresponding to the plurality of initial frames respectively; According to the corresponding coding reference information, exclude invalid coding modes to obtain candidate coding modes corresponding to the plurality of initial frames respectively; According to the target Lagrange multiplier corresponding to the intra-frame coded frame, determine the target coding mode corresponding to the intra-frame coded frame from the corresponding candidate coding modes.

12. The method of claim 10, wherein, According to the target Lagrange multiplier corresponding to the intra-frame coded frame, determine the target coding mode corresponding to the intra-frame coded frame, including: In the case where the intra-frame coded frame includes a plurality of frames, according to the picture group size corresponding to the other frames, determine the target intra-frame coded frame corresponding to the other frames from the plurality of intra-frame coded frames; According to the target coding mode corresponding to the target intra-frame coded frame and the coding mode corresponding to the other frames in pre-encoding, obtain the coding mode corresponding to the other frames.

13. The method of claim 12, wherein, According to the target coding mode corresponding to the target intra-frame coded frame and the coding mode corresponding to the other frames in pre-encoding, obtain the coding mode corresponding to the other frames, including: In the case where the intra-frame coded frame includes a plurality of frames, according to the picture group size corresponding to the other frames, determine the target intra-frame coded frame corresponding to the other frames from the plurality of intra-frame coded frames; According to the target coding mode corresponding to the target intra-frame coded frame and the coding mode corresponding to the other frames in pre-encoding, obtain the coding mode corresponding to the other frames. According to the target coding mode corresponding to the target intra-frame coded frame and the coding mode corresponding to the other frames in pre-encoding, obtain the coding mode corresponding to the other frames, including: encoding the intra-coded frame according to a target coding mode corresponding to the intra-coded frame, outputting a first single-frame encoded stream, and encoding the other frames according to coding modes corresponding to the other frames, obtaining a second single-frame encoded stream; integrating the first single-frame encoded stream and the second single-frame encoded stream, obtaining a target video frame.

14. The method according to any one of claims 1 to 13, characterized in that, After receiving the initial video stream, further comprising: dividing the initial video stream, obtaining a plurality of image group structures; sequentially pre-encoding the plurality of image group structures according to frame sequence arrangement order, obtaining corresponding encoding statistical data; updating a target Lagrange multiplier corresponding to a current image group structure according to the corresponding encoding statistical data; taking the target Lagrange multiplier corresponding to the current image group structure as a reference distortion tradeoff index of a next group of image group structures.

15. A method of encoding an encoded frame in a video stream, characterized by, comprising: receiving an initial video stream, wherein the initial video stream comprises a plurality of initial frames, and the plurality of initial frames comprise an intra-coded frame and other frames; determining target global quantization parameters corresponding to the plurality of initial frames respectively; controlling pre-encoding of corresponding initial frames using corresponding target global quantization parameters, obtaining encoding statistical data, wherein the encoding statistical data comprises loss values corresponding to a plurality of reconstructed frames respectively; determining a target loss ratio value according to loss values corresponding to the plurality of reconstructed frames respectively, wherein the target loss ratio value is a ratio of a first loss value and a second loss value, the first loss value is a loss value corresponding to the intra-coded frame, and the second loss value is a loss value corresponding to the other frames; updating a predetermined Lagrange multiplier according to the target loss ratio value, obtaining a target Lagrange multiplier corresponding to the intra-coded frame, wherein the predetermined Lagrange multiplier comprises one of the following: a historical Lagrange multiplier, a reference Lagrange multiplier; determining a target coding mode corresponding to the intra-coded frame according to the target Lagrange multiplier corresponding to the intra-coded frame; encoding the intra-coded frame according to the target coding mode corresponding to the intra-coded frame, outputting a single-frame encoded stream corresponding to the intra-coded frame.

16. A computing device, comprising: comprising a processing component and a storage component; the storage component stores one or more computer instructions; the one or more computer instructions are used to be called and executed by the processing component to implement the method of any one of claims 1 to 15.

17. A computer program product, characterised in that, comprising computer programs / instructions, which are executed by the processing component to implement the method of any one of claims 1 to 15.