Image encoding method, apparatus, and storage medium

By performing semantic recognition and scene awareness on video frames, the comprehensive importance score of video frames is determined, and the QP of image blocks is adjusted. This solves the storage and transmission pressure problem caused by the large amount of video data, and achieves efficient compression and quality optimization of video encoding.

CN122093565BActive Publication Date: 2026-07-21HANGZHOU HIKVISION DIGITAL TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
HANGZHOU HIKVISION DIGITAL TECHNOLOGY CO LTD
Filing Date
2026-04-22
Publication Date
2026-07-21

Smart Images

  • Figure CN122093565B_ABST
    Figure CN122093565B_ABST
Patent Text Reader

Abstract

The application provides an image coding method, device and storage medium. In one example, the method comprises: performing semantic recognition on a current video frame to determine semantic importance scores of targets; determining frame-level semantic importance scores according to the semantic importance scores of the targets, and determining frame-level scene dynamic scores according to scene perception information of the current video frame; thereby determining frame-level comprehensive importance scores; determining a reference QP of the current video frame according to the frame-level comprehensive importance scores; determining QP offsets of image blocks in the current video frame according to signal statistical features of the image blocks, and determining QPs of the image blocks according to the reference QP and the QP offsets; and coding the current video frame according to the QPs of the image blocks. The method can optimize the balance between compression efficiency and video quality.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of video coding technology, and in particular to an image coding method, device and storage medium. Background Technology

[0002] With the widespread adoption of high-definition and ultra-high-definition video applications, video data volume has exploded, putting enormous pressure on storage space and network transmission bandwidth. Against this backdrop, video encoding technology has become a key means of reducing data volume and alleviating storage and transmission pressure, with compression efficiency being paramount. As a crucial component of the video encoder, the bitrate control module is responsible for optimizing the subjective or objective quality of the reconstructed video as much as possible by rationally allocating encoded bits under a given bitrate constraint, thereby achieving an effective balance between compression efficiency and video quality. Summary of the Invention

[0003] In view of this, this application provides an image encoding method, device and storage medium.

[0004] Specifically, this application is implemented through the following technical solution: According to a first aspect of the embodiments of this application, an image encoding method is provided, comprising: Perform semantic recognition on the current video frame to determine the semantic importance score of each target in the current video frame; Based on the semantic importance scores of each target in the current video frame, the frame-level semantic importance score of the current video frame is determined, and based on the scene perception information of the current video frame, the frame-level scene dynamics score of the current video frame is determined. Based on the frame-level semantic importance score of the current video frame and the frame-level scene dynamism score of the current video frame, the frame-level comprehensive importance score of the current video frame is determined. The baseline QP of the current video frame is determined based on the frame-level comprehensive importance score of the current video frame; Based on the signal statistical characteristics of each image block in the current video frame, the QP offset of each image block in the current video frame is determined, and based on the reference QP of the current video frame and the QP offset of each image block in the current video frame, the QP of each image block in the current video frame is determined. The current video frame is encoded based on the QP of each image block.

[0005] According to a second aspect of the embodiments of this application, an image encoding apparatus is provided, comprising: The first determining unit is used to perform semantic recognition on the current video frame and determine the semantic importance score of each target in the current video frame; The second determining unit is used to determine the frame-level semantic importance score of the current video frame based on the semantic importance scores of each target in the current video frame, and to determine the frame-level scene dynamics score of the current video frame based on the scene perception information of the current video frame. The second determining unit is further configured to determine the frame-level comprehensive importance score of the current video frame based on the frame-level semantic importance score of the current video frame and the frame-level scene dynamics score of the current video frame. The third determining unit is used to determine the baseline QP of the current video frame based on the frame-level comprehensive importance score of the current video frame. The fourth determining unit is used to determine the QP offset of each image block in the current video frame based on the signal statistical characteristics of each image block in the current video frame, and to determine the QP of each image block in the current video frame based on the reference QP of the current video frame and the QP offset of each image block in the current video frame. The encoding unit is used to encode the current video frame based on the QP of each image block of the current video frame.

[0006] According to a third aspect of the present application, an electronic device is provided, including a processor and a memory, the memory storing machine-executable instructions executable by the processor, the processor being configured to execute the machine-executable instructions to implement the method provided in the first aspect.

[0007] According to a fourth aspect of the embodiments of this application, a machine-readable storage medium is provided, wherein machine-executable instructions are stored therein, and when the machine-executable instructions are executed by a processor, the method provided in the first aspect is implemented.

[0008] According to a fifth aspect of the embodiments of this application, a monitoring front-end device is provided, comprising: a lens, an image sensor, and a processor, wherein the processor includes a scene perception module, a target semantic recognition module, and a signal analysis module; wherein: The lens is used to receive incident light; The image sensor is used to receive an optical image focused by the lens and convert the optical image into a raw image electrical signal; The scene perception module includes an image processing unit, which processes the original image electrical signal to generate a digital image signal, and transmits the digital image signal to the target semantic recognition module and the signal analysis module respectively. The scene perception module further includes a background modeling unit. The image processing unit and the background modeling unit are used to acquire scene perception information and transmit it to the signal analysis module. The target semantic recognition module includes an intelligent processing unit, which is used to perform semantic recognition on the digital image signal, extract target semantic information, and transmit the target semantic information to the signal analysis module; The signal analysis module includes a video encoding processing unit, which performs video encoding using the method provided in the first aspect to obtain a video encoded bitstream; wherein the video encoded bitstream is used for local file storage and / or network transmission.

[0009] According to a sixth aspect of the embodiments of this application, a storage device is provided, comprising: a processor and a storage unit, wherein the processor includes a scene perception module, a target semantic recognition module, and a signal analysis module; wherein: The scene perception module includes a decoding processing unit, which is used to decode the received video stream, transmit the obtained video image data to the signal analysis module and the target semantic recognition module respectively, and acquire scene perception information and transmit the scene perception information to the signal analysis module. The target semantic recognition module includes an intelligent processing unit, which is used to perform semantic recognition on the video image data, extract target semantic information, and transmit the target semantic information to the signal analysis module; The signal analysis module includes a video encoding processing unit, which is used to perform video encoding by executing the method provided in the first aspect to obtain a video encoded bitstream; The storage unit is used to store the video encoded bitstream.

[0010] The technical solution provided in this application can bring at least the following beneficial effects: By performing semantic recognition on the current video frame, the semantic importance score of each target in the current video frame is determined. Based on the semantic importance scores of each target in the current video frame, the frame-level semantic importance score of the current video frame is determined. Furthermore, based on the scene perception information of the current video frame, the frame-level scene dynamics score of the current video frame is determined. Based on the frame-level semantic importance score and the frame-level scene dynamics score of the current video frame, the frame-level comprehensive importance score of the current video frame is determined. Based on the frame-level comprehensive importance score of the current video frame, the baseline QP of the current video frame is determined. Then, based on the current... By analyzing the signal statistical characteristics of each image block in the previous video frame, the QP offset of each image block in the current video frame is determined. Based on the baseline QP of the current video frame and the QP offset of each image block in the current video frame, the QP of each image block in the current video frame is determined. Based on the QP of each image block in the current video frame, the current video frame is encoded. This constructs a multi-layer rate-distortion optimization decision system from the "semantic layer" to the "signal layer," realizing refined and adaptive allocation of bitrate resources from the target level to the frame level and then to the image block level, thus optimizing the balance between compression efficiency and video quality. Attached Figure Description

[0011] Figure 1 This is an example diagram of the temporal reference relationship that includes I-frames, P-frames, and B-frames; Figure 2 This is a schematic flowchart illustrating an image encoding method according to an exemplary embodiment of this application; Figure 3 This is a schematic diagram illustrating the implementation process of a multi-layer rate-distortion optimization coding scheme according to an exemplary embodiment of this application; Figure 4 This is a schematic diagram illustrating a scene perception layer processing flow according to an exemplary embodiment of this application; Figure 5 This is a schematic diagram of a signal statistics layer control process illustrated in an exemplary embodiment of this application; Figure 6 This is a schematic diagram illustrating the process of determining the QP of an image block according to an exemplary embodiment of this application; Figure 7 This is a schematic diagram illustrating the structure of a monitoring front-end device according to an exemplary embodiment of this application; Figure 8 This is a schematic diagram illustrating the structure of a storage device according to an exemplary embodiment of this application; Figure 9 This is a schematic diagram illustrating the structure of an image encoding device according to an exemplary embodiment of this application; Figure 10 This is a schematic diagram of the hardware structure of an electronic device provided in an embodiment of this application. Detailed Implementation

[0012] To enable those skilled in the art to better understand the technical solutions provided in the embodiments of this application, some technical terms involved in the embodiments of this application will be explained below.

[0013] 1. Video encoding: Compressing video data to remove as much redundant data as possible and reduce the amount of compressed (i.e., encoded) data.

[0014] 2. GOP (Group of Pictures) format: A video sequence consists of several time-series consecutive images. When compressing it, the video sequence is first divided into several small groups of images, each of which is called a GOP.

[0015] 3. I-frame: Intra-coded frame, an independent frame containing all its own information, which can be decoded independently without referring to other images. The first video coded frame in a GOP is usually an I-frame.

[0016] 4. P-frame: Inter-frame predictive coded frame, which performs motion estimation by referencing previous frames. Decoding requires superimposing the differences defined in this frame onto previously buffered frames to generate the final image.

[0017] 5. B-frame: Bidirectional predictive coded frame, which records the difference between the current frame and the frames before and after it. Decoding a B-frame requires obtaining the previously buffered frame and the decoded frame. The final frame is obtained by superimposing the data from the previous and current frames with the data from the previous and current frames.

[0018] For example, the reference relationship between I-frames, P-frames, and B-frames can be as follows: Figure 1 As shown, the arrows point to the reference frames used for each frame.

[0019] 6. Breathing effect: Within a GOP, the further away a P-frame is from an I-frame, the larger the coding error and the lower the image quality. When a new GOP appears, the image quality suddenly improves due to the insertion of I-frames, and the image quality of subsequent P-frames gradually decreases again, resulting in a periodic change in image quality from good to bad. This phenomenon is called the breathing effect.

[0020] 7. PSNR (Peak Signal to Noise Ratio): The most basic video quality assessment method. Its value is generally between 20 and 50; the higher the value, the closer the damaged image is to the original image. PSNR calculates the error between pixels in the two images by comparing the original and distorted images pixel by pixel, and ultimately determines the quality score of the distorted image based on these errors.

[0021] 8. Encoding CTU: Tree-shaped encoding unit in the H.265 protocol, with sizes supporting 16×16, 32×32, and 64×64.

[0022] 9. Encoding MB (Micro Block): The basic processing unit for prediction and transform coding in the H.264 protocol.

[0023] 10. Encoding CU Block: The CU is the basic unit of predictive coding in the H.265 protocol, and its size supports 8×8, 16×16, 32×32 and 64×64.

[0024] 11. Skip coding mode: A type of inter-frame coding mode. Skip mode does not transmit residuals, but only skip_flag and merge_index. Skip coding can save bitrate, but the image quality will be degraded.

[0025] 12. Encoding QP (Quantization Parameter): During the video encoding quantization process, the transformed frequency domain coefficients (such as DCT coefficients) are reduced by a certain step size, discarding high-frequency details and reducing the amount of data. Adjusting the QP can control the quality and compression ratio.

[0026] Quantization is the main step in the encoding process that causes compression loss (distortion). QP controls the step size of quantization.

[0027] For example, a smaller QP value means a smaller quantization step size, retaining more details and resulting in higher image quality, but also a larger file size (requiring a higher bitrate). A larger QP value means a larger quantization step size, discarding more detail information and resulting in higher compression ratio, but also a more significant loss of image quality (greater distortion).

[0028] For example, the process of determining the optimal QP is itself a form of rate-distortion optimization.

[0029] 13. Rate-Distortion Optimization (RDO): A technique that seeks the optimal coding decision to minimize the distortion of the reconstructed signal relative to the original signal given a code rate (bit rate) budget.

[0030] The goal of rate-distortion optimization is to achieve the best image quality given a bitrate budget; or, to achieve the target image quality using the least bitrate.

[0031] To make the above-mentioned objectives, features and advantages of the embodiments of this application more apparent and understandable, the technical solutions of the embodiments of this application will be further described in detail below with reference to the accompanying drawings.

[0032] It should be noted that the sequence number of each step in the embodiments of this application does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.

[0033] Please see Figure 2 This is a flowchart illustrating an image encoding method provided in an embodiment of this application. For example, this image encoding method can be applied to a video front-end device (such as an IPC). The video front-end device can use this image encoding method to encode the acquired video and transmit the encoded video data via a network to a video back-end device (such as an NVR or CVR) and / or store it in a local storage device; or, this image encoding method can be applied to a video back-end device (such as an NVR or CVR). The video back-end device can decode the received bitstream sent by the video front-end device, encode it using the image encoding method, and store the encoded video data, such as... Figure 2As shown, the image encoding method may include the following steps: Step S200: Perform semantic recognition on the current video frame to determine the semantic importance score of each target in the current video frame.

[0034] For example, under normal circumstances, video frames in a video can be divided into GOPs according to a fixed number of frames, such as one GOP every 750 frames. A GOP includes one I-frame (usually the first video frame); where the first video frame is usually an I-frame.

[0035] In this embodiment of the application, considering that different types / states of targets (also called objects) may exist in different video frames, and that the importance of targets of different types / states will also differ in the encoding process, for B-frames or P-frames, in order to improve the rationality of the determination of encoding parameters, the semantic importance of each target in the video frame can be determined based on the type / state of each target in the video frame. This semantic importance can be used to determine the encoding parameters of the video frame.

[0036] For example, the semantic importance of a target can be characterized by a semantic importance score. The semantic importance of a target can be positively correlated with its semantic importance score.

[0037] For example, the type of target may include, but is not limited to, foreground or background. Foreground targets may be further categorized, such as pedestrians, vehicles, or animals. The state of the target may include, but is not limited to, movement or stillness.

[0038] Step S210: Based on the semantic importance scores of each target in the current video frame, determine the frame-level semantic importance score of the current video frame, and based on the scene perception information of the current video frame, determine the frame-level scene dynamics score of the current video frame.

[0039] Step S220: Determine the overall frame importance score of the current video frame based on the frame-level semantic importance score and the frame-level scene dynamism score of the current video frame.

[0040] In this embodiment, the semantic importance of each target in a video frame can be used to determine the importance of the video frame.

[0041] In determining the importance of video frames, scene-aware information of the video frames can also be referenced.

[0042] For example, scene-aware information may include, but is not limited to, some or all of the information such as brightness, motion vector magnitude, global motion percentage, and texture complexity.

[0043] For example, for any video frame, the brightness information of the video frame may include the mean brightness and / or the variance of brightness of all pixels in the video frame. The mean brightness reflects the overall brightness of the frame, and the variance of brightness reflects the contrast of the lighting within the frame.

[0044] The motion vector amplitude of a video frame can be estimated by comparing the current video frame with the previous reference frame, obtaining the motion vector amplitude of each image block in the current video frame, and calculating the average amplitude and histogram of all non-zero motion vectors. A larger average amplitude indicates more intense motion in the video frame.

[0045] The global motion percentage of a video frame can be determined based on the percentage of pixels in the current video frame whose motion vectors relative to the previous video frame exceed a preset motion vector threshold.

[0046] The texture complexity of a video frame can be calculated by measuring the gradient variance or the variance of the DCT (Discrete Cosine Transform) coefficients of the current video frame.

[0047] It should be noted that since the encoding processing unit usually encodes each image block individually, it cannot directly obtain the scene perception information of the entire video frame. Therefore, before the encoding processing unit performs the encoding processing, the scene perception information of the video frame can be determined first, and the determined scene perception information can be used as the input information of the encoding processing unit, so that the encoding processing unit can determine the QP of each image block in the video frame according to the scheme provided in the embodiments of this application based on the input scene perception information.

[0048] For example, the importance of a video frame can be characterized by its frame-level comprehensive importance score. The importance of a video frame can be positively correlated with its frame-level comprehensive importance score.

[0049] Accordingly, for the current video frame, on the one hand, the frame-level semantic importance score of the current video frame can be determined based on the semantic importance scores of each target in the current video frame; on the other hand, the frame-level scene dynamics score of the current video frame can be determined based on the scene perception information of the current video frame. Furthermore, the frame-level comprehensive importance score of the current video frame can be determined based on the frame-level semantic importance score and the frame-level scene dynamics score of the current video frame.

[0050] For example, the frame-level comprehensive importance score of a video frame can be positively correlated with the frame-level semantic importance score and the frame-level scene dynamics score of the video frame, respectively.

[0051] Step S230: Determine the baseline QP of the current video frame based on the frame-level comprehensive importance score of the current video frame.

[0052] Step S240: Based on the signal statistical characteristics of each image block in the current video frame, determine the QP offset of each image block in the current video frame, and based on the reference QP of the current video frame and the QP offset of each image block in the current video frame, determine the QP of each image block in the current video frame.

[0053] In this embodiment of the application, for a video frame, the encoding parameters (such as QP) of each image block (also called a coding block) in the video frame can be determined in the form of image blocks.

[0054] For example, the image block may include a CU, or sub-blocks further subdivided from the CU, or a macroblock of a specified size.

[0055] The size of the exemplary image patch can be preset, such as 4×4 or 8×8, etc.

[0056] For example, the QP of each image block in a video frame can be determined by taking the frame-level QP of the video frame as a benchmark and combining it with the signal statistical characteristics of each image block.

[0057] For example, the QP offset of each image block in the current video frame can be determined based on the signal statistical characteristics of each image block in the current video frame, and the QP of each image block in the current video frame can be determined based on the reference QP of the current video frame and the QP offset of each image block in the current video frame.

[0058] For example, for any image block in the current video frame, the QP of that image block can be the sum of the reference QP of the current video frame and the QP offset of that image block.

[0059] For example, signal statistical features may include local spatial features (such as texture complexity) and / or temporal features (such as motion vectors) of image patches.

[0060] For example, the frame-level QP (which may be called the baseline QP) of a video frame can be determined based on the overall frame-level importance score of the video frame.

[0061] Step S250: Encode the current video frame according to the QP of each image block in the current video frame.

[0062] In this embodiment of the application, when the QP of each image block of the current video frame is determined, the current video frame can be encoded based on the QP of each image block of the current video frame.

[0063] It can be seen that, in Figure 2In the illustrated method flow, semantic recognition is performed on the current video frame to determine the semantic importance score of each target in the current video frame. Based on the semantic importance scores of each target in the current video frame, the frame-level semantic importance score of the current video frame is determined. Furthermore, based on the scene perception information of the current video frame, the frame-level scene dynamism score of the current video frame is determined. Based on the frame-level semantic importance score and the frame-level scene dynamism score of the current video frame, the frame-level comprehensive importance score of the current video frame is determined. Finally, based on the frame-level comprehensive importance score of the current video frame, the baseline QP of the current video frame is determined. Based on the signal statistical characteristics of each image block in the current video frame, the QP offset of each image block in the current video frame is determined. Based on the baseline QP of the current video frame and the QP offset of each image block in the current video frame, the QP of each image block in the current video frame is determined. Based on the QP of each image block in the current video frame, the current video frame is encoded. This constructs a multi-layer rate-distortion optimization decision system from the "semantic layer" to the "signal layer", realizing fine-grained and adaptive allocation of bitrate resources from the target level to the frame level and then to the image block level, thus optimizing the balance between compression efficiency and video quality.

[0064] In some embodiments, the above-described semantic recognition of the current video frame to determine the semantic importance score of each target in the current video frame may include: For any target in the current video frame, determine the semantic importance score of the target based on its multi-dimensional features.

[0065] For example, the semantic importance of targets in a video frame can be evaluated from multiple dimensions.

[0066] For example, the multidimensional features of a target may include, but are not limited to, at least two of the following features: Foreground / background, motion state, semantic category, and relative area.

[0067] The foreground / background is used to characterize whether the target belongs to the foreground or background (such as sky, buildings, or trees).

[0068] Motion state is used to characterize whether a target is in a static or dynamic state.

[0069] Semantic categories are used to represent the category of a target, such as pedestrians, vehicles, or animals.

[0070] Relative area is used to characterize the proportion of the target's bounding box area in the area of ​​the video frame (the area of ​​the entire screen).

[0071] In one example, determining the semantic importance score of the target based on its multi-dimensional features could include: Based on the multi-dimensional characteristics of the target, the importance score of each dimension of the target is determined respectively; The semantic importance score of the objective is determined by the weighted sum of the importance scores of each dimension.

[0072] For example, the weighting coefficients of the importance scores of each dimension of the target are determined according to the application scenario. Different application scenarios may allow different weighting coefficients of the importance scores of each dimension of the target.

[0073] For example, in traffic monitoring scenarios, the weighting coefficient of the importance score of the semantic category dimension can be increased; in perimeter security scenarios, the weighting coefficient of the importance score of the motion state dimension can be increased.

[0074] As an example, determining the importance score of each dimension of the target based on its multi-dimensional characteristics can include: When multi-dimensional features include relative area, the relative area importance score of the target is determined based on the absolute value of the difference between the area ratio of the target bounding box in the current video frame and the area ratio of the most concerned area; wherein, the relative area importance score of the target is negatively correlated with the absolute value.

[0075] For example, during the encoding of video frames, we can focus on targets with a medium area ratio and reduce attention to targets that are too large or too small, so as to better balance the compression efficiency and video quality of video frames.

[0076] Accordingly, the importance score for the relative area dimension (which can be called the relative area importance score) can be determined based on the absolute value of the difference between the target area percentage and the most concerned area percentage (which can be set according to the actual scenario, and its specific value can be an empirical value, such as 10%).

[0077] For example, the greater the difference between the actual area percentage of a target and the area percentage of the most concerned area (the larger the absolute value of the difference), the lower the relative area importance score of the target.

[0078] For example, the relative area importance score of an objective can be determined as follows: Where r is the proportion of the target bounding box area to the total screen area (i.e., the actual area proportion), where 0 < r ≤ 1; μ is the proportion of the most attention area, for example, μ = 0.1; σ is used to control the width of the attention interval, for example, σ is set to 0.05.

[0079] In some embodiments, determining the frame-level semantic importance score of the current video frame based on the semantic importance scores of each target in the current video frame may include: The frame-level semantic importance score of the current video frame is determined based on the semantic importance score of each target in the current video frame and the area of ​​the target bounding box of each target.

[0080] The above-mentioned determination of the overall frame importance score of the current video frame based on the frame-level semantic importance score and the frame-level scene dynamism score of the current video frame may include: Based on the set fusion weights, the frame-level semantic importance score and the frame-level scene dynamics score of the current video frame are fused, and the resulting fusion score is determined as the frame-level comprehensive importance score of the current video frame.

[0081] For example, in order to avoid a single small target dominating the decision of the whole frame, in the process of determining the frame-level semantic importance score of the video frame based on the semantic importance score of each target in the current video frame, the semantic importance score of each target can be weighted and adjusted by combining the target bounding box area of ​​each target.

[0082] In one example, determining the frame-level semantic importance score of the current video frame based on the semantic importance scores of each object in the current video frame and the area of ​​the bounding box of each object can include: For any target in the current video frame, the weighting coefficient of the semantic importance score of the target is determined based on the ratio of the area of ​​the target's bounding box to the sum of the areas of the bounding boxes of all targets in the current video frame. The weighted sum of the semantic importance scores of each target in the current video frame is used to determine the frame-level semantic importance score of the current video frame.

[0083] For example, the frame-level semantic importance score of a video frame can be determined in the following way: in, The semantic importance score of target i. Let N be the area of ​​the bounding box of target i, and N be the number of targets detected in the current video frame.

[0084] It should be noted that when N=0, that is, when no target is detected in the current video frame, This is the default value. For example, the default value can be a small value, such as 0.1.

[0085] In some embodiments, scene-aware information includes global motion percentage and texture complexity; The above-mentioned determination of the frame-level scene dynamics score of the current video frame based on the scene perception information of the current video frame may include: The proportion of pixels in the current video frame whose motion vectors relative to the previous video frame exceed a preset motion vector threshold is determined as the global motion proportion of the current video frame; and, Determine the texture complexity of the current video frame; The weighted sum of the global motion percentage and texture complexity of the current video frame is used to determine the frame-level scene dynamics score of the current video frame.

[0086] For example, taking scene-aware information including global motion percentage and texture complexity as an example, the frame-level scene dynamics score of a video frame can be determined based on the weighted sum of the global motion percentage and texture complexity of the video frame.

[0087] For example, the global motion percentage of the current video frame can be the percentage of pixels in the current video frame whose motion vectors relative to the previous video frame exceed a preset motion vector threshold.

[0088] For example, the texture complexity of the current video frame can be characterized by calculating the gradient variance or DCT coefficient variance of the current video frame.

[0089] For example, the frame-level scene dynamics score of the current video frame is positively correlated with the global motion ratio and texture complexity of the current video frame, respectively.

[0090] For example, the frame-level scene dynamics score of the current video frame can be determined in the following way: Where α and β are weighting coefficients used to balance the effects of motion and texture (for example, they can both be set to 0.5).

[0091] In some embodiments, determining the baseline QP of the current video frame based on the frame-level comprehensive importance score of the current video frame may include: If the current video frame is a B-frame or a P-frame, the number of bits for the current video frame is determined by allocating bits based on the frame-level comprehensive importance score of the current video frame. The baseline QP of the current video frame is determined based on the number of bits in the current video frame.

[0092] For example, for B-frames or P-frames, bits can be allocated to the video frame based on the frame-level comprehensive importance score, and the baseline QP of the video frame can be determined based on the number of bits in the video frame.

[0093] In one example, the bit allocation for the current video frame based on its frame-level comprehensive importance score can include: Bit allocation is performed on the current video frame based on its frame-level comprehensive importance score, the bitrate budget of the image group to which the current video frame belongs, and the bitrate surplus status of the image group to which the current video frame belongs.

[0094] For example, for any video frame, bit allocation can be performed based on the frame-level comprehensive importance score of the video frame, the bitrate budget of the group of pictures (GOP) to which the video frame belongs, and the bitrate balance status of the group of pictures to which the current video frame belongs.

[0095] As an example, the bit allocation for the current video frame based on its frame-level comprehensive importance score, the bitrate budget of the image group to which the current video frame belongs, and the bitrate surplus status of the image group to which the current video frame belongs can include: Based on the remaining bitrate budget of the image group to which the current video frame belongs and the number of remaining uncoded video frames, determine the average remaining bitrate of the image group to which the current video frame belongs; The first number of bits is determined based on the average remaining bitrate of the image group to which the current video frame belongs, and the frame-level comprehensive importance score of the current video frame; The second number of bits is determined based on the bitrate balance of the image group to which the current video frame belongs; The sum of the first number of bits and the second number of bits is used to determine the number of bits allocated to the current video frame.

[0096] For example, the number of bits allocated to a video frame may include two parts: one part is the number of bits determined based on the average remaining bitrate in the GOP to which the video frame belongs (which may be referred to as the first number of bits), and the other part is the number of bits determined based on the bitrate balance status of the GOP to which the video frame belongs (which may be referred to as the second number of bits).

[0097] For example, the bitrate balance can be used to characterize the instantaneous deviation between the current encoder output bitrate and the target bitrate, and can be used to measure past coding accuracy. For instance, the bitrate balance can be determined by comparing the difference between the ideal code buffer saturation and the current code buffer saturation. If the difference is positive, it means the actual buffer is lower than the ideal value, indicating that the bitrate allocated in the previous coding was too low; conversely, it indicates that the bitrate allocated in the previous coding was too high.

[0098] The remaining bitrate budget of the picture group to which the current video frame belongs can be used to make a static estimate of future available resources. For example, the remaining bitrate budget can be determined based on a fixed target bitrate, frame rate, and the number of encoded frames in the GOP.

[0099] For example, the second bit number can be a positive or negative value, used to adjust the number of bits allocated to the current video frame.

[0100] As an example, determining the first number of bits based on the average remaining bitrate of the image group to which the current video frame belongs, and the frame-level comprehensive importance score of the current video frame, may include: Based on the frame-level comprehensive importance score of the current video frame, the bit allocation adjustment parameters are determined; wherein, the bit allocation adjustment parameters are positively correlated with the frame-level comprehensive importance score of the current video frame; The first number of bits is determined by multiplying the bit allocation adjustment parameter by the average remaining bit rate of the image group to which the current video frame belongs.

[0101] As an example, the bitrate balance of the image group to which the current video frame belongs is determined based on the difference between the ideal coding buffer saturation and the current coding buffer saturation.

[0102] For example, the number of bits in the current video frame can be determined in the following way: in, Rtotal The remaining bitrate budget for the GOP to which the current video frame belongs. Nremain This represents the number of remaining uncoded frames in the GOP to which the current video frame belongs. The average remaining bitrate of the GOP to which the previous video frame belongs. F ( Sframe The bit allocation adjustment parameters are determined based on the frame-level comprehensive importance score of the current video frame. Btarget For ideal encoding buffer saturation, Bcurrent The current encoding buffer saturation. δ The feedback control strength coefficient (usually 0 < 0) δ <1).

[0103] For example, F ( Sframe The importance score can be determined using the following importance adjustment function based on the overall frame-level importance score of the current video frame: Where k is the sensitivity coefficient, For historical frames in the GOP group to which the current video frame belongs. Sframe The average value.

[0104] For example, F ( Sframe This can be used to ensure that frames of high importance receive a higher-than-average number of bits.

[0105] For example, This is a key indicator of the remaining bitrate, including: If the buffer is almost full ( Bcurrent > BtargetThis option is negative, reducing the number of bits allocated to the current video frame to prevent overflow.

[0106] If the buffer zone is empty ( Bcurrent < Btarget This value is positive, increasing the number of bits allocated to the current frame to avoid underflow and fully utilize bandwidth.

[0107] In some embodiments, determining the baseline QP of the current video frame based on the frame-level comprehensive importance score of the current video frame may include: If the current video frame is an I-frame and is not the first I-frame, the baseline QP of the current video frame is determined based on the frame-level comprehensive importance score and the QP of historical video frames.

[0108] For example, the QP of the first I-frame is determined based on the frame-level comprehensive importance score and encoder context information, or based on configuration instructions.

[0109] For example, the QP of the first I-frame in a video can be determined based on the frame-level comprehensive importance score and encoder context information, or based on configuration instructions (i.e., it can be configured by the user according to actual needs).

[0110] For example, encoder context information may include, but is not limited to, some or all of the information such as I-frame interval, video encoding bitrate, frame rate, and resolution.

[0111] For example, for a non-first I-frame in a video, its QP can be determined based on the frame-level comprehensive importance score and the QP of historical video frames (such as historical video frames in the same image group).

[0112] For example, the baseline QP of the current video frame can be determined based on the frame-level composite importance score of the current video frame (I-frame), the QP of the previous B-frame or P-frame, and the QP of the previous I-frame.

[0113] For example, the baseline QP of the current video frame (I-frame) can be determined based on the QP of historical video frames, and the baseline QP of the current video frame (I-frame) can be adjusted based on the frame-level comprehensive importance score of the current video frame (I-frame) to determine the final baseline QP of the current video frame (I-frame).

[0114] For example, the baseline QP of the current video frame (I-frame) can be the average of the QP of the previous P-frame and the QP of the previous I-frame.

[0115] For example, the higher the frame-level composite importance score of the current video frame (I-frame), the smaller the adjusted baseline QP.

[0116] In some embodiments, before determining the baseline quantization parameter QP of the current video frame based on the frame-level comprehensive importance score of the current video frame, the following may also be included: If the absolute value of the difference between the brightness of the current video frame and the brightness of the previous video frame exceeds a preset brightness threshold, or if the difference between the motion vector amplitude of the current video frame and the motion vector amplitude of the previous video frame exceeds a preset amplitude threshold, the current video frame is determined to be an I-frame.

[0117] For example, considering that there may be scene changes or a sharp increase in motion complexity in the video, if I-frames are still encoded at fixed intervals in this case, the overall image quality may be degraded. Therefore, in order to improve the overall image quality, I-frames can be dynamically inserted when scene changes or a sharp increase in motion complexity are determined based on scene perception information.

[0118] Accordingly, if the absolute value of the difference between the brightness of the current video frame and the brightness of the previous video frame exceeds a preset brightness threshold (which can include situations where the brightness of the current video frame is significantly greater than that of the previous video frame, or situations where the brightness of the current video frame is significantly less than that of the previous video frame, for example, when the video image changes from bright to dark, or from dark to bright), or if the difference between the motion vector amplitude of the current video frame and the motion vector amplitude of the previous video frame exceeds a preset amplitude threshold (i.e., the motion vector amplitude of the current video frame is significantly greater than that of the previous video frame, such as when the motion complexity increases significantly), it can be determined that a scene change or a sharp increase in motion complexity has occurred, and an I-frame can be dynamically inserted.

[0119] For example, the brightness of a video frame can be the average brightness of the video frame.

[0120] For example, the motion vector magnitude of a video frame can be the ratio of the number of moving pixels in the video frame to the total number of pixels in the video frame.

[0121] For example, for static or simple scenarios, high-quality I-frames can be encoded (reducing the QP of the I-frame) and the I-frame interval can be extended.

[0122] In some embodiments, determining the QP offset of each image block in the current video frame based on the signal statistical characteristics of each image block in the current video frame may include: For any image patch, determine the motion intensity level of the image patch based on its SAD; and / or, determine the texture complexity level of the image patch based on its texture complexity information. The QP offset of the image patch is determined based on the motion intensity level of the image patch and / or the texture complexity level of the image patch.

[0123] For example, if the baseline QP of the current video frame is determined, the QP of each image block in the current video frame can be determined based on the baseline QP of the current video frame.

[0124] For example, the signal statistical characteristics of an image patch may include the motion intensity level of the image patch, and / or the texture complexity of the image patch.

[0125] For any image patch, the QP offset of the image patch can be determined based on the motion intensity level of the image patch and / or the texture complexity level of the image patch, and the baseline QP can be adjusted based on the QP offset of the image patch to determine the QP of the image patch.

[0126] For example, the motion intensity level and / or texture complexity of an image patch can be determined by the encoding processing unit during the encoding process of the image patch.

[0127] In one example, determining the motion intensity level of an image patch based on its SAD (Sensitive Aspect Ratio) can include: The motion intensity level of the image patch is determined based on the SAD of the image patch and the average SAD of the previous video frame.

[0128] For example, for any image block in the current video frame, the motion intensity level of the image block can be determined based on the SAD of the image block and the average SAD of the video frame.

[0129] For example, the motion intensity level of the image patch is positively correlated with the SAD (Sum of Absolute Differences) ratio, which is the ratio of the SAD of the image patch to the average SAD of the previous video frame.

[0130] In one example, for any image patch, the texture complexity information of that image patch is based on the texture complexity representation parameters of that image patch.

[0131] For example, texture complexity characterization parameters include at least one of MAD (Mean Absolute Difference), SATD (Sum of Absolute Transformed Differences), or MR_SATD (Mean of Residual – Sum of Absolute Transformed Difference).

[0132] For example, the texture complexity level of an image patch is positively correlated with the texture complexity representation parameter of the image patch.

[0133] For example, SATD is a metric used to measure the texture complexity of an image or video patch. The calculation of SATD for any image patch can include: transforming the image patch (such as Hadamard transform or DCT transform); and calculating the sum of the absolute differences between the transformed coefficients and the reference coefficients.

[0134] For example, for any image block in the current video frame, the QP of that image block can be obtained by offsetting the QP from the reference QP of the current video frame.

[0135] For example, the QP of the image block can be obtained by adding the QP offset of the image block to the base QP of the current video frame.

[0136] In some embodiments, determining the QP offset of each image block in the current video frame based on the signal statistical characteristics of each image block in the current video frame may include: Based on the semantic importance score of the target corresponding to the image patch, the basic QP offset of the image patch is determined; Based on the signal statistical characteristics of the image block and the basic QP offset of the image block, the QP offset of the image block is determined.

[0137] For example, the QP offset of an image patch can be determined based on the semantic importance score of the target corresponding to the image patch, as well as the signal statistical characteristics of the image patch.

[0138] For example, the signal statistical characteristics of an image patch may include the motion intensity level of the image patch, and / or the texture complexity of the image patch.

[0139] For example, the basic QP offset of an image patch can be determined based on the semantic importance score of the target corresponding to the image patch, and the basic QP offset of the image patch can be adjusted based on the motion intensity level and / or texture complexity level of the image patch to determine the final QP offset of the image patch.

[0140] In one example, a complexity amplification factor for adjusting the base QP offset can be determined based on the motion intensity level and / or texture complexity level of the image patch. Then, the QP offset of the image patch can be determined based on the base QP offset of the image patch and the complexity amplification factor of the image patch.

[0141] For example, the QP offset of an image patch can be determined by multiplying the base QP offset of the image patch by the complexity amplification factor of the image patch.

[0142] For example, the complexity scaling factor of an image patch can be determined in the following way: Among them, W mv and W mad These are the weighted weights for the motion intensity level and texture complexity level of the image patch, respectively, and W mv +W mad =1. For example, W mv and W mad The default value can be set to 0.5, indicating that the motion intensity level and texture complexity level of the image patch are equally important.

[0143] COMPLEXITY_RANGE is the total range of complexity.

[0144] For example, assuming that the values ​​of mv_level and mad_level are both from 1 to 4, then the values ​​of (mv_level-1) and (mad_level-1) are both from 0 to 3. The theoretical maximum value of ((mv_level-1)×Wmv+(mad_level-1)×Wmad) is 3. Therefore, COMPLEXITY_RANGE=3.

[0145] K_strength is the strength coefficient of the influence of complexity.

[0146] For example, K_strength can be set to 0.5, and the complexity scaling factor can be in the range of [1, 1.5], to prevent the complexity scaling factor from exceeding the dominance of the importance scaling factor.

[0147] For example, for any image block in the current video frame, given that the QP offset of the image block has been determined, the QP of the image block can be determined based on the reference QP of the current video frame and the QP offset of the image block.

[0148] For example, the QP of the current video frame can be determined by the sum of the offset of the reference QP of the current video frame and the QP of the image block.

[0149] To enable those skilled in the art to better understand the technical solutions provided in the embodiments of this application, the technical solutions provided in the embodiments of this application are described below with reference to specific examples.

[0150] In this embodiment, a multi-level rate-distortion optimization coding scheme is proposed. By constructing a multi-level rate-distortion optimization decision system from the "semantic layer" to the "signal layer", advanced machine vision task analysis is deeply coupled with low-level coding parameter control, thereby achieving fine-grained and adaptive allocation of bitrate resources.

[0151] like Figure 3 As shown, in this embodiment, the implementation process of the multi-layer rate-distortion optimized coding scheme may include the following steps: Step S300: Perform real-time target detection and tracking on the input video frame, construct a target importance evaluation model based on the multi-dimensional feature information of target detection, and determine the semantic importance score of each target in the input video frame.

[0152] For example, multi-dimensional feature information may include foreground / background, motion state, target category, and relative area, etc.

[0153] Step S310: Based on the semantic importance scores of each target in the input video frame and combined with scene-aware feature information, determine the baseline QP of the input video frame.

[0154] For example, scene-aware feature information may include information such as global motion percentage and brightness (e.g., average screen brightness).

[0155] Step S320: Based on the baseline QP of the input video frame and combined with the signal statistical characteristics of each image block in the input video frame, determine the QP of each image block in the input video frame.

[0156] For example, signal statistical features may include local spatial features (such as texture complexity) and temporal features (such as motion vectors) of image patches.

[0157] As can be seen, a top-down, multi-layered bitrate decision-making system is constructed through target semantic layer analysis, scene perception layer processing, and signal statistics layer regulation, deeply coordinating semantic understanding, scene analysis, and signal coding. The semantic importance extracted from high-level visual tasks (such as target recognition) is integrated with mid-level scene dynamic features (such as motion and texture complexity) to accurately map and guide the optimization of low-level coding parameters (such as block-level QP), thereby achieving adaptive bitrate allocation for visual perception (i.e., rate-distortion optimization) under given bitrate constraints.

[0158] The following section, with reference to the accompanying diagram, explains some implementation details of the above implementation process.

[0159] I. Target semantic layer analysis.

[0160] For example, the target semantic layer aims to build a comprehensive video surveillance target importance assessment model. The target semantic recognition module inputs target feature information, and the target importance assessment model will convert the target's multi-dimensional features into an importance score (semantic importance score) through a weighted summation method.

[0161] For example, the target semantic importance score (S) can be decomposed into a weighted sum of four core dimensions: Target motion state (M): distinguishes whether the target is moving or stationary.

[0162] Relationship between target and context (B): Distinguish between target and context.

[0163] Target semantic category (C): Different weights are assigned based on the target category (such as pedestrian, license plate).

[0164] Relative area of ​​target (A): Considers the size of the target in the image.

[0165] For example, the target semantic importance score S can be calculated using the following formula: S=W_m×S_M+W_b× S_B+W_c×S_C+W_a×S_A Where W_m, W_b, W_c, and W_a are the weight coefficients for each dimension, which can be adjusted according to the actual application scenario. S_M, S_B, S_C, and S_A are the scores for each dimension.

[0166] For example, the target semantic importance score can be determined in the following way: 1.1 Definition of scores for each dimension.

[0167] 1.1.1 Motion status score (S_M).

[0168] This dimension is used to measure the dynamic characteristics of the target. For example: Foreground of motion: S_M=1.0; Static foreground: S_M=0.7 (Static but with anomalies, still of high interest); Background of the movement: S_M=0.3 (such as swaying trees, low attention level); Static background: S_M=0.0 (basically not of concern).

[0169] 1.1.2 Background Relationship Score (S_B).

[0170] This dimension further refines the hierarchy of the target within the scene.

[0171] Foreground: S_B=1.0; Background: S_B=0.0.

[0172] 1.1.3 Semantic category score (S_C).

[0173] Based on the social significance of the target and the purpose of monitoring, this dimension divides common prospective targets into three levels of concern: high, medium, and low.

[0174] High-concern categories (S_C=1.0): For example, pedestrians, cars, and license plates. These are core targets for security and traffic monitoring.

[0175] Category of interest (S_C=0.6): For example, buses, trucks, motorcycles, bicycles, tricycles. These have transportation significance, but are of slightly lower importance.

[0176] Low-concern categories (S_C=0.2): For example, animals (cats, dogs), unidentified objects / packages. These have some significance, but are usually not primary monitoring targets.

[0177] Background category (S_C=0.0): Sky, ground, trees, buildings. These are typically considered part of the environment and their importance is not calculated separately.

[0178] 1.1.4 Relative Area Score (S_A).

[0179] This dimension addresses the impact of target size on attention.

[0180] For example, a target area that is too large (such as a pedestrian nearby) or too small (such as a vehicle in the distance) will reduce attention.

[0181] For example, a Gaussian function can be used to simulate the pattern that "medium-sized targets receive the most attention": Where r is the proportion of the target bounding box area to the total screen area (i.e., the actual area proportion), where 0 < r ≤ 1; μ is the proportion of the most attention area, for example, μ = 0.1; σ is used to control the width of the attention interval, for example, σ is set to 0.05.

[0182] 1.2 Definition of semantic importance score S.

[0183] For example, the semantic importance score S can be determined based on the scores of each dimension as follows: S=W_m×S_M+W_b×S_B+W_c×S_C+W_a×S_A For example, the weighting coefficients of the importance scores of each dimension can be determined according to the application scenario. In different application scenarios, the weighting coefficients of the importance scores of each dimension of the target are allowed to be different.

[0184] For example, in traffic monitoring scenarios, the weighting coefficient of the importance score of the semantic category dimension can be increased; in perimeter security scenarios, the weighting coefficient of the importance score of the motion state dimension can be increased.

[0185] Examples of semantic importance scores for different objectives can be shown in Table 1: Table 1 For example, the target importance assessment model provides a quantifiable framework that can effectively convert human subjective attention into machine-calcifiable scores, thereby improving the intelligence level of video surveillance systems.

[0186] II. Scene Awareness Layer Processing.

[0187] For example, the task of the scene perception layer is to receive target semantic layer information and input information (scene perception information) from the scene perception module.

[0188] For example, the scene perception information obtained by the scene perception module may include, but is not limited to, global motion percentage, brightness, and motion vector amplitude information.

[0189] The target bit count Tframe and frame-level baseline quantization parameter QP_base for the current frame can be determined by combining the current encoder system status (bitrate balance status).

[0190] like Figure 4 As shown, the scene perception layer processing flow may include the following steps: Step S400: Dynamically insert I-frames based on scene perception information.

[0191] For example, when scene switching or a sharp increase in motion complexity is determined based on scene perception information, I-frames are dynamically inserted.

[0192] For example, if the absolute value of the difference between the brightness of the current video frame and the brightness of the previous video frame exceeds a preset brightness threshold, or if the difference between the motion vector amplitude of the current video frame and the motion vector amplitude of the previous video frame exceeds a preset amplitude threshold, the current video frame is determined to be an I-frame.

[0193] For example, for static or simple scenarios, high-quality I-frames can be encoded (reducing the QP of the I-frame) and the I-frame interval can be extended.

[0194] S410, Calculate the frame-level comprehensive importance score. Sframe .

[0195] For example, frame-level comprehensive importance score Sframe Semantic importance and scenario complexity can be combined as the main basis for bit allocation.

[0196] 2.1 Semantic Importance Components Ssemantic (Also known as frame-level semantic importance score).

[0197] Input: Importance scores S of all targets from the target semantic layer i and its corresponding bounding box area a i .

[0198] Calculation: Calculate the weighted average semantic importance of the entire frame to avoid a single small target dominating the overall frame decision.

[0199] For example, the frame-level semantic importance score of a video frame can be determined in the following way: Where N is the number of targets detected in the current video frame.

[0200] It should be noted that when N=0, that is, when no target is detected in the current video frame, This is the default value. For example, the default value can be a small value, such as 0.1.

[0201] 2.2 Scene Dynamics Components (Also known as scene dynamism score).

[0202] For example, scene dynamism scores can be determined based on the proportion of motion and texture complexity.

[0203] For example, motion percentage Rmotion: calculates the percentage of pixels whose motion vectors exceed a preset motion vector threshold between the current video frame and the previous video frame.

[0204] A high percentage of motion indicates significant scene changes, requiring a higher bitrate to maintain smoothness.

[0205] Texture complexity Ctexture: Calculates the gradient variance or DCT coefficient variance of video frames.

[0206] Frames with complex textures are more difficult to compress and require a higher bitrate.

[0207] Calculation: Normalize and combine the motion ratio Rmotion and texture complexity Ctexture: Here, α and β are weighting coefficients used to balance the effects of motion and texture.

[0208] 2.3 Frame-level Comprehensive Importance Score Sframe .

[0209] For example, the frame-level comprehensive importance score can be determined in the following way. Sframe : Here, γ is the fusion weight (0 < γ < 1), used to control the relative importance of semantic information and low-level statistical information. For example, in a monitoring scenario, a larger γ (such as 0.7) can be set to give more weight to semantics.

[0210] Step S420: Based on the frame-level comprehensive importance score Sframe Perform bit allocation.

[0211] For example, the target number of bits for video frame allocation can be determined in the following ways. Tframe : in, Rtotal The remaining bitrate budget for the GOP to which the current video frame belongs. Nremain This represents the number of remaining uncoded frames in the GOP to which the current video frame belongs. The average remaining bitrate of the GOP to which the previous video frame belongs. F ( Sframe The bit allocation adjustment parameters are determined based on the frame-level comprehensive importance score of the current video frame. Btarget For ideal encoding buffer saturation, Bcurrent The current encoding buffer saturation. δ The feedback control strength coefficient (usually 0 < 0) δ <1).

[0212] For example, F ( Sframe The importance score can be determined using the following importance adjustment function based on the overall frame-level importance score of the current video frame: Where k is the sensitivity coefficient, For historical frames in the GOP group to which the current video frame belongs. Sframe The average value.

[0213] For example, F ( Sframe This can be used to ensure that frames of high importance receive more bits than average.

[0214] For example, This is a key indicator of the remaining bitrate, including: If the buffer is almost full ( Bcurrent > Btarget This option is negative, reducing the number of bits allocated to the current video frame to prevent overflow.

[0215] If the buffer zone is empty ( Bcurrent < Btarget This value is positive, increasing the number of bits allocated to the current frame to avoid underflow and fully utilize bandwidth.

[0216] Step S430: Determine the baseline quantization parameter QP_base.

[0217] For example, taking a video frame as a B-frame or a P-frame, the reference quantization parameters of the video frame can be determined based on the target number of bits of the video frame.

[0218] For example, the relationship between the baseline quantization parameters of a video frame and the target bit count can be determined using the rate-distortion model R-lamda. That is, the baseline quantization parameters of a video frame can be determined using the rate-distortion model R-lamda based on the target bit count of the video frame.

[0219] III. Signal Statistical Layer Control.

[0220] like Figure 5 As shown, the signal statistics layer control process may include the following steps: Step S500: For any image block in the current video frame, determine the motion intensity level mv_level of the image block based on the SAD of the image block and the average SAD of the previous video frame.

[0221] For example, the SAD of an image patch can be determined in the following way: Where f(x,y) is the original pixel value at (x,y), and g(x,y) is the reconstructed pixel value or predicted pixel value at (x,y).

[0222] Step S510: Determine the texture complexity level mad_level of the image block based on the texture complexity information of the image block.

[0223] For example, the texture complexity information of an image patch can be based on the MAD representation of the image patch.

[0224] For example, the MAD of an image patch can be determined in the following way: Step S520: Determine the QP of the image patch based on the motion intensity level, texture complexity level, and semantic importance score of the target corresponding to the image patch.

[0225] For example, the implementation process for determining the QP of an image patch based on its motion intensity level, texture complexity level, and the semantic importance score of the target corresponding to the image patch can be found in [reference needed]. Figure 6 .

[0226] like Figure 6 As shown, the process for determining the QP of an image patch may include: Step S521: Calculate the QP base offset (QP_BaseOffset) of the image patch.

[0227] For example, for any image block in the current video frame, the QP base offset of the image block can be determined based on the semantic importance score of the target corresponding to the image block.

[0228] It should be noted that when an image patch corresponds to multiple targets (i.e., the pixels of the image patch are included within the bounding boxes of the multiple targets), the semantic importance score of the target with the highest semantic importance score among the multiple targets can be determined as the semantic importance score of the target corresponding to the image patch.

[0229] For example, the QP base offset of an image patch can be determined in the following way: QP_BaseOffset=-MAX_OFFSET×S MAX_OFFSET is the maximum allowed QP reduction value. Its specific value can be set according to actual needs. For example, setting it to 10 means that when the semantic importance score S of the target corresponding to the image patch is 1, the QP base offset is -10.

[0230] Step S522: Calculate the complexity magnification factor of the image patch.

[0231] For example, for more complex image patches, the QP of the image patch usually needs to be set to a smaller value, that is, the absolute value of the QP offset of the image patch is larger (the QP offset is negative, the larger the absolute value, the smaller the value of the QP offset), so as to allocate a higher bitrate to the more complex image patch.

[0232] For example, the complexity scaling factor of an image patch can be determined in the following way: Among them, W mv and W mad These are the weighted weights for the motion intensity level and texture complexity level of the image patch, respectively, and W mv +W mad =1. For example, W mv and W mad The default value can be set to 0.5, indicating that the motion intensity level and texture complexity level of the image patch are equally important.

[0233] COMPLEXITY_RANGE is the total range of complexity.

[0234] For example, assuming that the values ​​of mv_level and mad_level are both from 1 to 4, then the values ​​of (mv_level-1) and (mad_level-1) are both from 0 to 3. The theoretical maximum value of ((mv_level-1)×Wmv+(mad_level-1)×Wmad) is 3. Therefore, COMPLEXITY_RANGE=3.

[0235] K_strength is the strength coefficient of the influence of complexity.

[0236] For example, K_strength can be set to 0.5, and the complexity scaling factor can be in the range of [1, 1.5], to prevent the complexity scaling factor from exceeding the dominance of the importance scaling factor.

[0237] Step S523: Calculate the final QP offset of the image block, and determine the QP of the image block based on the final QP offset of the image block.

[0238] For example, the final QP offset of an image patch can be determined based on the QP baseline offset of the image patch and the complexity amplification factor.

[0239] For example, the final QP offset of an image patch can be determined in the following way: QP_Offset = QP_BaseOffset × K complexity That is, the final QP offset of the image block is the product of the QP baseline offset of the image block and the complexity amplification factor.

[0240] For example, the QP of an image block can be determined based on the baseline QP of the current video frame and the final QP offset of the image block.

[0241] For example, the QP of an image patch can be determined in the following way: QP_final = QP_base + QP_Offset That is, the QP of an image block is the sum of the reference QP of the current video frame and the final QP offset of the image block.

[0242] As can be seen, this embodiment proposes a closed-loop decision-making process that spans the semantic layer, scene layer, and signal layer. First, at the semantic layer, a machine vision model is used to analyze video content in real time, generating a bitrate allocation priority based on target importance (e.g., foreground > background). This priority information is then passed down to the scene layer, where it is fused and weighted with real-time calculated dynamic features (e.g., motion intensity, texture details, and other scene-aware information) to form a more refined bitrate control strategy that adapts to scene changes. Finally, this strategy, combined with block-level statistical information, is derived at the signal layer as a block-level quantization parameter (QP), enabling precise quantization control of regions of different importance. This achieves refined and adaptive optimization of bitrate resource allocation, realizing multi-layer adaptive rate-distortion optimization.

[0243] The technical solutions provided in the embodiments of this application will be described by way of example below in conjunction with specific application scenarios.

[0244] Scenario 1: Application in a monitoring scenario.

[0245] In surveillance scenarios, the front-end surveillance equipment can apply the multi-level rate-distortion optimized encoding scheme provided in the embodiments of this application to reduce the bitstream transmission bandwidth and save storage space while ensuring video quality.

[0246] like Figure 7 As shown, in a monitoring scenario, after the image passes through the lens, image sensor, and image processing unit, the data stream is split into two paths and sent to the video encoding processing unit and the intelligent processing unit respectively. The intelligent processing unit performs semantic recognition on the input image, extracts the target semantic information in the input image, and inputs the extracted target semantic information into the video encoding processing unit. The video encoding processing unit can use the multi-level rate-distortion optimized encoding scheme provided in the embodiments of this application to perform video encoding, thereby reducing the video encoding bitrate, reducing transmission bandwidth, and saving storage while ensuring the subjective effect of the image.

[0247] like Figure 7 As shown, the image processing unit and the background modeling unit, as scene perception modules, acquire scene perception information, including the proportion of motion information and changes in screen brightness, and input it into the video encoding unit.

[0248] The video encoding processing unit can combine the scene perception information input by the scene perception module and the target semantic information input by the target semantic recognition module, and use the multi-level rate-distortion optimized encoding scheme provided in the embodiments of this application to perform video encoding to obtain a video encoded bitstream.

[0249] For example, video encoded streams can be used for local file storage and network transmission.

[0250] Scenario 2: Application in backend storage business scenarios.

[0251] In back-end storage scenarios, storage devices (such as CVRs or NVRs) can apply the multi-level rate-distortion optimized encoding scheme provided in the embodiments of this application to save storage space while ensuring video quality.

[0252] like Figure 8 As shown, in backend storage scenarios, storage devices CVR / NVR receive bitstreams from frontend devices via the network.

[0253] For example, the bitstream of the front-end device can be a bitstream obtained by encoding using a traditional scheme, or a bitstream obtained by encoding using a multi-level rate-distortion optimized encoding scheme provided in the embodiments of this application.

[0254] The storage device can decode the received bitstream and perform secondary encoding using the multi-level rate-distortion optimized encoding scheme provided in this application embodiment, thereby reducing the bitstream size and saving storage while ensuring video quality.

[0255] like Figure 8 As shown, the decoding processing unit can be used to decode the received bitstream and input the obtained video image data (such as video YUV data stream) into the video encoding processing unit and the intelligent processing unit respectively; on the other hand, it can be used as a scene perception module to acquire scene perception information and input it into the video encoding processing unit.

[0256] The intelligent processing unit can serve as a target semantic recognition module, performing semantic recognition on the input video image data, extracting target semantic information from the video image data, and inputting the extracted target semantic information into the video encoding processing unit.

[0257] The video encoding processing unit can combine the scene perception information input by the scene perception module and the target semantic information input by the target semantic recognition module, and use the multi-level rate-distortion optimized encoding scheme provided in this application embodiment to encode the input video image data to obtain the video encoded bitstream, and input it into the storage unit for storage.

[0258] The method provided in this application has been described above. The apparatus provided in this application is described below: Please see Figure 9 This is a schematic diagram of the structure of an image encoding device provided in an embodiment of this application, as shown below. Figure 9 As shown, the image encoding device may include: The first determining unit is used to perform semantic recognition on the current video frame and determine the semantic importance score of each target in the current video frame; The second determining unit is used to determine the frame-level semantic importance score of the current video frame based on the semantic importance scores of each target in the current video frame, and to determine the frame-level scene dynamics score of the current video frame based on the scene perception information of the current video frame. The second determining unit is further configured to determine the frame-level comprehensive importance score of the current video frame based on the frame-level semantic importance score of the current video frame and the frame-level scene dynamics score of the current video frame. The third determining unit is used to determine the baseline QP of the current video frame based on the frame-level comprehensive importance score of the current video frame. The fourth determining unit is used to determine the QP offset of each image block in the current video frame based on the signal statistical characteristics of each image block in the current video frame, and to determine the QP of each image block in the current video frame based on the reference QP of the current video frame and the QP offset of each image block in the current video frame. The encoding unit is used to encode the current video frame based on the QP of each image block of the current video frame.

[0259] For example, the specific implementation process of the image encoding scheme by each functional unit in the image encoding device can be found in the relevant descriptions in the above embodiments.

[0260] This application provides an electronic device including a processor and a memory, wherein the memory stores machine-executable instructions that can be executed by the processor, and the processor executes the machine-executable instructions to implement the image encoding method described above.

[0261] Please see Figure 10 This is a schematic diagram of the hardware structure of an electronic device provided in an embodiment of this application. The electronic device may include a processor 1001 and a memory 1002 storing machine-executable instructions. The processor 1001 and the memory 1002 can communicate via a system bus 1003. Furthermore, by reading and executing the machine-executable instructions corresponding to the image encoding logic in the memory 1002, the processor 1001 can execute the image encoding method described above.

[0262] The memory 1002 mentioned in this document can be any electronic, magnetic, optical, or other physical storage device that can contain or store information, such as executable instructions, data, etc. For example, machine-readable storage media can be: RAM (Random Access Memory), volatile memory, non-volatile memory, flash memory, storage drives (such as hard disk drives), solid-state drives, any type of storage disk (such as optical discs, DVDs, etc.), or similar storage media, or combinations thereof.

[0263] In some embodiments, a machine-readable storage medium, such as Figure 10 The memory 1002 in the memory, which is a machine-readable storage medium, stores machine-executable instructions that, when executed by a processor, implement the image encoding method described above. For example, the storage medium may be ROM, RAM, CD-ROM, magnetic tape, floppy disk, or optical data storage device.

[0264] This application embodiment also provides a monitoring front-end device, including: a lens, an image sensor, and a processor, wherein the processor includes a scene perception module, a target semantic recognition module, and a signal analysis module; wherein: The lens is used to receive incident light; The image sensor is used to receive an optical image focused by the lens and convert the optical image into a raw image electrical signal; The scene perception module includes an image processing unit, which processes the original image electrical signal to generate a digital image signal, and transmits the digital image signal to the target semantic recognition module and the signal analysis module respectively. The scene perception module further includes a background modeling unit. The image processing unit and the background modeling unit are used to acquire scene perception information and transmit it to the signal analysis module. The target semantic recognition module includes an intelligent processing unit, which is used to perform semantic recognition on the digital image signal, extract target semantic information, and transmit the target semantic information to the signal analysis module; The signal analysis module includes a video encoding processing unit, which performs video encoding using the image encoding method described above to obtain a video encoded bitstream; wherein the video encoded bitstream is used for local file storage and / or network transmission.

[0265] This application embodiment also provides a storage device, including: a processor and a storage unit, wherein the processor includes a scene perception module, a target semantic recognition module, and a signal analysis module; wherein: The scene perception module includes a decoding processing unit, which is used to decode the received video stream, transmit the obtained video image data to the signal analysis module and the target semantic recognition module respectively, and acquire scene perception information and transmit the scene perception information to the signal analysis module. The target semantic recognition module includes an intelligent processing unit, which is used to perform semantic recognition on the video image data, extract target semantic information, and transmit the target semantic information to the signal analysis module; The signal analysis module includes a video encoding processing unit, which performs video encoding by executing the image encoding method described above to obtain a video encoded bitstream; The storage unit is used to store the video encoded bitstream.

Claims

1. An image encoding method, characterized in that, include: Perform semantic recognition on the current video frame to determine the semantic importance score of each target in the current video frame; Based on the semantic importance scores of each target in the current video frame, the frame-level semantic importance score of the current video frame is determined, and based on the scene perception information of the current video frame, the frame-level scene dynamics score of the current video frame is determined. Based on the frame-level semantic importance score of the current video frame and the frame-level scene dynamism score of the current video frame, the frame-level comprehensive importance score of the current video frame is determined. Based on the frame-level comprehensive importance score of the current video frame, the baseline quantization parameter QP of the current video frame is determined; Based on the signal statistical characteristics of each image block in the current video frame, the QP offset of each image block in the current video frame is determined, and based on the reference QP of the current video frame and the QP offset of each image block in the current video frame, the QP of each image block in the current video frame is determined. The current video frame is encoded based on the QP of each image block in the current video frame; The scene perception information includes global motion percentage and texture complexity; The step of determining the frame-level scene dynamics score of the current video frame based on the scene perception information of the current video frame includes: The proportion of pixels in the current video frame whose motion vectors relative to the previous video frame exceed a preset motion vector threshold is determined as the global motion proportion of the current video frame; and, Determine the texture complexity of the current video frame; The weighted sum of the global motion percentage and texture complexity of the current video frame is determined as the frame-level scene dynamics score of the current video frame. The step of determining the QP offset of each image block in the current video frame based on the signal statistical characteristics of each image block in the current video frame includes: Based on the semantic importance score of the target corresponding to the image patch, the basic QP offset of the image patch is determined; Based on the signal statistical characteristics of the image block and the basic QP offset of the image block, the QP offset of the image block is determined.

2. The method according to claim 1, characterized in that, The step of performing semantic recognition on the current video frame and determining the semantic importance score of each target in the current video frame includes: For any target in the current video frame, the semantic importance score of the target is determined based on the multi-dimensional features of the target; The multi-dimensional features include at least two of the following features: Foreground / background, motion state, semantic category, and relative area.

3. The method according to claim 2, characterized in that, The determination of the semantic importance score of the target based on its multi-dimensional features includes: Based on the multi-dimensional characteristics of the target, the importance score of each dimension of the target is determined respectively; The weighted sum of the importance scores of each dimension of the objective is used to determine the semantic importance score of the objective. The weighting coefficients for the importance scores of each dimension of the objective are determined based on the application scenario. Different application scenarios may allow for different weighting coefficients for the importance scores of each dimension of the objective.

4. The method according to claim 3, characterized in that, The determination of importance scores for each dimension of the target based on its multi-dimensional characteristics includes: When the multi-dimensional features include relative area, the relative area importance score of the target is determined based on the absolute value of the difference between the area ratio of the target bounding box area in the current video frame and the area ratio of the most concerned area; wherein the relative area importance score of the target is negatively correlated with the absolute value.

5. The method according to claim 1, characterized in that, Determining the frame-level semantic importance score of the current video frame based on the semantic importance scores of each target in the current video frame includes: The frame-level semantic importance score of the current video frame is determined based on the semantic importance score of each target in the current video frame and the area of ​​the target bounding box of each target. The step of determining the overall frame importance score of the current video frame based on the frame-level semantic importance score and the frame-level scene dynamism score of the current video frame includes: Based on the set fusion weights, the frame-level semantic importance score and the frame-level scene dynamism score of the current video frame are fused, and the resulting fusion score is determined as the frame-level comprehensive importance score of the current video frame.

6. The method according to claim 5, characterized in that, The step of determining the frame-level semantic importance score of the current video frame based on the semantic importance score of each target in the current video frame and the area of ​​the target bounding box of each target includes: For any target in the current video frame, the weighting coefficient of the semantic importance score of the target is determined based on the ratio of the area of ​​the target's bounding box to the sum of the areas of the bounding boxes of all targets in the current video frame. The weighted sum of the semantic importance scores of each target in the current video frame is determined as the frame-level semantic importance score of the current video frame.

7. The method according to claim 1, characterized in that, The step of determining the baseline quantization parameter QP of the current video frame based on the frame-level comprehensive importance score of the current video frame includes: If the current video frame is a B-frame or a P-frame, the number of bits in the current video frame is determined by allocating bits to the current video frame based on the frame-level comprehensive importance score of the current video frame. The baseline QP of the current video frame is determined based on the number of bits in the current video frame.

8. The method according to claim 7, characterized in that, The step of allocating bits to the current video frame based on its frame-level comprehensive importance score includes: Bit allocation is performed on the current video frame based on its frame-level comprehensive importance score, the bitrate budget of the image group to which the current video frame belongs, and the bitrate surplus status of the image group to which the current video frame belongs.

9. The method according to claim 8, characterized in that, The step of allocating bits to the current video frame based on its frame-level comprehensive importance score, the bitrate budget of the image group to which the current video frame belongs, and the bitrate surplus status of the image group to which the current video frame belongs includes: Based on the remaining bitrate budget of the image group to which the current video frame belongs and the number of remaining uncoded video frames, the average remaining bitrate of the image group to which the current video frame belongs is determined; The first number of bits is determined based on the average remaining bitrate of the image group to which the current video frame belongs, and the frame-level comprehensive importance score of the current video frame; The second number of bits is determined based on the bitrate balance status of the image group to which the current video frame belongs; The sum of the first number of bits and the second number of bits is determined as the number of bits allocated to the current video frame.

10. The method according to claim 9, characterized in that, The step of determining the first number of bits based on the average remaining bitrate of the image group to which the current video frame belongs, and the frame-level comprehensive importance score of the current video frame, includes: Based on the frame-level comprehensive importance score of the current video frame, a bit allocation adjustment parameter is determined; wherein the bit allocation adjustment parameter is positively correlated with the frame-level comprehensive importance score of the current video frame; The first number of bits is determined by multiplying the bit allocation adjustment parameter by the average remaining bit rate of the image group to which the current video frame belongs. And / or, The bitrate balance of the image group to which the current video frame belongs is determined based on the difference between the ideal coding buffer saturation and the current coding buffer saturation.

11. The method according to claim 1, characterized in that, The step of determining the baseline quantization parameter QP of the current video frame based on the frame-level comprehensive importance score of the current video frame includes: If the current video frame is an I-frame and is not the first I-frame, the baseline QP of the current video frame is determined based on the frame-level comprehensive importance score and the QP of historical video frames. The QP of the first I-frame is determined based on the frame-level comprehensive importance score and encoder context information, or based on configuration instructions.

12. The method according to claim 1, characterized in that, Before determining the baseline quantization parameter QP of the current video frame based on the frame-level comprehensive importance score of the current video frame, the method further includes: If the absolute value of the difference between the brightness of the current video frame and the brightness of the previous video frame exceeds a preset brightness threshold, or if the difference between the motion vector amplitude of the current video frame and the motion vector amplitude of the previous video frame exceeds a preset amplitude threshold, the current video frame is determined as an I-frame.

13. The method according to claim 1, characterized in that, The step of determining the QP offset of each image block in the current video frame based on the signal statistical characteristics of each image block in the current video frame includes: For any image patch, determine the motion intensity level of the image patch based on its absolute error and SAD; and / or, determine the texture complexity level of the image patch based on its texture complexity information. The QP offset of the image patch is determined based on the motion intensity level of the image patch and / or the texture complexity level of the image patch.

14. The method according to claim 13, characterized in that, The determination of the motion intensity level of the image patch based on its absolute error and SAD includes: Based on the SAD of the image patch and the average SAD of the previous video frame, the motion intensity level of the image patch is determined; The motion intensity level of the image block is positively correlated with the SAD ratio, which is the ratio of the SAD of the image block to the average SAD of the previous video frame. And / or, For any image patch, the texture complexity information of the image patch is determined based on the texture complexity representation parameters of the image patch; wherein, the texture complexity representation parameters include at least one of mean absolute difference (MAD), absolute transform difference (SATD), or absolute transform difference after removing the residual mean and MR_SATD, and the texture complexity level of the image patch is positively correlated with the texture complexity representation parameters.

15. An image encoding device, characterized in that, include: The first determining unit is used to perform semantic recognition on the current video frame and determine the semantic importance score of each target in the current video frame; The second determining unit is used to determine the frame-level semantic importance score of the current video frame based on the semantic importance scores of each target in the current video frame, and to determine the frame-level scene dynamics score of the current video frame based on the scene perception information of the current video frame. The second determining unit is further configured to determine the frame-level comprehensive importance score of the current video frame based on the frame-level semantic importance score of the current video frame and the frame-level scene dynamics score of the current video frame. The third determining unit is used to determine the baseline quantization parameter QP of the current video frame based on the frame-level comprehensive importance score of the current video frame. The fourth determining unit is used to determine the QP offset of each image block in the current video frame based on the signal statistical characteristics of each image block in the current video frame, and to determine the QP of each image block in the current video frame based on the reference QP of the current video frame and the QP offset of each image block in the current video frame. The encoding unit is used to encode the current video frame according to the QP of each image block of the current video frame; The scene perception information includes global motion percentage and texture complexity; The second determining unit determines the frame-level scene dynamics score of the current video frame based on the scene perception information of the current video frame, including: The proportion of pixels in the current video frame whose motion vectors relative to the previous video frame exceed a preset motion vector threshold is determined as the global motion proportion of the current video frame; and, Determine the texture complexity of the current video frame; The weighted sum of the global motion percentage and texture complexity of the current video frame is determined as the frame-level scene dynamics score of the current video frame. The fourth determining unit determines the QP offset of each image block in the current video frame based on the signal statistical characteristics of each image block in the current video frame, including: Based on the semantic importance score of the target corresponding to the image patch, the basic QP offset of the image patch is determined; Based on the signal statistical characteristics of the image block and the basic QP offset of the image block, the QP offset of the image block is determined.

16. An electronic device, characterized in that, The method includes a processor and a memory, the memory storing machine-executable instructions that can be executed by the processor, the processor executing the machine-executable instructions to implement the method as described in any one of claims 1-14.

17. A machine-readable storage medium, characterized in that, The machine-readable storage medium stores machine-executable instructions, which, when executed by a processor, implement the method as described in any one of claims 1-14.

18. A monitoring front-end device, characterized in that, include: The system comprises a lens, an image sensor, and a processor, wherein the processor includes a scene perception module, a target semantic recognition module, and a signal analysis module; wherein: The lens is used to receive incident light; The image sensor is used to receive an optical image focused by the lens and convert the optical image into a raw image electrical signal; The scene perception module includes an image processing unit, which processes the original image electrical signal to generate a digital image signal, and transmits the digital image signal to the target semantic recognition module and the signal analysis module respectively. The scene perception module further includes a background modeling unit. The image processing unit and the background modeling unit are used to acquire scene perception information and transmit it to the signal analysis module. The target semantic recognition module includes an intelligent processing unit, which is used to perform semantic recognition on the digital image signal, extract target semantic information, and transmit the target semantic information to the signal analysis module; The signal analysis module includes a video encoding processing unit, used to perform video encoding by executing the method described in any one of claims 1-14 to obtain a video encoded bitstream; wherein the video encoded bitstream is used for local file storage and / or network transmission.

19. A storage device, characterized in that, include: The processor and storage unit, wherein the processor includes a scene perception module, a target semantic recognition module, and a signal analysis module; wherein: The scene perception module includes a decoding processing unit, which is used to decode the received video stream, transmit the obtained video image data to the signal analysis module and the target semantic recognition module respectively, and acquire scene perception information and transmit the scene perception information to the signal analysis module. The target semantic recognition module includes an intelligent processing unit, which is used to perform semantic recognition on the video image data, extract target semantic information, and transmit the target semantic information to the signal analysis module; The signal analysis module includes a video encoding processing unit, used to perform video encoding by executing the method described in any one of claims 1-14 to obtain a video encoded bitstream; The storage unit is used to store the video encoded bitstream.