Video processing method and video processing device

By selecting an appropriate bitstream to synthesize video footage, the problem of poor playback quality was solved, achieving the effect of saving device decoding performance while ensuring picture quality.

CN116193048BActive Publication Date: 2026-07-24HANGZHOU HIKVISION DIGITAL TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
HANGZHOU HIKVISION DIGITAL TECHNOLOGY CO LTD
Filing Date
2022-12-29
Publication Date
2026-07-24

AI Technical Summary

Technical Problem

The actual playback effect of synthesized video images in existing technologies is not good, and it cannot balance saving the decoding performance of playback devices with ensuring the quality of image display.

Method used

By acquiring the data stream configuration information of the video channel and the encoding resolution and sub-picture configuration information of the image to be synthesized, the target bitstream is selected to synthesize the image to be synthesized, so that it meets the preset display clarity requirements and the resolution size of the target bitstream meets the preset resolution requirements. The target bitstream is either the main bitstream or the sub-bitstream.

Benefits of technology

While ensuring the clarity of the synthesized video, the sub-streams are used reasonably to reduce the decoding performance requirements of the playback device and improve the playback effect.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116193048B_ABST
    Figure CN116193048B_ABST
Patent Text Reader

Abstract

Embodiments of the present application disclose a video processing method and a video processing device. The video processing method comprises: obtaining data stream configuration information of a video channel, the data stream configuration information comprising a main code stream resolution and a resolution supported by a sub code stream, the main code stream resolution being higher than the resolution supported by the sub code stream; obtaining an encoding resolution of a to-be-combined picture and sub picture configuration information, the sub picture configuration information representing a size of a picture corresponding to the video channel in the to-be-combined picture; and selecting a target code stream to combine the to-be-combined picture according to the data stream configuration information of the video channel, the encoding resolution of the to-be-combined picture and the sub picture configuration information, so that the picture corresponding to the video channel in the to-be-combined picture meets a preset display definition requirement and a resolution of the target code stream meets a preset resolution requirement. The video processing method solves the problem of poor actual playing effect of a combined video picture in the related art.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of video processing technology, and more specifically, to a video processing method and a video processing apparatus. Background Technology

[0002] Currently, mainstream high-definition network cameras typically employ dual-stream technology. The camera's encoder has two encoding formats: a main stream and a sub-stream. The video stream encoded with the main stream boasts high resolution, high image quality, but consumes significant network bandwidth. Conversely, the video stream encoded with the sub-stream has relatively lower resolution and image quality, but consumes less network bandwidth during transmission. By applying dual-stream technology, the high-bitrate stream can be used for local high-definition video storage and preview, ensuring video quality, while the low-bitrate stream is used for network transmission (e.g., remote preview), thereby saving network bandwidth and improving video transmission speed and playback smoothness.

[0003] In some applications, it is necessary to display video from one or more video channels together. This requires decoding the streams from multiple video channels and encoding them according to a specific segmentation method to synthesize a single stream for recording or burning. For example, combining streams from two cameras into a single frame splits the image into two parts, each displaying an image captured by one of the two cameras.

[0004] In related technologies, when performing video image compositing, the main bitstream is selected for image compositing for any data stream from any video channel. This causes the display device to consume more performance for decoding when displaying the composite image, and in extreme cases, decoding may even fail due to insufficient decoding performance of the display device itself. Moreover, since the image from each video channel in the composite image is scaled down, the actual display effect of the image composited using the main bitstream may not be good.

[0005] Therefore, the scheme of using the main bitstream for image synthesis in related technologies cannot simultaneously achieve the goals of saving decoding performance of playback devices and ensuring image display quality, resulting in poor actual playback effect of the synthesized video images.

[0006] The information disclosed in the background section is only intended to enhance the understanding of the background art described herein. Therefore, the background art may contain information that would not be considered part of the prior art by those skilled in the art. Summary of the Invention

[0007] This invention provides a video processing method and a video processing apparatus to at least solve the problem of poor actual playback effect of synthesized video images in related technologies.

[0008] According to one aspect of the present invention, a video processing method is provided, comprising: acquiring data stream configuration information of a video channel, the data stream configuration information including a main bitstream resolution and a resolution supported by a sub-bitstream, wherein the main bitstream resolution is higher than the resolution supported by the sub-bitstream; acquiring the encoding resolution and sub-picture configuration information of a picture to be synthesized, wherein the sub-picture configuration information characterizes the size of the picture corresponding to the video channel in the picture to be synthesized; and selecting a target bitstream to synthesize the picture to be synthesized based on the data stream configuration information of the video channel, the encoding resolution and sub-picture configuration information of the picture to be synthesized, such that the picture corresponding to the video channel in the picture to be synthesized meets a preset display clarity requirement and the resolution of the target bitstream meets a preset resolution requirement, wherein the target bitstream is a main bitstream or a sub-bitstream.

[0009] Optionally, selecting a target bitstream to synthesize the image to be synthesized based on the data stream configuration information of the video channel, the encoding resolution of the image to be synthesized, and the sub-image configuration information includes: determining a target segmentation mode from multiple preset segmentation modes based on the data stream configuration information of the video channel, the encoding resolution of the image to be synthesized, and the sub-image configuration information, each preset segmentation mode having corresponding data stream configuration information, encoding resolution of the image to be synthesized, and sub-image configuration information; determining the target bitstream corresponding to the target segmentation mode based on a preset correspondence; and synthesizing the image to be synthesized using the target bitstream.

[0010] Optionally, selecting a target bitstream to synthesize the image to be synthesized based on the data stream configuration information of the video channel, the encoding resolution of the image to be synthesized, and the sub-image configuration information includes: determining the display resolution of the image corresponding to the video channel based on the encoding resolution of the image to be synthesized and the sub-image configuration information; determining the target bitstream based on the display resolution of the image corresponding to the video channel and the data stream configuration information; and using the target bitstream to synthesize the image to be synthesized.

[0011] Optionally, determining the target bitstream based on the display resolution of the image corresponding to the video channel and the data stream configuration information includes: determining whether the sub-bitstream of the video channel meets the display requirements, wherein if there is a resolution among the resolutions supported by the sub-bitstream that is greater than or equal to the target resolution, the display requirements are determined to be met; if there is no resolution among the resolutions supported by the sub-bitstream that is greater than or equal to the target resolution, the display requirements are determined not to be met; the target resolution is not less than 80% of the display resolution of the image corresponding to the video channel; if the sub-bitstream of the video channel meets the display requirements, the target bitstream is determined to be a sub-bitstream.

[0012] Optionally, if the sub-stream of the video channel meets the display requirements, determining the target stream as a sub-stream includes: selecting a resolution greater than or equal to the target resolution from the resolutions supported by the sub-stream, and using that resolution as the resolution of the sub-stream.

[0013] Optionally, if the sub-stream of the video channel meets the display requirements, determining the target stream as a sub-stream includes: if there exists a resolution among the sub-streams that is equal to the display resolution of the image corresponding to the video channel, adjusting the resolution of the sub-stream to be equal to the display resolution of the image corresponding to the video channel; and / or, if there exists a resolution among the sub-streams that is equal to the display resolution of the image corresponding to the video channel, selecting the minimum value from the resolutions supported by the sub-stream that are greater than the display resolution of the image corresponding to the video channel, and using that as the resolution of the sub-stream.

[0014] Optionally, determining the target bitstream based on the display resolution and data stream configuration information of the video channel includes: determining the target bitstream as the main bitstream when the sub-bitstream of the video channel does not meet the display requirements.

[0015] Optionally, determining the target bitstream based on the display resolution of the image corresponding to the video channel and the data stream configuration information includes: if the sub-bitstream of the video channel does not meet the display requirements, reducing at least a portion of the image corresponding to the main bitstream to a target size to obtain a first image; enlarging at least a portion of the image with the highest resolution supported by the sub-bitstream to a target size to obtain a second image, where the target size is the size occupied by the image corresponding to the video channel in the composite image; comparing the clarity of the first image and the second image; if the clarity of the first image is higher than that of the second image, determining the target bitstream as the main bitstream; and if the clarity of the second image is higher than that of the first image, determining the target bitstream as a sub-bitstream.

[0016] Optionally, the main bitstream has multiple supported resolutions. Based on the display resolution of the video channel and the data stream configuration information, the target bitstream is determined by: determining the resolution of the main bitstream as the minimum value among the candidate resolutions, where the candidate resolution is greater than or equal to the resolution of the video channel.

[0017] According to another aspect of the present invention, a video processing apparatus is also provided, comprising: a first acquisition unit, configured to acquire data stream configuration information of a video channel, the data stream configuration information including a main stream resolution and a resolution supported by a sub-stream, wherein the main stream resolution is higher than the resolution supported by the sub-stream; a second acquisition unit, configured to acquire the encoding resolution of a frame to be synthesized and sub-frame configuration information, wherein the sub-frame configuration information characterizes the size of the frame corresponding to the video channel in the frame to be synthesized; and a synthesis unit, configured to select a target stream to synthesize the frame to be synthesized based on the data stream configuration information of the video channel, the encoding resolution of the frame to be synthesized, and the sub-frame configuration information, so that the frame corresponding to the video channel in the frame to be synthesized meets a preset display clarity requirement and the resolution of the target stream meets a preset resolution requirement, wherein the target stream is a main stream or a sub-stream.

[0018] The video processing method in this embodiment of the invention includes: acquiring data stream configuration information of a video channel, the data stream configuration information including the main stream resolution and the resolution supported by the sub-stream, wherein the main stream resolution is higher than the resolution supported by the sub-stream; acquiring the encoding resolution and sub-picture configuration information of the image to be synthesized, wherein the sub-picture configuration information represents the size occupied by the image corresponding to the video channel in the image to be synthesized; and selecting a target stream to synthesize the image to be synthesized based on the data stream configuration information of the video channel, the encoding resolution and sub-picture configuration information of the image to be synthesized, so that the image corresponding to the video channel in the image to be synthesized meets the preset display clarity requirements and the resolution of the target stream meets the preset resolution requirements, wherein the target stream is the main stream or the sub-stream. In the process of image compositing using the above video processing method, information such as the video stream configuration information of the video channel, the encoding resolution of the image to be composited, and the sub-image configuration information are first obtained. The encoding resolution of the image to be composited is the resolution of the image to be composited after encoding. This resolution needs to be determined before image compositing. Once the encoding resolution of the image to be composited and the size occupied by the image corresponding to a certain video channel in the composite image are obtained, it is equivalent to knowing the resolution requirement of the video channel bitstream. By combining the main bitstream resolution and the resolution supported by the sub-bitstream in the data stream configuration information, the main bitstream or sub-bitstream can be selected to composite the image to be composited, so that the image corresponding to the video channel in the composite image meets the preset display clarity requirements and the resolution size of the target bitstream meets the preset resolution requirements. In this way, by combining resolution information to select either the main bitstream or the sub-bitstream to synthesize the image, instead of directly selecting the main bitstream, the clarity of the synthesized video image can be guaranteed while making reasonable use of the sub-bitstream. This helps to reduce the decoding performance requirements of the synthesized image on the display device while ensuring the display quality of the synthesized image. It balances the goals of saving decoding performance of the playback device and ensuring image display quality, thereby improving the playback effect of the synthesized image and solving the technical problem of poor actual playback effect of synthesized video images in related technologies. Attached Figure Description

[0019] The accompanying drawings, which are included to provide a further understanding of the invention and form part of this application, illustrate exemplary embodiments of the invention and, together with their description, serve to explain the invention and do not constitute an undue limitation thereof. In the drawings:

[0020] Figure 1 This is a flowchart illustrating an optional embodiment of the video processing method according to the present invention;

[0021] Figure 2 This is a schematic diagram of an optional embodiment of the video processing apparatus according to the present invention;

[0022] Figure 3 This is a schematic diagram of a synthesized image synthesized using the video processing method of this invention.

[0023] Figure 4 This is a schematic diagram of a video processing method according to an embodiment of the present invention when synthesizing images;

[0024] Figure 5 This is a schematic diagram of different preset segmentation modes of a video processing method according to an embodiment of the present invention;

[0025] Figure 6 This is a schematic diagram of a video processing method according to another embodiment of the present invention when synthesizing images. Detailed Implementation

[0026] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.

[0027] It should be noted that the terms "first," "second," etc., in the specification, claims, and drawings of this invention are used to distinguish different objects, rather than to limit a specific order.

[0028] Figure 1 This is a video processing method according to an embodiment of the present invention, such as... Figure 1 As shown, the method includes the following steps:

[0029] Step S102: Obtain the data stream configuration information of the video channel. The data stream configuration information includes the main stream resolution and the resolution supported by the sub-stream. The main stream resolution is higher than the resolution supported by the sub-stream.

[0030] Step S104: Obtain the encoding resolution and sub-picture configuration information of the picture to be composited. The sub-picture configuration information represents the size of the picture corresponding to the video channel in the picture to be composited.

[0031] Step S106: Based on the data stream configuration information of the video channel, the encoding resolution of the image to be synthesized, and the sub-image configuration information, select the target bitstream to synthesize the image to be synthesized, so that the image corresponding to the video channel in the image to be synthesized meets the preset display clarity requirements and the resolution of the target bitstream meets the preset resolution requirements. The target bitstream is either the main bitstream or the sub-bitstream.

[0032] The video processing method using the above scheme includes: obtaining the data stream configuration information of the video channel, which includes the main stream resolution and the resolution supported by the sub-stream, wherein the main stream resolution is higher than the resolution supported by the sub-stream; obtaining the encoding resolution and sub-picture configuration information of the image to be synthesized, wherein the sub-picture configuration information represents the size of the image corresponding to the video channel in the image to be synthesized; and selecting a target stream to synthesize the image to be synthesized based on the data stream configuration information of the video channel, the encoding resolution of the image to be synthesized, and the sub-picture configuration information, so that the image corresponding to the video channel in the image to be synthesized meets the preset display clarity requirements and the resolution of the target stream meets the preset resolution requirements, wherein the target stream is either the main stream or the sub-stream. In the process of image compositing using the above video processing method, the video stream configuration information of the video channel, the encoding resolution of the image to be composited, and the sub-image configuration information are first obtained. The encoding resolution of the image to be composited is the resolution of the image to be composited after encoding. This resolution needs to be determined before image compositing. With the encoding resolution of the image to be composited and the size occupied by the image corresponding to a certain video channel in the composite image, it is equivalent to knowing the resolution requirement of the video channel bitstream. By combining the main bitstream resolution and the resolution supported by the sub-bitstream in the data stream configuration information, the main bitstream or sub-bitstream can be selected to composite the image to be composited in a targeted manner, so that the image corresponding to the video channel in the image to be composited meets the preset display clarity requirements and the resolution size of the target bitstream meets the preset resolution requirements. On the basis of ensuring the clarity of the image to be composited, the bitstream with a lower resolution is selected. In this way, by combining resolution information to select either the main bitstream or the sub-bitstream to synthesize the image, instead of directly selecting the main bitstream, the clarity of the synthesized video image can be guaranteed while making reasonable use of the sub-bitstream. This helps to reduce the decoding performance requirements of the synthesized image on the display device while ensuring the display quality of the synthesized image. It balances the goals of saving decoding performance of the playback device and ensuring image display quality, thereby improving the playback effect of the synthesized image and solving the technical problem of poor actual playback effect of synthesized video images in related technologies.

[0033] The size mentioned here refers to the aspect ratio. For example, the frame corresponding to one video channel is called a sub-frame. The frame to be composited will contain sub-frames, and each sub-frame occupies a certain size within the composite frame. The sub-frame configuration information at least represents this size. Of course, the sub-frame configuration information can also represent other content, such as the position of the sub-frame within the composite frame. The video channel mentioned above can be understood as a video source. Different video channels represent different video sources. For example, if two cameras are used to access two video streams, then the data streams of two video channels are being acquired. The frame to be composited and the composited frame mentioned above can be understood as the same frame. Before being composited, this frame is called the frame to be composited; after being composited, it is called the composited frame.

[0034] The aforementioned requirement ensures that the video channels in the composite image meet preset display clarity requirements. This means that after compositing the image using the target bitstream, the clarity of the sub-images corresponding to the video channels in the composite image must meet the preset requirements. Since the main bitstream and sub-bitstreams have different resolutions, the clarity of the sub-images corresponding to the video channels may differ after compositing using different bitstreams. Therefore, when selecting the target bitstream, it is necessary to choose from the main and sub-bitstreams based on the preset display clarity requirements. Additionally, the target bitstream's resolution must meet preset resolution requirements. Since a higher resolution target bitstream results in a larger storage and bandwidth requirement for the composite video, a lower resolution bitstream should be selected as the target bitstream while still meeting display clarity requirements. Therefore, a preset resolution requirement is added here; that is, when determining the target bitstream, it is necessary to consider not only the display clarity of the sub-images but also the resolution requirements of the target bitstream. The aforementioned display clarity and resolution requirements can be flexibly set according to actual conditions.

[0035] There can be one or more video channels. When there are multiple channels, each one can be processed according to this video processing method. For example... Figure 3 As shown, Figure 3 The image to be composited consists of a background and image 1. At this time, there is one video channel, corresponding to the image of image 1. Figure 4 In this process, the image to be composited is composed of images from two video channels, namely IP channel 1 and IP channel 2. Each IP channel has a main stream and a sub-stream available for selection. Figure 4 In the embodiments described above, after employing the video processing method, IP channel 1 selects the main bitstream for image compositing, and IP channel 2 selects the sub-bitstream for image compositing. The encoding resolution of the image to be composited, i.e., the resolution of the newly composited image, is a preset value, such as a resolution value input by the user or a fixed resolution value.

[0036] In practice, based on the data stream configuration information of the video channel, the encoding resolution of the image to be composited, and the sub-image configuration information, there are several specific implementation methods for selecting the main bitstream or sub-bitstream of the video channel to composite the image to be composited:

[0037] For example, in one implementation, selecting a target bitstream to synthesize the image to be synthesized based on the data stream configuration information of the video channel, the encoding resolution of the image to be synthesized, and the sub-image configuration information includes: determining a target segmentation mode from multiple preset segmentation modes based on the data stream configuration information of the video channel, the encoding resolution of the image to be synthesized, and the sub-image configuration information, each preset segmentation mode having corresponding data stream configuration information, encoding resolution of the image to be synthesized, and sub-image configuration information; determining the target bitstream corresponding to the target segmentation mode based on a preset correspondence; and synthesizing the image to be synthesized using the target bitstream.

[0038] like Figure 5 As shown, it includes multiple fixed segmentation modes (preset segmentation modes), each with corresponding data stream configuration information, encoding resolution of the image to be composited, and sub-image configuration information. By setting multiple preset segmentation modes, each corresponding to its specific data stream configuration information, encoding resolution of the image to be composited, and sub-image configuration information, during implementation, after obtaining the aforementioned data stream configuration information, encoding resolution of the image to be composited, and sub-image configuration information, the corresponding segmentation mode (i.e., the target segmentation mode) can be matched. Knowing the target segmentation mode, the specific selection method of the target bitstream can be determined according to the preset correspondence. In other words, in this embodiment, some correspondences between segmentation modes and target bitstream selection methods are predefined, and subsequent selection only requires segmentation mode matching, effectively facilitating the target bitstream selection process.

[0039] In another implementation, selecting a target bitstream to synthesize the image to be synthesized based on the data stream configuration information of the video channel, the encoding resolution of the image to be synthesized, and the sub-image configuration information includes: determining the display resolution of the image corresponding to the video channel based on the encoding resolution of the image to be synthesized and the sub-image configuration information; determining the target bitstream based on the display resolution of the image corresponding to the video channel and the data stream configuration information; and using the target bitstream to synthesize the image to be synthesized.

[0040] After obtaining the encoding resolution and sub-picture configuration information of the image to be composited, since the sub-picture configuration information represents the size occupied by the image corresponding to a video channel in the image to be composited, the actual display resolution of the image corresponding to the video channel can be obtained (i.e., the resolution of the sub-picture; for example, if the sub-picture configuration information represents that the image of a certain video channel occupies 1 / 2 of the image to be composited in both length and width, then the resolution of the sub-picture corresponding to that video channel in both length and width is half of the corresponding encoding resolution of the image to be composited). Once the display resolution of the image corresponding to a certain video channel is obtained, the resolution requirement for displaying the sub-picture corresponding to that video channel is known. Based on this resolution requirement, the main bitstream or sub-bitstream can be selected specifically to composite the image. In this embodiment, by determining the display resolution of the image corresponding to a video channel and then selecting the main bitstream or sub-bitstream based on that resolution, more flexible selection of bitstream type can be achieved. There is no need to predefine the correspondence between the segmentation mode and the target bitstream selection method, making it more flexible in actual use. Users can adjust the size of the image corresponding to each video channel in the image to be composited according to actual needs.

[0041] In one embodiment, obtaining sub-screen configuration information includes: receiving an adjustment instruction from a user, the adjustment instruction being used to adjust the size of the screen corresponding to a video channel in the composite screen; and determining screen configuration information based on the adjustment instruction. That is, the user can adjust the size of the screen corresponding to any video channel in the composite screen, and based on the user's adjustment instruction, the size of the screen corresponding to the video channel in the composite screen can be obtained, thereby acquiring the sub-screen configuration information.

[0042] Based on the display resolution and data stream configuration information of the video channel's corresponding image, determining the target bitstream includes: determining whether the video channel's sub-bitstreams meet the display requirements. Specifically, if any sub-bitstream supports a resolution greater than or equal to the target resolution, the display requirements are met; otherwise, they are not met. The target resolution must be no less than 80% of the display resolution of the video channel's corresponding image. If the video channel's sub-bitstreams meet the display requirements, the target bitstream is determined to be a sub-bitstream. As mentioned above, the selection of the target bitstream requires balancing the goals of saving decoding performance on the playback device and ensuring image display quality, ensuring that the final composite image has a smaller footprint in terms of storage space and network bandwidth, while also achieving a better display effect. To more rationally select the target bitstream, in this embodiment, if any sub-bitstream supports a resolution greater than or equal to the display resolution of the image corresponding to the video channel (the resolution of the sub-image), it indicates that using the sub-bitstream can meet the display requirements of the image corresponding to the video channel. In this case, selecting the sub-bitstream as the target bitstream can reduce the size of the composite image without compromising its display clarity, thereby saving decoding performance on the display device and reducing the composite image's footprint on storage space and network bandwidth. Specifically, the target resolution is not less than 80% of the display resolution of the image corresponding to the video channel; that is, the number of pixels in both the length and width directions of the target resolution is not less than 80% of the number of pixels in both the length and width directions of the display resolution of the image corresponding to the video channel.

[0043] In practical implementation, determining the target bitstream as a sub-bitstream, assuming the sub-bitstream of the video channel meets the display requirements, involves selecting a resolution greater than or equal to the target resolution from the resolutions supported by the sub-bitstream. If the sub-bitstream meets the display requirements, any resolution greater than or equal to the target resolution supported by the sub-bitstream can be selected. This ensures the display effect of the sub-pictures in the composite image, and by using sub-pictures for compositing, the decoding performance of the playback device can be effectively saved.

[0044] In practice, sub-streams may support multiple resolutions. To balance saving decoding performance on playback devices and ensuring image display quality, if a sub-stream supports a resolution equal to the display resolution of the image corresponding to the video channel (the sub-image's resolution), then that sub-stream resolution can be directly selected to satisfy both objectives. If no sub-stream supports a resolution equal to the display resolution of the image corresponding to the video channel (the sub-image's resolution), then the minimum value among the sub-stream's supported resolutions that are greater than the display resolution of the image corresponding to the video channel is selected as the sub-stream's resolution. This approach minimizes image quality while ensuring the clarity of the composite image, thereby saving decoding performance on the playback device. Specifically, when the sub-stream of the video channel meets the display requirements, determining the target stream as a sub-stream includes: if there exists a resolution among the sub-streams that is equal to the display resolution of the image corresponding to the video channel, adjusting the resolution of the sub-stream to be equal to the display resolution of the image corresponding to the video channel; and / or, if there exists a resolution among the sub-streams that is equal to the display resolution of the image corresponding to the video channel, selecting the minimum value from the resolutions supported by the sub-stream that are greater than the display resolution of the image corresponding to the video channel, and using that as the resolution of the sub-stream.

[0045] When the resolutions supported by the sub-streams of a certain video channel are all smaller than the display resolution of the corresponding screen of that video channel, there can be different target bitstream selection methods.

[0046] In one embodiment, to ensure the clarity of the composite image, the main bitstream is selected as the target bitstream. Specifically, determining the target bitstream based on the display resolution of the image corresponding to the video channel and the data stream configuration information includes: if the sub-bitstreams of the video channel do not meet the display requirements, the target bitstream is determined to be the main bitstream.

[0047] At least a portion of the image corresponding to the main bitstream is reduced to a target size to obtain a first image. At least a portion of the image with the highest resolution supported by the sub-bitstream is enlarged to the target size to obtain a second image. By comparing the clarity of the first and second images, the scheme with higher clarity is selected as the target bitstream. Specifically, determining the target bitstream based on the display resolution of the screen corresponding to the video channel and the data stream configuration information includes: if the sub-bitstream of the video channel does not meet the display requirements, at least a portion of the image corresponding to the main bitstream is reduced to a target size to obtain a first image; at least a portion of the image with the highest resolution supported by the sub-bitstream is enlarged to the target size to obtain a second image, where the target size is the size occupied by the screen corresponding to the video channel in the composite screen; the clarity of the first and second images is compared; if the clarity of the first image is higher than that of the second image, the target bitstream is determined to be the main bitstream; if the clarity of the second image is higher than that of the first image, the target bitstream is determined to be a sub-bitstream. By downscaling at least a portion of the image corresponding to the main bitstream and upscaling at least a portion of the image with the highest resolution supported by the sub-bitstream, and comparing the display clarity of the two, the bitstream with higher clarity is selected as the target bitstream. This makes the image corresponding to that video channel in the composite image clearer. In practice, the objects used for scaling and comparison can be flexibly adjusted. That is, at least a portion of the image corresponding to the main bitstream and at least a portion of the image corresponding to the sub-bitstream can be flexibly selected, as long as they are comparable. For example, the entire image corresponding to the main bitstream and the entire image corresponding to the sub-bitstream can be scaled to the target size for comparison. Another example is scaling 90% of the image corresponding to the main bitstream and 90% of the image corresponding to the sub-bitstream to the target size for comparison, as long as it can be determined which image is clearer after compositing using the main bitstream and the sub-bitstream.

[0048] During the process of compositing the images to be composited, if the sub-stream does not support a resolution equal to the display resolution of the image corresponding to the video channel, the image corresponding to the sub-stream needs to be scaled to the target size. Similarly, if the main stream is selected as the target stream, the image corresponding to the main stream also needs to be scaled down to the target size.

[0049] In actual implementation, the main bitstream can also support multiple resolutions. In this case, there are more options available for image compositing. In order to ensure the display clarity of the sub-images in the image to be composited and to save the decoding performance of the playback device, in a preferred embodiment, the target bitstream is determined according to the display resolution of the image corresponding to the video channel and the data stream configuration information, including: determining the resolution of the main bitstream as the minimum value among the candidate resolutions, wherein the candidate resolution is greater than or equal to the resolution of the image corresponding to the video channel.

[0050] The video processing method of the present invention will be described below with reference to a specific embodiment:

[0051] Cameras employing dual-stream technology have encoders with two encoding formats: a main stream and a sub-stream. The main stream is characterized by high bitrate, high bandwidth usage, high resolution, and high image quality. Conversely, the sub-stream is characterized by low bitrate, low bandwidth usage, low resolution, and low image quality. In local transmission scenarios (such as recording on a video recorder) and local previewing on a video recorder, the main stream can be used to ensure display clarity. However, for remote transmission scenarios (such as remote previewing), bandwidth limitations can significantly impact smoothness. Therefore, reducing image quality and bandwidth can improve performance, making the sub-stream a more suitable choice.

[0052] In some scenarios, it's necessary to encode and synthesize multiple channel streams into a single composite stream for recording or burning. Currently, the composite stream uses the main stream, regardless of the actual encoding parameters of the sub-channels or their size within the composite image. This approach has the following problems: First, with a large number of sub-channels, the device's decoding performance is wasted. In extreme cases, the device's decoding performance may be insufficient to decode all channels simultaneously, resulting in a black screen for each channel. Second, the main stream typically needs to be decoded, scaled, and then encoded into the composite channel. If the main stream has a high resolution, scaling the source stream can also degrade image quality, resulting in a final video quality in the composite image that is even worse than the video synthesized using sub-streams.

[0053] To make it easier to understand, an example is given below, such as... Figure 3 As shown, assuming the composite image encoding resolution is 720P, the image segmentation configuration window display size is 1024*768, and the sub-image occupies 1 / 9 of the entire image (1 / 3 in both length and width), then:

[0054] The virtual size of screen 1 (the size of screen 1 in the split configuration window, in pixels) is: (1024 / 3)*(768 / 3)≈342*256;

[0055] The physical size of the sub-picture in the composite image (i.e., the actual size displayed in picture 1, in pixels) is: (1280 / 3)*(720 / 3)

[0056] ≈427*240;

[0057] Since the physical size of the sub-screen is 427*240, a sub-bitstream with a resolution of 2cif (704*288) is sufficient for its use.

[0058] If the resolution of the composite image is changed to 2560*1440, then the physical size of the sub-image will simultaneously become (2560 / 3)*(1440 / 3).

[0059] ≈852*480; If the resolution of the composite image is switched to 4096*2160, then the physical size of the sub-image will simultaneously become (4096 / 3)*(2160 / 3)≈1366*720. It can be seen that the choice of sub-image resolution is inseparable from the encoding channel resolution parameter.

[0060] The video processing method of this invention can select different bitstream selection schemes when performing video processing:

[0061] In one implementation, such as Figure 4 and Figure 5 As shown, the composite image adopts a fixed segmentation mode (preset segmentation mode). For commonly used fixed segmentation modes, since the size ratio of each image is fixed, a set of calculated parameter selection logic can be pre-selected based on the resolution supported by the composite image, directly realizing the selection of main and sub-bitstreams. That is, by setting multiple preset segmentation modes, each preset segmentation mode corresponds to its specific data stream configuration information, the encoding resolution of the image to be composited, and the sub-image configuration information. During implementation, after obtaining the above-mentioned data stream configuration information, the encoding resolution of the image to be composited, and the sub-image configuration information, the corresponding segmentation mode (i.e., the target segmentation mode) can be matched. Knowing the target segmentation mode, the specific selection method of the target bitstream can be determined according to the preset correspondence. In other words, in this embodiment, some correspondences between segmentation modes and target bitstream selection methods are predefined. Subsequent selection only requires segmentation mode matching, effectively facilitating the target bitstream selection process.

[0062] In another implementation, such as Figure 6As shown, users can manually adjust the size of the sub-screen. At this point, it's necessary to dynamically calculate parameters based on the adjusted sub-screen size and then select the bitstream. The specific implementation process is as follows: First, determine if the camera channel (video channel) supports sub-bitstreams. If the camera channel doesn't support sub-bitstreams, directly select the main bitstream. If the camera channel supports sub-bitstreams, obtain the encoding resolution of the compositing channel (the encoding resolution of the image to be composited), and calculate the physical size (W*H) of the sub-screen based on the size relationship between the image to be composited and the sub-screen (sub-screen configuration information). Combine this with the resolutions supported by the sub-bitstream in the data stream configuration information. If the sub-bitstream meets the image quality requirements, select it for compositing encoding, and modify the sub-bitstream resolution to the corresponding parameters. If no resolution meets the requirements in the sub-bitstream, differentiate based on the actual device situation: the main bitstream can be used directly for compositing; alternatively, the image quality can be compared between scaling the main bitstream and enlarging the sub-screen, and the bitstream with the least impact on image quality can be selected for compositing. For example, in this embodiment, by determining whether the encoding capability of the sub-stream supports the physical size (W*H) of the sub-picture, that is, by determining whether there is a resolution among the resolutions supported by the sub-stream that is equal to the display resolution of the picture corresponding to the video channel, if so, the sub-stream can be directly used for picture compositing. If not, the first encoding parameter greater than W*H will be selected from the sub-stream encoding capability from low to high. If this parameter exists, the sub-stream will be used for picture compositing. In other words, if there is no resolution among the resolutions supported by the sub-stream that is equal to the display resolution (sub-picture resolution) of the picture corresponding to the video channel, the minimum value among the resolutions supported by the sub-stream that are greater than the display resolution of the picture corresponding to the video channel will be selected as the resolution of the sub-stream. This can reduce the picture quality as much as possible while ensuring the display clarity requirements of the composite picture, thereby saving the decoding performance of the playback device. If the resolution supported by the sub-streams of a certain video channel is less than the display resolution of the corresponding video channel, the encoding parameters of the IPC main stream (i.e., the camera main stream) and the maximum sub-stream encoding parameters will be obtained. The image quality of the two methods, namely, reducing the sub-video main stream and increasing the sub-stream, will be compared to select the appropriate stream (target stream).

[0063] In other words, when the resolutions supported by all sub-streams of a video channel are lower than the display resolution of the corresponding image, different target bitstream selection methods are possible. One feasible method is to select the main bitstream as the target bitstream to ensure the clarity of the composite image. Another feasible method involves scaling down the image corresponding to the main bitstream to the target size to obtain the first image, and scaling up the image with the highest resolution supported by the sub-streams to the target size to obtain the second image. By comparing the clarity of the first and second images, the one with higher clarity is selected as the target bitstream.

[0064] In addition, such as Figure 2As shown, embodiments of the present invention also provide a video processing apparatus, comprising: a first acquisition unit, configured to acquire data stream configuration information of a video channel, the data stream configuration information including the main stream resolution and the resolution supported by the sub-stream, wherein the main stream resolution is higher than the resolution supported by the sub-stream; a second acquisition unit, configured to acquire the encoding resolution of the image to be synthesized and sub-image configuration information, wherein the sub-image configuration information characterizes the size of the image corresponding to the video channel in the image to be synthesized; and a synthesis unit, configured to select a target stream to synthesize the image to be synthesized based on the data stream configuration information of the video channel, the encoding resolution of the image to be synthesized, and the sub-image configuration information, so that the image corresponding to the video channel in the image to be synthesized meets a preset display clarity requirement and the resolution of the target stream meets a preset resolution requirement, wherein the target stream is a main stream or a sub-stream. In the process of image compositing using the aforementioned video processing device, the first acquisition unit and the second acquisition unit first acquire information such as the video stream configuration information of the video channel, the encoding resolution of the image to be composited, and the sub-image configuration information. The encoding resolution of the image to be composited is the resolution of the image to be composited after encoding. This resolution needs to be determined before image compositing. Once the encoding resolution of the image to be composited and the size occupied by the image corresponding to a certain video channel in the composite image are obtained, it is equivalent to knowing the resolution requirement of the video channel bitstream. At this time, the compositing unit can selectively choose the main bitstream or sub-bitstream to composite the image to be composited by combining the main bitstream resolution and the resolution supported by the sub-bitstream in the data stream configuration information, so that the image corresponding to the video channel in the image to be composited meets the preset display clarity requirements and the resolution size of the target bitstream meets the preset resolution requirements. In this way, by combining resolution information to select either the main bitstream or the sub-bitstream to synthesize the image, instead of directly selecting the main bitstream, the clarity of the synthesized video image can be guaranteed while making reasonable use of the sub-bitstream. This helps to reduce the decoding performance requirements of the synthesized image on the display device while ensuring the display quality of the synthesized image. It balances the goals of saving decoding performance of the playback device and ensuring image display quality, thereby improving the playback effect of the synthesized image and solving the technical problem of poor actual playback effect of synthesized video images in related technologies.

[0065] In one embodiment, the compositing unit is configured to: determine a target segmentation mode from multiple preset segmentation modes based on the data stream configuration information of the video channel, the encoding resolution of the image to be composited, and the sub-image configuration information, wherein each preset segmentation mode has corresponding data stream configuration information, encoding resolution of the image to be composited, and sub-image configuration information; determine the target bitstream corresponding to the target segmentation mode based on a preset correspondence; and composite the image to be composited using the target bitstream.

[0066] In another embodiment, the compositing unit includes: a first determining module, configured to determine the display resolution of the image corresponding to the video channel based on the encoding resolution of the image to be composited and the sub-image configuration information; a second determining module, configured to determine the target bitstream based on the display resolution of the image corresponding to the video channel and the data stream configuration information; and a compositing module, configured to composite the image to be composited using the target bitstream.

[0067] The second determining module includes: a first determining submodule for determining whether a sub-stream of the video channel meets the display requirements, wherein if there is a resolution greater than or equal to the target resolution among the resolutions supported by the sub-stream, the display requirements are determined to be met; if there is no resolution greater than or equal to the target resolution among the resolutions supported by the sub-stream, the display requirements are determined not to be met, and the target resolution is not less than 80% of the display resolution of the image corresponding to the video channel; and a second determining submodule for determining the target stream as a sub-stream if the sub-stream of the video channel meets the display requirements.

[0068] The second determining submodule is used to: select a resolution greater than or equal to the target resolution from the resolutions supported by the sub-bitstream, as the resolution of the sub-bitstream.

[0069] The second determining submodule is used to: adjust the resolution of the sub-stream to be equal to the display resolution of the image corresponding to the video channel if there exists a resolution among the resolutions supported by the sub-stream that is equal to the display resolution of the image corresponding to the video channel; and / or, if there exists a resolution among the resolutions supported by the sub-stream that is equal to the display resolution of the image corresponding to the video channel, select the minimum value from the resolutions supported by the sub-stream that are greater than the display resolution of the image corresponding to the video channel, and use it as the resolution of the sub-stream.

[0070] In an optional embodiment, the second determining module is used to: determine the target bitstream as the main bitstream when the sub-bitstream of the video channel does not meet the display requirements.

[0071] In another optional embodiment, the second determining module is configured to: reduce at least a portion of the image corresponding to the main stream to a target size to obtain a first image when the sub-stream of the video channel does not meet the display requirements; enlarge at least a portion of the image with the highest resolution supported by the sub-stream to a target size to obtain a second image, wherein the target size is the size occupied by the image corresponding to the video channel in the composite image; compare the clarity of the first image and the second image; determine the target stream as the main stream when the clarity of the first image is higher than that of the second image, and determine the target stream as a sub-stream when the clarity of the second image is higher than that of the first image.

[0072] In one embodiment, the second determining module is further configured to: determine the resolution of the main bitstream as the minimum value among the candidate resolutions, wherein the candidate resolution is greater than or equal to the resolution of the frame corresponding to the video channel.

[0073] The above description is merely a preferred embodiment of the present invention and is not intended to limit the scope of protection of the present invention.

Claims

1. A video processing method, characterized in that, include: Obtain the data stream configuration information of the video channel, the data stream configuration information including the main stream resolution and the resolution supported by the sub-stream, wherein the main stream resolution is higher than the resolution supported by the sub-stream; Obtain the encoding resolution and sub-picture configuration information of the image to be composited, wherein the sub-picture configuration information represents the size occupied by the image corresponding to the video channel in the image to be composited; Based on the data stream configuration information of the video channel, the encoding resolution of the image to be synthesized, and the sub-image configuration information, a target bitstream is selected to synthesize the image to be synthesized, so that the image corresponding to the video channel in the image to be synthesized meets the preset display clarity requirements and the resolution of the target bitstream meets the preset resolution requirements. The target bitstream is the main bitstream or the sub-bitstream. Selecting a target bitstream to synthesize the image to be synthesized based on the data stream configuration information of the video channel, the encoding resolution of the image to be synthesized, and the sub-image configuration information includes: determining the display resolution of the image corresponding to the video channel based on the encoding resolution of the image to be synthesized and the sub-image configuration information; determining the target bitstream based on the display resolution of the image corresponding to the video channel and the data stream configuration information; and synthesizing the image to be synthesized using the target bitstream. Determining the target bitstream based on the display resolution of the image corresponding to the video channel and the data stream configuration information includes: determining whether the sub-bitstream of the video channel meets the display requirements; if the sub-bitstream of the video channel does not meet the display requirements, reducing the image corresponding to the main bitstream to a target size to obtain a first image, and enlarging the image with the highest resolution supported by the sub-bitstream to the target size to obtain a second image, wherein the target size is the size occupied by the image corresponding to the video channel in the image to be synthesized; comparing the clarity of the first image and the second image; if the clarity of the first image is higher than that of the second image, determining the target bitstream as the main bitstream, and if the clarity of the second image is higher than that of the first image, determining the target bitstream as the sub-bitstream.

2. The video processing method according to claim 1, characterized in that, If any resolution supported by the sub-stream is greater than or equal to the target resolution, the display requirement is determined to be met; if no resolution supported by the sub-stream is greater than or equal to the target resolution, the display requirement is determined to be unmet. The target resolution is not less than 80% of the display resolution of the image corresponding to the video channel. If the sub-stream of the video channel meets the display requirements, the target stream is determined to be the sub-stream.

3. The video processing method according to claim 2, characterized in that, Determining the target stream as the sub-stream when the sub-stream of the video channel meets the display requirements includes: Select a resolution greater than or equal to the target resolution from the resolutions supported by the sub-stream, and use it as the resolution of the sub-stream.

4. The video processing method according to claim 2, characterized in that, Determining the target stream as the sub-stream when the sub-stream of the video channel meets the display requirements includes: If, among the resolutions supported by the sub-stream, there exists a resolution equal to the display resolution of the image corresponding to the video channel, the resolution of the sub-stream is adjusted to be equal to the display resolution of the image corresponding to the video channel; and / or, If there is no resolution among the resolutions supported by the sub-stream that is equal to the display resolution of the image corresponding to the video channel, the minimum value is selected from the resolutions supported by the sub-stream that are greater than the display resolution of the image corresponding to the video channel, and this minimum value is used as the resolution of the sub-stream.

5. The video processing method according to claim 1, characterized in that, The main bitstream supports multiple resolutions. Based on the display resolution of the video channel and the data stream configuration information, the target bitstream is determined to include: The resolution of the main stream is determined to be the minimum value among the candidate resolutions, wherein the candidate resolution is greater than or equal to the resolution of the frame corresponding to the video channel.

6. A video processing apparatus, characterized in that, include: The first acquisition unit is used to acquire data stream configuration information of the video channel. The data stream configuration information includes the main stream resolution and the resolution supported by the sub-stream. The main stream resolution is higher than the resolution supported by the sub-stream. The second acquisition unit is used to acquire the encoding resolution and sub-picture configuration information of the picture to be synthesized, wherein the sub-picture configuration information represents the size occupied by the picture corresponding to the video channel in the picture to be synthesized; The compositing unit is configured to select a target bitstream to compose the image to be composed based on the data stream configuration information of the video channel, the encoding resolution of the image to be composed, and the sub-image configuration information, so that the image corresponding to the video channel in the image to be composed meets the preset display clarity requirements and the resolution of the target bitstream meets the preset resolution requirements. The target bitstream is the main bitstream or the sub-bitstream. The compositing unit includes: a first determining module, configured to determine the display resolution of the image corresponding to the video channel based on the encoding resolution of the image to be composited and the sub-image configuration information; a second determining module, configured to determine the target bitstream based on the display resolution of the image corresponding to the video channel and the data stream configuration information; and a compositing module, configured to composite the image to be composited using the target bitstream. The second determining module includes a first determining submodule, which is used to determine whether the sub-bitstream of the video channel meets the display requirements; the second determining module is further used to: when the sub-bitstream of the video channel does not meet the display requirements, reduce the image corresponding to the main bitstream to a target size to obtain a first image, and enlarge the image with the highest resolution supported by the sub-bitstream to the target size to obtain a second image, wherein the target size is the size occupied by the image corresponding to the video channel in the image to be synthesized; compare the clarity of the first image and the second image; when the clarity of the first image is higher than the clarity of the second image, determine the target bitstream as the main bitstream, and when the clarity of the second image is higher than the clarity of the first image, determine the target bitstream as the sub-bitstream.