A video processing method, apparatus, device, and storage medium
By dividing video images into blocks and adjusting quantization parameters, the problem of inaccurate ROI determination in existing technologies is solved, thereby improving video compression efficiency and quality.
Patent Information
- Application Number
- CN202210827116.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-07-13
- Publication Date
- 2026-02-06
- Estimated Expiration
- 2042-07-13
AI Technical Summary
Existing video compression methods cannot accurately determine the region of interest, resulting in poor video compression performance and low efficiency.
By dividing the video image into blocks, the ROI and non-ROI are determined, and the quantization parameters are dynamically adjusted according to the target bit count and location information to perform differentiated compression coding.
Accurately identify ROI and non-ROI, reduce the image quality of non-ROI to save bandwidth, while preserving image details of ROI to improve video visual effects.
Smart Images

Figure CN117440160B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The embodiment of the present disclosure relates to the technical field of video processing, in particular to a video processing method, device, equipment and storage medium. BACKGROUND
[0002] In the application implementation of video transmission, how to reduce the video code rate while ensuring the video quality has become a problem focused on by the video technology field.
[0003] Since the spherical video often has a high resolution (such as 2K, 4K or even 8K), when transmitting such a video, if the video quality is to be ensured, the video code rate required is also often high. Therefore, before video transmission, the demand for subjective lossless compression processing of such a video (that is, compressing the video while the subjective perception of the human eye remains unchanged) is even more intense.
[0004] The existing video compression method, for example, cannot accurately determine the region of interest (ROI) (for example, the determined ROI area is too large or even the determined ROI is incorrect) and has a low compression efficiency, resulting in poor video compression effect. SUMMARY
[0005] The present disclosure provides a video processing method, device, equipment and storage medium to effectively determine the important image region in the video and effectively improve the video compression effect.
[0006] In a first aspect, the embodiment of the present disclosure provides a video processing method, which comprises:
[0007] dividing an image in a to-be-processed video into blocks to obtain a plurality of block images;
[0008] determining a region of interest (ROI) and a non-ROI of the image based on the image and the plurality of block images;
[0009] determining a first quantization parameter of the ROI and a second quantization parameter of the non-ROI according to a preset target bit number, position information of the ROI and the non-ROI;
[0010] compressing and encoding the image according to the first quantization parameter and the second quantization parameter to obtain a target video.
[0011] In a second aspect, the embodiment of the present disclosure further provides a video processing device, which comprises:
[0012] an image division module configured to divide an image in a to-be-processed video into blocks to obtain a plurality of block images;
[0013] The first determining module is configured to determine the ROI and the non-ROI of the image based on the image and the plurality of block images.
[0014] The second determining module is configured to determine a first quantization parameter of the ROI and a second quantization parameter of the non-ROI according to a preset target bit number, the ROI, and position information of the non-ROI.
[0015] The graph encoding module is configured to compress and encode the image according to the first quantization parameter and the second quantization parameter, to obtain a target video.
[0016] In a third aspect, the embodiments of the present disclosure further provide an electronic device, which comprises:
[0017] one or more processors;
[0018] a storage device configured to store one or more programs,
[0019] When the one or more programs are executed by the one or more processors, the one or more processors implement the video processing method provided in the first aspect of the embodiments of the present disclosure.
[0020] In a fourth aspect, the embodiments of the present disclosure further provide a storage medium containing computer executable instructions, which, when executed by a computer processor, are used to perform the video processing method provided in the first aspect of the embodiments of the present disclosure.
[0021] The embodiments of the present disclosure provide a video processing method, device, equipment and storage medium. The method comprises: first dividing an image in a to-be-processed video into blocks to obtain a plurality of block images, then determining the ROI and the non-ROI of the image based on the image and the plurality of block images, then determining the first quantization parameter of the ROI and the second quantization parameter of the non-ROI according to a preset target bit number, the ROI and the position information of the non-ROI, and finally compressing and encoding the image according to the first quantization parameter and the second quantization parameter to finally obtain a target video of the to-be-processed video. The above technical solution can effectively predict the ROI and the non-ROI of the image from the image of the video, realize the differential determination of the ROI, and further effectively determine the quantization parameters of the ROI and the non-ROI in the compression and encoding by the ROI and the non-ROI, so as to realize the differential compression and encoding of the image. The above technical solution can accurately determine the ROI and the non-ROI, effectively reduce the image quality of the non-ROI by the compression and encoding based on the second quantization parameter, thereby greatly reducing the code rate and saving the bandwidth flow in the video transmission, and can better retain the image details of the ROI by the compression and encoding based on the first quantization parameter, thereby greatly improving the visual effect of the video and effectively improving the user viewing experience. BRIEF DESCRIPTION OF DRAWINGS
[0022] The above and other features, aspects and advantages of embodiments of the present disclosure will become more apparent from the following detailed description when taken in conjunction with the accompanying drawings. Like or similar elements and / or features throughout the drawings are denoted by identical reference numbers. It should be understood that the drawings are schematic and elements and features not be necessarily to scale.
[0023] Figure 1 A flowchart of a video processing method provided by an embodiment of the present disclosure;
[0024] Figure 2 An example implementation effect diagram of target heat map determination in a video processing method provided by an embodiment of the present disclosure
[0025] Figure 3 A structural diagram of a video processing apparatus provided by an embodiment of the present disclosure;
[0026] Figure 4 A structural diagram of an electronic device provided by an embodiment of the present disclosure. DETAILED DESCRIPTION
[0027] Embodiments of the present disclosure will be described in more detail with reference to the drawings. While certain embodiments of the present disclosure will be shown and described below, it is to be understood that the present disclosure can be embodied in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided as part of the disclosure to convey the principles and subtleties of the present disclosure to those skilled in the art.
[0028] It should be understood that each step recited in the method embodiments of the present disclosure can be executed in different orders and / or in parallel. In addition, the method embodiments can include additional steps and / or omit the execution of the steps shown. The scope of the present disclosure is not limited in this respect.
[0029] The term "comprising" and variations thereof as used herein are used inclusively, i.e., "comprising but not limited to." The term "based on" is "based at least in part on." The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments." Related terms are defined in the description that follows.
[0030] It should be noted that the terms "first", "second", and the like in the present disclosure are used only to distinguish different devices, modules or units, and do not imply the order or interdependence of the functions performed by these devices, modules or units.
[0031] It should be noted that the modification of "one" and "multiple" mentioned in the present disclosure is illustrative rather than limiting, and those skilled in the art should understand that "one or more" should be understood unless otherwise explicitly indicated in the context.
[0032] The names of the messages or information exchanged between the plurality of devices in the embodiments of the present disclosure are only for illustrative purposes, and are not used to limit the scope of the messages or information.
[0033] It can be understood that, before using the technical solutions disclosed in the embodiments of the present disclosure, the type, use range, use scenario, etc. of the personal information involved in the present disclosure should be informed to the user and the authorization of the user should be obtained through appropriate means according to relevant laws and regulations.
[0034] For example, in response to receiving the active request of the user, the prompt information is sent to the user to explicitly prompt the user that the operation requested to be performed will require obtaining and using the personal information of the user. Thus, the user can voluntarily choose whether to provide the personal information to the software or hardware such as electronic device, application program, server or storage medium, etc. that performs the operation of the technical solutions of the present disclosure according to the prompt information.
[0035] As an optional but non-limiting implementation manner, in response to receiving the active request of the user, the prompt information can be sent to the user in the form of a pop-up window, and the prompt information can be presented in the form of text in the pop-up window. In addition, the pop-up window can also carry selection controls for the user to select "agree" or "disagree" to provide the personal information to the electronic device.
[0036] It can be understood that the above notification and user authorization process is only illustrative, and does not limit the implementation manner of the present disclosure. Other manners that meet the relevant laws and regulations can also be applied to the implementation manner of the present disclosure.
[0037] Figure 1 A flowchart of a video processing method provided by the embodiments of the present disclosure is applicable to the case of encoding and compressing the images in the video. The method can be executed by a video processing device, which can be realized in the form of software and / or hardware. Optionally, the video processing device can be realized by an electronic device as an execution terminal. The electronic device can be a mobile terminal, a PC terminal or a server, etc.
[0038] As shown in FIG. 1, the video processing method provided by the embodiments of the present disclosure specifically includes the following operations: Figure 1
[0039] S110, block division is performed on the images in the to-be-processed video to obtain a plurality of block images.
[0040] In the embodiment, the video to be processed can be a short video or a live video to be transmitted by a video server to a viewer user terminal.
[0041] In the embodiment, the video to be processed can be a short video or a live video to be transmitted by a video server to a viewer user terminal.
[0042] In the embodiment, for the original video in the three-dimensional space, the equidistant cylindrical projection manner can be used for the optimization processing to the two-dimensional space, and finally the video to be processed in the two-dimensional space is obtained. Therefore, in the embodiment, the video to be processed can be a spherical video after the equidistant cylindrical projection processing, and correspondingly, the image contained in the video to be processed obtained after the equidistant cylindrical projection processing can be a full-view image or a partial-view image.
[0043] It can be known that the processing to be performed on the video to be processed in the embodiment can be video compression processing, and the compression processing of the video to be processed in the embodiment is specifically converted into the compression processing of each frame image in the video to be processed, and the compression processing of the image can be specifically implemented through each execution step provided in the embodiment.
[0044] In the specific implementation, the image in the video to be processed can be first divided into blocks through the step, and thus a plurality of block images are obtained relative to the image. The block division manner used can be to divide the image according to the given division row and column values, so as to form a plurality of rectangular image blocks, which are respectively denoted as block images. For the plurality of block images formed by division, the embodiment can preferably have the same block size.
[0045] S120, determining the ROI and the non-ROI of the image based on the image and the plurality of block images.
[0046] In the embodiment, the ROI of the image can be a region of interest in the image, and correspondingly, the non-ROI can be a region not of interest in the image.
[0047] The implementation manner of determining the ROI can include directly determining a specified region in the image as the ROI of the image, wherein the specified region is often determined through historical experience, for example, for a full-view / partial-view image, the entire equatorial region in the image is currently taken as the ROI. However, the area of the determined ROI (the entire equatorial region) is too large, or the determined ROI is not the region of interest of the user, for example, the region of interest of the user can be the north and south polar regions of the image.
[0048] The problems of the above-mentioned existing ROI determination manner directly affect the image compression processing effect. For example, an excessively large ROI area limits the space for reducing the code rate in the compression processing stage; for another example, when the determined ROI is not the region of interest of the user, the image compression processing will reduce the quality of the region of interest of the user, and the lower quality will affect the video viewing effect of the user.
[0049] Therefore, in the implementation of the present embodiment, the ROI determination is realized by S110 and S120. Compared with the existing ROI determination, the present embodiment can detect the ROI of the entire image and the ROI of each block image respectively by S110 and S120, and then fuse the detection results of the ROI detection of the entire image and the detection results of the ROI detection of each block image to determine the final ROI and non-ROI of the image.
[0050] As an implementation manner, the present embodiment can use an RIO detection network model to detect the RIO of the entire image and each region. The ROI determination manner used in the present embodiment realizes the differential determination of the ROI in different images, and ensures that the predicted ROI is more matched with the image region of interest of the user.
[0051] S130, determining a first quantization parameter of the ROI and a second quantization parameter of the non-ROI according to the preset target bit number, the ROI, and the position information of the non-ROI.
[0052] In the present embodiment, the target bit number can be understood as the number of bits expected to be used for compression encoding of the image. The ROI and the non-ROI are the region of interest and the non-region of interest in the image, respectively. After the ROI and the non-ROI are determined, the position information of the ROI and the non-ROI in the image is determined.
[0053] In the present embodiment, the quantization parameter (quantization parameter, QP) reflects the spatial detail compression condition in the compression processing. The quantization parameter can be the serial number of the quantization step. Generally, the quantization step can have different length ranges according to different encoding objects. For example, for the encoding of image brightness, the corresponding quantization step is 0-52, and then the value of the quantization parameter is 0-51; for the encoding of image chroma, the corresponding quantization step can be 0-39, and then the value of the quantization parameter is also adjusted to 0-38.
[0054] It can be known that the smaller the value of the quantization parameter corresponding to a certain image region is, the more refined the quantization for compressing the image region is, most details in the image region are retained, the higher the image quality of the image region obtained after compression is, and the longer the code stream generated is; on the contrary, if the value of the quantization parameter is larger, some details will be lost during compression, the code rate will be relatively reduced, but the image quality will be correspondingly reduced.
[0055] In the embodiment, the first quantization parameter of the ROI and the second quantization parameter of the non-ROI can be determined through the target bit number and the position information of the ROI and the non-ROI, and then the compression encoding of the ROI and the non-ROI in the image can be realized through the determined first quantization parameter and second quantization parameter.
[0056] It should be noted that through the above description, it can be known that a higher quantization parameter needs to be assigned to the non-ROI, and a lower quantization parameter needs to be assigned to the ROI, so as to ensure that the non-ROI has a lower image quality after compression encoding, and to ensure that the ROI has a higher image quality after compression encoding. However, in order to ensure the overall visual effect of the image after compression encoding, the embodiment considers controlling the display effect of the non-ROI and the ROI after compression encoding in a suitable range, that is, the difference between the quantization parameters of the non-ROI and the ROI needs to be controlled in a suitable range, and the quantization parameter setting of the non-ROI and the ROI can be considered based on this in the embodiment.
[0057] Specifically, for the determination process of the quantization parameters corresponding to the ROI and the non-ROI, for example, the first quantization parameter and the second quantization parameter can be dynamically adjusted based on the target bit number and the position information of the ROI and the non-ROI, combined with the quantization parameter difference of the first quantization parameter and the second quantization parameter, and finally the quantization parameter pair satisfying the set condition of the quantization parameter difference is found as the final quantization parameters of the ROI and the non-ROI.
[0058] The process of dynamically adjusting the first quantization parameter and the second quantization parameter can be: first, setting an initial ROI quantization parameter for the ROI, determining the first number of bits required to encode the ROI through the position information of the ROI, combining the known target number of bits to determine the second number of bits required to encode the non-ROI, then the current quantization parameter of the non-ROI can be back calculated through the second number of bits and the position information of the non-ROI, then the quantization parameter difference between the ROI quantization parameter and the non-ROI quantization parameter at this time can be determined, if the quantization parameter difference does not meet the set condition, the ROI quantization parameter can be adjusted, and the non-ROI quantization parameter can be determined again through the above operation after the adjustment of the ROI quantization parameter, and the quantization parameter difference is judged, and the above operation is repeated until the ROI quantization parameter and the non-ROI quantization parameter that meet the set condition are finally determined, the ROI quantization parameter can be recorded as the first quantization parameter, and the non-ROI quantization parameter can be recorded as the second quantization parameter.
[0059] S140, according to the first quantization parameter and the second quantization parameter, compressively encoding the image to obtain a target video.
[0060] It should be noted that the quantization parameter is an effective parameter for image compression encoding, in the embodiment, the image is divided into ROI and non-ROI through the above operation, after obtaining the first quantization parameter and the second quantization parameter, the ROI and the non-ROI in the image can be respectively compressed and encoded through the given compression encoding execution logic combined with the first quantization parameter and the second quantization parameter, thereby realizing the compression encoding of the image, and finally combining the compressed image can form a target video, which is a video after compression processing of the video to be processed.
[0061] The video processing method provided in the embodiment can effectively predict the ROI and the non-ROI of the image from the image of the video, and through the ROI and the non-ROI, the quantization parameters of the ROI and the non-ROI during compression encoding can be further effectively determined, thereby realizing the differential compression encoding of the image. By using the above method, the ROI and the non-ROI can be accurately determined, and through the compression encoding based on the second quantization parameter, the image quality of the non-ROI can be effectively reduced, thereby greatly reducing the code rate and saving the bandwidth flow during video transmission, and through the compression encoding based on the first quantization parameter, the image details of the ROI can be better preserved, thereby greatly improving the subjective visual effect of the video and effectively improving the user viewing experience.
[0062] In some embodiments, the image in the video to be processed can be divided into blocks to obtain a plurality of block images, which are specifically:
[0063] a1) determining the field of view angle of the video to be processed, and obtaining the division row number and the division column number corresponding to the field of view angle.
[0064] Generally, a video can be obtained in the form of capturing images by an image capturing device, and different field angles can be used when capturing images, such as a 360-degree panoramic field angle or a 180-degree half-panoramic field angle. In this embodiment, the video to be processed can be directly captured by the image capturing device, or can be a video processed by optimizing the video captured by the image capturing device. When the image capturing device captures images to form a video using different field angles, the video to be processed associated with the video also has different field angles.
[0065] In this embodiment, the association between the field angle of the video and the image block division can be established in advance when the image in the video to be processed is divided into blocks. For example, the number of division rows and the number of division columns used when dividing blocks of images with different field angles can be different, and the larger the field angle, the larger the values of the number of division rows and the number of division columns.
[0066] Therefore, this step can first obtain the field angle of the video to be processed, and then determine the number of division rows and the number of division columns used for image block division according to the direct association between the field angle and the image block division established in advance. For example, when the video to be processed is captured with a 360-degree panoramic field angle, 3 and 4 can be used as the number of division rows and the number of division columns, respectively; when the video to be processed is captured with a 180-degree half-panoramic field angle, 2 and 3 can be used as the number of division rows and the number of division columns, respectively.
[0067] b1) dividing the image into blocks according to the number of division rows and the number of division columns to obtain a plurality of block images.
[0068] After obtaining the number of division rows and the number of division columns through the above steps, the image can be divided into blocks according to the number of division rows m and the number of division columns n, thereby forming m*n rectangular block images. The size of each block image can be the same.
[0069] In the implementation of the above embodiment, the relationship between the field angle and the block division is considered, that is, the larger the field angle, the more image content contained in the corresponding image, and the more division numbers required for block division. Through this division method, effective division of block images is realized, and effective determination of ROI and non-ROI in subsequent execution is ensured.
[0070] In some embodiments, determining the ROI and the non-ROI of the image based on the image and the plurality of block images can be specifically implemented as:
[0071] a2) determining a first heat map of the image which is the same size as the image.
[0072] The step considers ROI detection on the whole image, and a first heat map can be used to represent the detection result after ROI detection on the whole image. The first heat map can be determined according to a trained ROI detection network model.
[0073] In the embodiment, to facilitate effective fusion of the ROI detection result (first heat map) of the whole image and the ROI detection result (second heat map) of each block image, the first heat map and the second heat map are preferably the same size as the image.
[0074] In the implementation of the embodiment, the first heat map can be determined according to a trained ROI detection network model. For example, determining the first heat map of the image which is the same size as the image can be embodied as: inputting a down-sampled image obtained by down-sampling the image into the ROI detection network model; and performing up-sampling processing on the heat map output by the ROI detection network model to obtain the first heat map which is the same size as the image.
[0075] The down-sampling of the image can be used to obtain an input image suitable for inputting into the ROI detection network model. The up-sampling processing of the output heat map can be used to obtain the first heat map which is the same size as the image.
[0076] In the embodiment, the ROI detection network model can preferably include a three-dimensional convolution layer or a two-dimensional convolution layer and a sub-cycle network model. It should be noted that for the ROI detection network model including the three-dimensional convolution layer, when performing ROI detection on a frame image in the video to be processed, the frame image is used as input data, and information of adjacent frame images before and after the frame image is also used as input. For the ROI detection network model including the two-dimensional convolution layer and the sub-cycle network model, when performing ROI detection on a frame image in the video to be processed, the frame image has hidden state information of the corresponding cycle network model in the ROI detection network model, so as to realize effective ROI detection.
[0077] b2) determining a second heat map of the block image which is the same size as the image.
[0078] In the embodiment, for the plurality of block images divided by the step, a second heat map which is the same size as the image can be determined.
[0079] In the implementation of the embodiment, the second heat map can also be determined according to a trained ROI detection network model. For example, the embodiment can embody determining the second heat map of the block image which is the same size as the image as:
[0080] The block image is input into the ROI detection network model respectively to obtain a block heat map of each block image; and the block heat maps are spliced to obtain a second heat map with the same size as the image.
[0081] As can be seen from the above description, in the determination of the second heat map, each block image can be input into the ROI detection network model, so that a block heat map can be output for each block image; and then the block heat maps can be spliced according to the division position of the block image in the image, and finally a second heat map with the same size as the image can be obtained.
[0082] It should be noted that the size of each block image formed after the block division is suitable for the size of the image that can be input into the ROI detection network model, so that no upsampling processing is required before input. Correspondingly, no downsampling processing is required for the output result.
[0083] In addition, for each block image, when the ROI detection is performed by using the ROI detection network model including the three-dimensional convolution layer, in addition to the block image as the input data, the block image at the same position in the front and rear frames of the image is also required as the input of the ROI detection network model; similarly, when the ROI detection is performed by using the ROI detection network model including the two-dimensional convolution layer and the sub-cycle network model, the block image also has the hidden state information of the corresponding cycle network model in the ROI detection network model, so as to realize the effective detection of the ROI.
[0084] c2) performing pixel-by-pixel multiplication on the first heat map and the second heat map to obtain a target heat map of the image.
[0085] The first heat map and the second heat map determined by the above steps are both the same size as the image, so that they can be directly subjected to pixel-by-pixel multiplication operation by the present step to realize image fusion, so as to obtain the target heat map of the image after fusion.
[0086] d2) determining the ROI and the non-ROI of the image based on the target heat map and a preset threshold.
[0087] In the present embodiment, each pixel in the target heat map is represented by a gray value, and the greater the gray value, the higher the brightness of the pixel, and the smaller the gray value, the smaller the brightness of the pixel. The pixels with brightness higher than a certain value can constitute the ROI.
[0088] In the embodiment, the preset threshold can be a gray critical value. After the gray values of the pixels in the target heat map are obtained, the gray values of the pixels are compared with the preset threshold, so that the pixels with the gray values greater than or equal to the preset threshold are determined, and finally the ROI of the image is formed based on the pixels meeting the above condition, and the area of the image other than the ROI is regarded as the non-ROI.
[0089] The preset threshold is the product of the highest gray value of the pixels in the target heat map and a preset percentage. The preset percentage can be an empirical value. For example, the preset percentage can be 10%.
[0090] In the implementation of the embodiment, the determination of the ROI and the non-ROI of the image based on the target heat map and the preset threshold can be specifically implemented as follows:
[0091] d21) determining the pixels with the gray values greater than or equal to the preset threshold in the target heat map as the ROI pixels; otherwise, determining as the non-ROI pixels.
[0092] In this step, the determination of the ROI pixels and the non-ROI pixels is realized by comparing the pixel gray values with the preset threshold.
[0093] d22) determining the ROI minimum circumscribed rectangle according to the pixel coordinates of the ROI pixels.
[0094] d23) determining the ROI minimum circumscribed rectangle as the ROI of the image.
[0095] d24) determining the area of the image other than the ROI as the non-ROI of the image.
[0096] It can be known that the determined ROI pixels can not be concentrated in the image, but can be dispersed, that is, the image can include multiple ROIs, and the ROIs are not connected. The embodiment can find the ROIs included in the image through the operation of determining the ROI minimum circumscribed rectangle.
[0097] For example, the image can include multiple ROIs, and the areas of the ROIs are not the same. The embodiment can perform secondary screening on the ROIs, and only keep the ROIs with the areas greater than a set value, so as to further optimize the determination of the ROIs.
[0098] Correspondingly, after the ROIs are determined, the area of the image other than the ROIs can be the non-ROI.
[0099] The implementation of the above embodiment gives a specific description of the determination of the ROI and the non-ROI in the image, thereby achieving effective determination of the ROI and the non-ROI in the image, and embodying differentiated determination of the ROI and the non-ROI between different images. The problem of excessively large ROI area in the image is avoided, and the problem of the determined ROI not being the actual region of interest of the user is also avoided, thereby providing basic information for subsequent determination of quantization parameters for compression encoding.
[0100] In the implementation of some embodiments, the determination of the first quantization parameter of the ROI and the second quantization parameter of the non-ROI according to the preset target bit number, the position information of the ROI, and the position information of the non-ROI can be specifically implemented as:
[0101] a4) According to the position information of the ROI and the position information of the non-ROI, the first encoding complexity of the ROI and the second encoding complexity of the non-ROI are counted.
[0102] In the present embodiment, the position information of the ROI can be understood as the position information of each pixel point in the ROI, or the position information of the smallest circumscribed rectangle constituting the ROI. The position information of the non-ROI can be understood as the position information of each pixel point in the non-ROI, or the position information of the positions other than the ROI in the image.
[0103] The encoding complexity can be used to describe the relationship between the encoding amount and the size of the problem to be solved. In the present embodiment, the encoding complexity can be understood as the complexity of the compression encoding processing of the ROI. Assuming that n is the size of the problem to be solved in the compression encoding of the ROI, the encoding complexity can be represented as C(n). Similarly, assuming that m is the size of the problem to be solved in the encoding of the non-ROI, the encoding complexity can be represented as C(m).
[0104] In the present embodiment, the size of the problem to be solved can be determined by the position information of the ROI, that is, the encoding complexity of the ROI can be determined by the position information of the ROI, and is recorded as the first encoding complexity. In the present embodiment, the size of the problem to be solved can also be determined by the position information of the non-ROI, and thus the encoding complexity of the non-ROI can also be determined by the position information of the non-ROI, and is recorded as the second encoding complexity.
[0105] b4) According to the preset target bit number, the first encoding complexity, and the second encoding complexity, the first quantization parameter of the ROI and the second quantization parameter of the non-ROI are determined.
[0106] As can be seen from the above description, the target bit number is the bit number expected to be used for compression encoding of the image, the bit number expected to be required for transmission of the image, in combination with the first encoding complexity and the second encoding complexity determined above, in combination with the set condition expected to be met by the difference between the quantization parameters of the ROI and the non-ROI in the image during compression encoding, and the initial quantization parameter set for the ROI, the dynamic adjustment of the first quantization parameter and the second quantization parameter can be realized, and finally the quantization parameter pair in which the difference between the quantization parameters meets the set condition can be obtained.
[0107] In the implementation of the above embodiment, the determination process of the first quantization parameter and the second quantization parameter is specifically given. It can be seen that the determination of the first quantization parameter and the second quantization parameter mainly involves the ROI and the non-ROI determined above, in addition to the determination logic setting of the first quantization parameter and the second quantization parameter. Through the effective determination of the ROI and the non-ROI in combination with the appropriate quantization parameter execution logic setting, the effectiveness of the determined first quantization parameter and the second quantization parameter is better ensured, and the effective execution of the subsequent compression encoding is provided with basic prerequisite information.
[0108] As one of the implementation manners of the above step b4), in the embodiment, the determination of the first quantization parameter of the ROI and the second quantization parameter of the non-ROI according to the preset target bit number, the first encoding complexity, and the second encoding complexity can be realized by the following steps, which give the specific implementation of the dynamic adjustment of the first quantization parameter and the second quantization parameter:
[0109] b41) obtaining a current first quantization parameter set for the ROI.
[0110] This step realizes the initial setting of the first quantization parameter corresponding to the ROI, and takes the initial set first quantization parameter as the first current first quantization parameter in the loop execution. The initial first quantization parameter can be a relatively small quantization value.
[0111] b42) determining a first bit number required for encoding the ROI according to the current first quantization parameter and the first encoding complexity.
[0112] In the embodiment, through the logical conversion relationship among the quantization parameter, the encoding complexity, and the bit number, the first bit number required for encoding the ROI can be determined after the current first quantization parameter and the first encoding complexity are known.
[0113] b43) determining a second bit number required for encoding the non-ROI as the difference between the target bit number and the first bit number.
[0114] In the embodiment, the second number of bits required for compressively encoding the non-ROI in the image can be determined by the difference between the target number of bits and the first number of bits.
[0115] b44) determining a current second quantization parameter of the non-ROI based on the second number of bits and the second encoding complexity.
[0116] Based on the logical conversion relationship among the quantization parameter, the encoding complexity and the number of bits as described above, the second quantization parameter of the non-ROI can be determined based on the determined second number of bits and the second encoding complexity, which can not be the final quantization parameter corresponding to the non-ROI, and can be referred to as the current second quantization parameter in the embodiment.
[0117] b45) if the difference between the current second quantization parameter and the current first quantization parameter is greater than a set threshold value, adjusting the current first quantization parameter and returning to step b42); otherwise, performing step b46).
[0118] In the embodiment, the threshold value for the loop ending condition is set by considering the absolute value of the difference between the quantization parameters of the ROI and the non-ROI. The set threshold value can be regarded as the maximum absolute difference of the quantization parameters allowed between the ROI and the non-ROI.
[0119] In the step, if the difference between the current second quantization parameter and the current first quantization parameter is greater than the set threshold value, the maximum absolute difference allowed has not been reached, and thus the current first quantization parameter of the ROI needs to be further adjusted and the step b42) needs to be returned for a new round of loop; otherwise, the difference between the current second quantization parameter and the current first quantization parameter satisfies the maximum absolute difference allowed, and thus the step b46) can be performed to end the loop.
[0120] In the step, the quantization value of the first quantization parameter can be increased by a set step to form a new current first quantization parameter.
[0121] b46) taking the current first quantization parameter and the current second quantization parameter as the first quantization parameter of the ROI and the second quantization parameter of the non-ROI, respectively.
[0122] After the condition for ending the loop is met, the final quantization values of the ROI and the non-ROI can be determined by the step, and can be referred to as the first quantization parameter and the second quantization parameter, respectively.
[0123] The above description of the cycle logic of the embodiment better illustrates the specific implementation of the embodiment considering controlling the display effect of the compressed and encoded ROI and non-ROI within a proper range. When the display effect of the ROI and non-ROI after compression and encoding is within a proper range, the image can be guaranteed to save code streams during transmission while also ensuring the best visual effect of the image.
[0124] Similarly, it should be noted that for the above technical implementation of the embodiment, before the image in the to-be-processed video is divided into blocks to obtain a plurality of block images, the above technical implementation can further include: when the image in the to-be-processed video is a binocular image, the binocular image is segmented, and each monocular image obtained after segmentation is the image.
[0125] In the embodiment, the image capturing device for capturing images to form the to-be-processed video can be a monocular image capturing device or a binocular image capturing device. When a binocular image capturing device is used to capture images, the captured image is a binocular image containing a left-eye image and a right-eye image, and the left-eye image and the right-eye image are two relatively independent images. For the to-be-processed video, if the images constituting the video are binocular images, the binocular images need to be segmented to obtain monocular images of the left eye and the right eye, respectively, and each monocular image can be used as an independent image in the to-be-processed video and participate in the logical execution of the video processing provided by the embodiment.
[0126] To better understand the video processing method provided by the embodiment, the embodiment exemplarily illustrates the specific determination of the target heat map in the video processing process through the following examples. Among them, Figure 2 An example implementation effect diagram of the determination of the target heat map in the video processing method provided by the embodiment of the disclosure. The example implementation effect diagram takes one frame of image in a to-be-processed video as a processing object and gives a specific processing implementation relative to the frame of image.
[0127] As Figure 2 shown, an image 20 in a to-be-processed video is given, and it is detected that the image 20 is a binocular image. Therefore, before the compression and encoding processing of the image is performed, the image 20 needs to be segmented to form two monocular images 201. The specific execution process of one monocular image 201 is taken as an example for illustration.
[0128] According to the method provided in the embodiment, the monocular image 201 is first divided into two branches. One branch is to divide the monocular image into blocks to obtain a certain number of block images, and the specific number is 12. Then each block image is input into an ROI detection network model to obtain a corresponding block heat map. The block heat maps can be spliced into a second heat map 202 according to the division position of the block image. The other branch is to first down-sample the monocular image, and then input the down-sampled image into the detection network model. Then the first heat map 203 is obtained by up-sampling the heat map output by the model. Then the first heat map 203 and the second heat map 202 can be multiplied pixel by pixel to obtain the fused target heat map 204.
[0129] According to the method provided in the embodiment, the monocular image 201 is first divided into two branches. One branch is to divide the monocular image into blocks to obtain a certain number of block images, and the specific number is 12. Then each block image is input into an ROI detection network model to obtain a corresponding block heat map. The block heat maps can be spliced into a second heat map 202 according to the division position of the block image. The other branch is to first down-sample the monocular image, and then input the down-sampled image into the detection network model. Then the first heat map 203 is obtained by up-sampling the heat map output by the model. Then the first heat map 203 and the second heat map 202 can be multiplied pixel by pixel to obtain the fused target heat map 204. Figure 2 It can be found that the block images after block division have higher brightness pixel points in the block heat maps after ROI detection. By this way of respectively performing ROI detection on the block images, the ROI effective detection of all regions in the image can be ensured, and the error rate of the determined ROI error is effectively reduced. The ROI optimization is realized again by splicing the block heat maps to form the second heat map 202 and fusing the second heat map with the first heat map 203 again, so that a more suitable ROI can be screened from the image, and the effective detection of the ROI is further strengthened.
[0130] Figure 3 The structure of the video processing device provided in the embodiment of the present disclosure is shown in FIG. 3. Figure 3 As shown in FIG. 3, the device includes an image division module 310, a first determination module 320, a second determination module 340, and a graph coding module 340.
[0131] The image division module 310 is configured to divide the images in the video to be processed into blocks to obtain a plurality of block images.
[0132] The first determination module 320 is configured to determine the ROI and the non-ROI of the image based on the image and the plurality of block images.
[0133] The second determination module 340 is configured to determine the first quantization parameter of the ROI and the second quantization parameter of the non-ROI according to the preset target bit number, the position information of the ROI, and the position information of the non-ROI.
[0134] The graph coding module 340 is configured to compress and encode the image according to the first quantization parameter and the second quantization parameter to obtain a target video.
[0135] The technical scheme provided by the embodiments of the present disclosure can effectively predict the ROI and non-ROI of an image from the image of a video, and through the ROI and non-ROI, the quantization parameters of the ROI and non-ROI during compression encoding can be further effectively determined, thereby realizing the differential compression encoding of the image. By using the above method, the ROI and non-ROI can be accurately determined, and through compression encoding based on the second quantization parameter, the image quality of the non-ROI can be effectively reduced, thereby greatly reducing the code rate and saving the bandwidth flow during video transmission. In addition, through compression encoding based on the first quantization parameter, the image details of the ROI can be better preserved, thereby greatly improving the subjective visual effect of the video and effectively improving the user viewing experience.
[0136] Further, the image division module 310 can be specifically configured to:
[0137] determine a field of view angle of the video to be processed, and obtain a division row number and a division column number corresponding to the field of view angle;
[0138] perform block division on the image according to the division row number and the division column number, and obtain the plurality of block images.
[0139] Further, the first determination module 320 can specifically include:
[0140] a first determination unit configured to determine a first heat map of the image, the first heat map having the same size as the image;
[0141] a second determination unit configured to determine a second heat map of the block image, the second heat map having the same size as the image;
[0142] a target obtaining unit configured to perform pixel-by-pixel multiplication on the first heat map and the second heat map, and obtain a target heat map of the image;
[0143] a result determination unit configured to determine the ROI and the non-ROI of the image based on the target heat map and a preset threshold.
[0144] Further, the first determination unit can be specifically configured to:
[0145] input a down-sampled image obtained by performing down-sampling processing on the image to an ROI detection network model;
[0146] perform up-sampling processing on a heat map output by the ROI detection network model, and obtain a first heat map having the same size as the image.
[0147] Further, the second determination unit can be specifically configured to:
[0148] input each of the block images to an ROI detection network model, and obtain a block heat map of each of the block images.
[0149] The heat maps are spliced to obtain a second heat map with the same size as the image.
[0150] Further, the ROI detection network model comprises a three-dimensional convolution layer or a two-dimensional convolution layer and a sub-cycle network model.
[0151] Further, the result determination unit can be specifically used for:
[0152] The pixel points in the target heat map with a gray value greater than or equal to the preset threshold value are determined as ROI pixel points; otherwise, as non-ROI pixel points.
[0153] According to the pixel coordinates of each ROI pixel point, a minimum ROI bounding rectangle is determined.
[0154] The minimum ROI bounding rectangle is determined as the ROI of the image.
[0155] The region of the image except the ROI is determined as the non-ROI of the image.
[0156] The preset threshold value is the product of the highest gray value among the gray values of each pixel of the target heat map and a preset percentage.
[0157] Further, the second determination module 330 can specifically include:
[0158] The complexity determination unit is configured to count a first encoding complexity of the ROI and a second encoding complexity of the non-ROI according to the position information of the ROI and the position information of the non-ROI.
[0159] The parameter determination unit is configured to determine a first quantization parameter of the ROI and a second quantization parameter of the non-ROI according to a preset target bit number, the first encoding complexity, and the second encoding complexity.
[0160] Further, the parameter determination unit can be specifically used for:
[0161] a) obtaining a current first quantization parameter set for the ROI;
[0162] b) determining a first bit number required for encoding the ROI according to the current first quantization parameter and the first encoding complexity;
[0163] c) determining a second bit number required for encoding the non-ROI as the difference between the target bit number and the first bit number;
[0164] d) Determine the current second quantization parameter of the non-ROI based on the second number of bits and the second encoding complexity;
[0165] e) If the difference between the current second quantization parameter and the current first quantization parameter is greater than a set threshold, then adjust the current first quantization parameter and return to step b); otherwise,
[0166] f) The current first quantization parameter and the current second quantization parameter are respectively used as the first quantization parameter of the ROI and the second quantization parameter of the non-ROI.
[0167] Furthermore, the video to be processed is a spherical video after undergoing equidistant cylindrical projection processing;
[0168] The images in the video to be processed are either full-view images or partial-view images.
[0169] The video processing apparatus provided in this disclosure can execute the video processing method provided in any embodiment of this disclosure, and has the corresponding functional modules and beneficial effects for executing the method.
[0170] It is worth noting that the various units and modules included in the above-mentioned device are only divided according to functional logic, but are not limited to the above division, as long as the corresponding functions can be realized; in addition, the specific names of each functional unit are only for easy differentiation and are not used to limit the protection scope of the embodiments of this disclosure.
[0171] Figure 4 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this disclosure. Reference is made below. Figure 4 It illustrates an electronic device suitable for implementing embodiments of the present disclosure (e.g., Figure 4 The diagram below shows the structure of the terminal device or server 400. The terminal device in this embodiment may include, but is not limited to, mobile terminals such as mobile phones, laptops, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), and vehicle terminals (e.g., vehicle navigation terminals), as well as fixed terminals such as digital TVs and desktop computers. Figure 4 The electronic device shown is merely an example and should not be construed as limiting the functionality and scope of the embodiments disclosed herein.
[0172] like Figure 4As shown, the electronic device 400 can include a processing device (e.g., a central processing unit, a graphics processing unit, etc.) 401 that can perform various appropriate actions and processes according to programs stored in a read-only memory (ROM) 402 or loaded into a random access memory (RAM) 403 from a storage device 408. Various programs and data required for the operation of the electronic device 400 are also stored in the RAM 403. The processing device 401, the ROM 402, and the RAM 403 are connected to each other through a bus 404. An input / output (I / O) interface 405 is also connected to the bus 404.
[0173] Generally, the following devices can be connected to the I / O interface 405: input devices 406 including, for example, a touch screen, a touch pad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; output devices 407 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; storage devices 408 including, for example, a magnetic tape, a hard disk, etc.; and communication devices 409. The communication devices 409 can allow the electronic device 400 to communicate wirelessly or wired with other devices to exchange data. Although Figure 4 The electronic device 400 is shown with various devices, but it should be understood that not all of the shown devices are required to be implemented or present. More or fewer devices can alternatively be implemented or present.
[0174] In particular, the processes described above with reference to the flowcharts can be implemented as a computer software program according to embodiments of the present disclosure. For example, embodiments of the present disclosure include a computer program product comprising a computer program carried on a non-transitory computer-readable medium, the computer program containing program code for performing the methods illustrated by the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network through the communication devices 409, or installed from the storage devices 408, or installed from the ROM 402. When the computer program is executed by the processing device 401, the above-mentioned functions defined in the methods of embodiments of the present disclosure are performed.
[0175] The names of the messages or information exchanged between the plurality of devices in the embodiments of the present disclosure are only for illustrative purposes, and are not intended to limit the scope of the messages or information.
[0176] The electronic device provided by the embodiments of the present disclosure and the resource allocation management method of the application software provided by the above embodiments belong to the same inventive concept, and the technical details not described in detail in the present embodiments can be referred to the above embodiments, and the present embodiments have the same beneficial effects as the above embodiments.
[0177] The embodiments of the present disclosure provide a computer storage medium having a computer program stored thereon, which, when executed by a processor, implements the video processing method provided by the above embodiments.
[0178] Note that the computer readable medium described above in the present disclosure can be a computer readable signal medium or a computer readable storage medium or any combination thereof. The computer readable storage medium can be, for example and without limitation, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus or device, or any suitable combination of the foregoing. More specific examples of the computer readable storage medium can include, but are not limited to, an electrical connection having one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing. In the present disclosure, a computer readable storage medium can be any tangible medium that contains or stores a program used by or in connection with an instruction execution system, apparatus or device. In the present disclosure, a computer readable signal medium can include a computer readable program code carried in a baseband or as a part of a carrier wave, in which the computer readable program code can be used by or in connection with an instruction execution system, apparatus or device. Such a propagated computer readable signal medium can take various forms, including but not limited to electro-magnetic, optical or any suitable combination of the foregoing. The computer readable signal medium can also be any computer readable medium that can be used to carry or transmit a program for use by or in connection with an instruction execution system, apparatus or device, except for the computer readable storage media described above. The program code carried by the computer readable media can be transmitted in any suitable media, including but not limited to wire, cable, fiber optic, RF (radio frequency), or any suitable combination of the foregoing.
[0179] In some embodiments, the client, server, or both can communicate using any current known or future developed network protocol, such as HTTP (HyperText Transfer Protocol), and can be interconnected with any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include local area networks ("LAN"), wide area networks ("WAN"), the Internet, and peer-to-peer networks (e.g., ad hoc peer-to-peer networks), as well as any current known or future developed networks.
[0180] The computer readable medium described above can be included in the electronic device described above; or can exist separately from the electronic device, and can be accessed via the electronic device.
[0181] The computer readable medium described above carries one or more programs, which when executed by the electronic device, cause the electronic device to:
[0182] The computer readable medium described above carries one or more programs, when the one or more programs are executed by the electronic device, the electronic device is caused to: determine current resource usage information of each resource item on the terminal in application software running; determine target resource items that currently satisfy resource allocation early warning conditions according to the current resource usage information; and adjust resource allocation logic of the target resource items in the application software running.
[0183] Computer program code for carrying out operations of the present disclosure can be written in one or more programming languages or combinations of languages including object oriented programming languages such as Java, Smalltalk, C++ or conventional procedural programming languages such as the "C" programming language or similar programming languages. The program code can execute entirely on the user's computer, partly on the user's computer, as a stand-alone software package, partly on the user's computer and partly on a remote computer or entirely on the remote computer or server. In the latter scenario, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or the connection can be made to an external computer (for example, through the Internet using an Internet Service Provider).
[0184] The flow diagrams and the block diagrams in the drawings are illustrations of architectures, functionalities, and operations of possible implementations of systems, methods, and computer program products according to various embodiments of present disclosure. In this regard, each block in the flow diagrams or block diagrams can represent a module, a procedure, or a part of code, which comprises one or more executable instructions for implementing the specified functions. It should also be noted that, in some alternative implementations, the functions noted in the blocks can occur in a different order than that noted in the figure. For example, two blocks noted in succession can in fact be executed substantially concurrently or in the opposite order, depending on the functionality involved. It should also be noted that each block in the block diagrams and / or flow diagrams, and combinations of blocks in the block diagrams and / or flow diagrams, can be implemented by dedicated hardware-based systems that perform the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.
[0185] The units described in the embodiments of the present disclosure can be implemented by software or by hardware. In some cases, the name of a unit does not constitute a limitation on the unit itself, for example, a first obtaining unit can also be an "obtaining at least two internet protocol addresses unit".
[0186] The functionality described above in this document can be performed, at least in part, by one or more hardware logic components. For example, and without limitation, non- transitory machine-readable media can include RAM, ROM, programmable ROM (EPROM, EEPROM or flash memory), or any other storage device(s) through which program instructions can be stored and executed by a processing unit. The above described functions can be implemented as software modules or software functions using object-oriented programming techniques (e.g., C++). The software modules or functions can be stored on one or more of the respective storage devices, in the RAM, or elsewhere by a processor executing at a device as instructions stored in non-transitory machine-readable media. The software modules or functions can include one or more of the above described functionality.
[0187] In the context of the present disclosure, a machine-readable medium can be a tangible medium that contains or stores a program for use by or in connection with an instruction execution system, apparatus, or device. The machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples of the machine-readable storage medium will include one or more of the following: a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0188] The above description is only preferred embodiments of the present disclosure and a description of principles of applied technology. It should be understood by those skilled in the art that the disclosed scope of the present disclosure is not limited to the technical solutions formed by the specific combinations of the above technical features, and also covers other technical solutions formed by the combinations of the above technical features or equivalent features without departing from the above disclosed concept. For example, the technical solutions formed by the mutual replacement of the above features and the technical features disclosed in the present disclosure (but not limited to) having similar functions.
[0189] Further, while operations are depicted in a particular order, this should not be understood as requiring the operations to be performed in the particular order shown or in sequential order, and that certain operations can be performed in parallel or in any suitably ordered order. Similarly, while specific implementation details are included in the above discussion, these should not be construed as limiting the scope of the disclosure. Certain features described in the context of separate embodiments can also be implemented together in a single embodiment. Conversely, various features described in the context of a single embodiment can also be implemented in multiple embodiments separately or in any suitable sub-combination.
[0190] Although the subject matter has been described in language specific to structural features and / or methodological acts, it is to be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or acts described above. Rather, the specific features and acts described above are disclosed as example forms of implementing the claims.
Claims
1. A method of video processing, the method comprising: The method comprises the following steps: block partitioning an image in a to-be-processed video to obtain a plurality of block images, the block images being rectangular image blocks formed by partitioning the image according to given partition row and column values; determining a ROI and a non-ROI of the image based on the image and the plurality of block images; determining a first quantization parameter of the ROI and a second quantization parameter of the non-ROI according to a preset target bit number, the ROI and position information of the non-ROI; compressively encoding the image according to the first quantization parameter and the second quantization parameter to obtain a target video; The method comprises the following steps: determining a first heat map of the image which has the same size as the image; determining a second heat map of the block image which has the same size as the image; pixel-by-pixel multiplying the first heat map and the second heat map to obtain a target heat map of the image; determining the ROI and the non-ROI of the image based on the target heat map and a preset threshold.
2. The method of claim 1, wherein, The method comprises the following steps: determining a field of view angle of the to-be-processed video to obtain a partition row number and a partition column number corresponding to the field of view angle; block partitioning the image according to the partition row number and the partition column number to obtain the plurality of block images.
3. The method of claim 1, wherein, The method comprises the following steps: inputting a down-sampled image obtained by down-sampling the image into an ROI detection network model; up-sampling a heat map output by the ROI detection network model to obtain a first heat map which has the same size as the image.
4. The method of claim 1, wherein, The method comprises the following steps: inputting each of the block images into an ROI detection network model to obtain a block heat map of each of the block images; splicing the block heat maps to obtain a second heat map which has the same size as the image.
5. The method according to claim 3 or 4, characterized in that, The ROI detection network model comprises a three-dimensional convolution layer or a two-dimensional convolution layer and a sub-cycle network model.
6. The method of claim 1, wherein, The method comprises the following steps: determining a ROI pixel point in the target heat map as a pixel point whose gray value is greater than or equal to the preset threshold; otherwise, determining a non-ROI pixel point; determining a ROI minimum circumscribed rectangle according to pixel coordinates of each of the ROI pixel points; determining the ROI minimum circumscribed rectangle as the ROI of the image; determining a region of the image other than the ROI as the non-ROI of the image; The preset threshold is a product of a highest gray value among gray values of each pixel of the target heat map and a preset percentage.
7. The method of claim 1, wherein, The method comprises the following steps: According to the position information of the ROI and the position information of the non-ROI, a first encoding complexity of the ROI and a second encoding complexity of the non-ROI are counted; According to the preset target bit number, the first encoding complexity, and the second encoding complexity, a first quantization parameter of the ROI and a second quantization parameter of the non-ROI are determined.
8. The method of claim 7, wherein, The method according to the preset target bit number, the first encoding complexity, and the second encoding complexity, to determine the first quantization parameter of the ROI and the second quantization parameter of the non-ROI, comprises: a) obtaining a current first quantization parameter set for the ROI; b) determining a first bit number required for encoding the ROI according to the current first quantization parameter and the first encoding complexity; c) determining a second bit number required for encoding the non-ROI as a difference between the target bit number and the first bit number; d) determining a current second quantization parameter of the non-ROI according to the second bit number and the second encoding complexity; e) if a difference between the current second quantization parameter and the current first quantization parameter is greater than a set threshold, adjusting the current first quantization parameter and returning to step b); otherwise, f) taking the current first quantization parameter and the current second quantization parameter as the first quantization parameter of the ROI and the second quantization parameter of the non-ROI respectively.
9. The method of claim 1, wherein, The video to be processed is a spherical video after equidistant cylindrical projection processing; The image in the video to be processed is a full-view image / partial-view image.
10. A video processing device, comprising: Comprise: An image division module is configured to divide an image in a video to be processed into blocks to obtain a plurality of block images, wherein the block images are rectangular image blocks formed by dividing the image according to given division row and column values; A first determination module is configured to determine a ROI and a non-ROI of the image based on the image and the plurality of block images; A second determination module is configured to determine a first quantization parameter of the ROI and a second quantization parameter of the non-ROI according to a preset target bit number, position information of the ROI, and position information of the non-ROI; A graph encoding module is configured to compress and encode the image according to the first quantization parameter and the second quantization parameter to obtain a target video. The first determination module comprises: A first determination unit is configured to determine a first heat map of the image which has the same size as the image; A second determination unit is configured to determine a second heat map of the block image which has the same size as the image; A target obtaining unit is configured to multiply the first heat map and the second heat map pixel by pixel to obtain a target heat map of the image; A result determination unit is configured to determine the ROI and the non-ROI of the image based on the target heat map and a preset threshold.
11. An electronic device, comprising: The electronic device comprises: One or more processors; A storage device configured to store one or more programs, When the one or more programs are executed by the one or more processors, the one or more processors implement the video processing method according to any one of claims 1-9.
12. A storage medium containing computer-executable instructions, wherein: The computer executable instructions, when executed by a computer processor, are for performing a video processing method as claimed in any one of claims 1-9.
Citation Information
Patent Citations
Video encoding method and device, equipment and storage medium
CN111918066A
Video high-speed frame difference backup system
CN112261386A
Video coding method and system
CN114567778A