Image processing method, electronic device, and readable storage medium
By identifying and distinguishing the ROI area in the encoding device, encoding with different code rates, and performing restoration processing based on ROI information on the decoding device side, the problem of blurring the ROI area in the prior art is solved, high definition and resolution of the image are achieved, and video encoding quality is improved.
Patent Information
- Application Number
- PCT/CN2024/111499
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-01-15
- Filing Date
- 2024-08-12
- Publication Date
- 2025-07-24
AI Technical Summary
In the process of video or picture transmission, when the recipient recovers through artificial intelligence, the area of interest (such as faces, animals, license plates, etc.) may become blurred, especially during the filtering process, which leads to a decrease in clarity.
By identifying the region of interest (ROI) in the image in the encoding device, encoding the ROI area with a higher code rate, encoding the non-ROI area with a lower code rate, and restoring the ROI area information is used on the decoding device side to ensure the clarity of the ROI area and avoiding blur caused by filtering.
The image quality of the area of interest is improved, the blurring of the ROI area in the restoration process is avoided, the overall clarity and resolution of the image is improved, and the problems such as breathing and tailing effects caused by video encoding compression are improved.
Smart Images

Figure CN2024111499_24072025_PF_FP_ABST
Abstract
Description
Image processing method, electronic device, and readable storage medium
[0001] This application claims priority to the Chinese patent application filed with the China Patent Office on January 15, 2024, with application number 202410060386.9 and application name “Image processing method, electronic device and readable storage medium”, the entire contents of which are incorporated by reference into this application. Technical Field
[0002] The present application relates to the field of image processing technology, and in particular to an image processing method, an electronic device, and a readable storage medium. Background Art
[0003] Referring to Figure 1, during the transmission of data such as video or images, to ensure transmission efficiency, the sender typically encodes the video or image before transmission and transmits the encoded video or image to the receiver. The receiver then decodes the received video or image and restores the decoded image to further improve the image quality of the video or image.
[0004] In existing technologies, after receiving a video or image, the receiver can restore the decoded video or image through an artificial intelligence (AI) restoration network to obtain the restored video or image. However, the restoration filtering efficiency may cause the foreground of the image (for example, regions of interest (ROI) such as faces, animals, and license plates) to become blurred. For example, excessive noise suppression during filtering can reduce the clarity of the foreground image.
[0005] Summary of the Invention
[0006] Some embodiments of the present application provide an image processing method, an electronic device, and a readable storage medium. The present application is introduced below from multiple aspects, and the embodiments and beneficial effects of the following aspects can be referenced to each other.
[0007] In a first aspect, an embodiment of the present application provides an image processing method for a first electronic device, the method comprising: acquiring a first image to be transmitted; determining an area of interest in the first image, and encoding the first image to obtain first coded data; sending the first coded data and first area information corresponding to the area of interest to a second electronic device, wherein the first coded data and the first area information are used to instruct the second electronic device to restore the second image after decoding the first coded data to obtain a third image.
[0008] The first electronic device may be the encoding device mentioned in this application. The second electronic device may be the decoding device mentioned in this application. The first region information may be the region information of the region of interest mentioned in this application, that is, ROI information.
[0009] After the decoding device determines the ROI region of the first image and encodes the first image to obtain encoded data, it can transmit the encoded data and regional information of the ROI region to the decoding device. After receiving the encoded data and regional information of the ROI region from the encoding device, the decoding device can restore the decoded image based on the regional information of the ROI region to obtain a restored image. During the restoration process of the decoded image, the decoding device accurately determines the ROI region in the decoded image based on the regional information of the ROI region, thereby protecting the ROI region and preventing the ROI region from being blurred due to filtering effects during the restoration process.
[0010] In some embodiments, encoding a first image to obtain first encoded data includes: encoding a region of interest in the first image using a first code rate, and encoding a non-region of interest outside the region of interest in the first image using a second code rate to obtain the first encoded data, wherein the first code rate is greater than the second code rate.
[0011] During the encoding process, image regions with higher bitrates retain more image data during encoding, lose less image information, are easier to restore, and have higher decoded image quality. Image regions with lower bitrates retain more image data during encoding, lose more image information, are more difficult to restore, and have lower decoded image quality. Using a higher bitrate to encode the ROI region can improve image quality in that region.
[0012] In some implementations, the bit rate is determined based on a region type of the region, where the region type includes at least one of the following types: a person region, an animal region, a license plate region, a sky region, an earth region, and a building region.
[0013] The importance of each region can be determined according to the label (type) of each region, and then the bit rate weight of each region can be determined according to the importance of each region. Finally, the bit rate of each region can be determined according to the bit rate weight and basic bit rate of each region.
[0014] In some embodiments, the third image is obtained by filtering the region of interest in the second image using the first filtering parameters and filtering the non-region of interest in the second image other than the region of interest using the second filtering parameters;
[0015] Alternatively, the third image is obtained by performing filtering processing on the non-interest region in the second image and not performing filtering processing on the interest region.
[0016] During the restoration process, using a lower filtering parameter to filter the ROI area in the image, or not filtering the ROI area in the image, can avoid the filtering effect in the restoration process causing blurring of the ROI area in the image.
[0017] In some embodiments, the resolution of the third image is higher than the resolution of the second image.
[0018] After the second image is restored, the resolution of the second image can be improved, so that the clarity of the third image obtained after the second image restoration is higher than the clarity of the second image.
[0019] In some embodiments, determining the region of interest in the first image includes: inputting the first image into a first model, and identifying the region of interest in the first image using the first model.
[0020] The first model can be the deep learning network model mentioned in this application.
[0021] In some embodiments, the first model includes at least one of the following: a fully convolutional network, a pyramid scene parsing network, and an image cascade network.
[0022] After the image is input into the deep learning network model, the deep learning network model can perform semantic-level ROI segmentation on the image, identify the ROI area in the image, and improve the accuracy of ROI area recognition.
[0023] In a second aspect, an embodiment of the present application provides an image processing method for a second electronic device, the method comprising: receiving first coded data of a first image sent by a first electronic device, and first area information corresponding to an area of interest in the first image; decoding the first coded data to obtain a second image; restoring the second image using the first image area in the second image as the area of interest to obtain a third image, wherein the second image area is the image area corresponding to the first area information.
[0024] After the decoding device determines the ROI region of the first image and encodes the first image to obtain encoded data, it can transmit the encoded data and regional information of the ROI region to the decoding device. After receiving the encoded data and regional information of the ROI region from the encoding device, the decoding device can restore the decoded image based on the regional information of the ROI region to obtain a restored image. During the restoration process of the decoded image, the decoding device accurately determines the ROI region in the decoded image based on the regional information of the ROI region, thereby protecting the ROI region and preventing the ROI region from being blurred due to filtering effects during the restoration process.
[0025] In some embodiments, restoring the second image using the first image region in the second image as the region of interest to obtain a third image includes: inputting the first region information and the second image into a second model, and restoring the second image based on the first region information using the second model to obtain the third image. The third image has a higher resolution than the second image.
[0026] The second model can be the AI restoration network mentioned in this application.
[0027] It is understood that the image restoration processing performed by the AI restoration network may include filtering, enhancement, inpainting, and super-resolution processing. Filtering is used to suppress (reduce) noise in the image. Enhancement is used to enhance the contrast and brightness of the image. Inpainting is used to repair defects and damaged areas in the image to restore the image's integrity. Super-resolution processing is used to increase the resolution of the image.
[0028] The image resolution can be improved by restoring the image through the AI restoration network. In the restoration process, the AI restoration network can determine the ROI area in the image based on the ROI area information, and perform regional inclusion on the ROI area to avoid blurring of the ROI area after restoration due to filtering processing.
[0029] In some embodiments, corresponding to the first image being an image in a video, the first image area in the second image is used as the region of interest to restore the second image to obtain a third image, including: inputting the first region information, the second image and the decoded image corresponding to the fourth image in the video into a second model, and restoring the second image based on the first region information and the decoded image corresponding to the fourth image by the second model to obtain the third image.
[0030] The first image may be the current frame mentioned in this application.
[0031] The decoded image corresponding to the fourth image may be the reference frame corresponding to the current frame in the video restoration stage mentioned in this application. The fourth image may be the image before the reference frame is decoded.
[0032] When restoring an image in a video, the regional information of the current frame, reference frame, and ROI area can be input into the AI restoration network. The AI restoration network then restores the current frame based on the regional information of the reference frame and ROI area to obtain the restored image of the current frame. During the AI restoration process of the current frame, the AI restoration network can transfer the information of the reference frame to the current frame to improve the background clarity of the image, and determine the ROI area in the image for regional protection based on the regional information of the ROI area.
[0033] It should be noted that the reference frame input to the AI restoration network can be a single reference frame or multiple reference frames. For example, two high-quality reference frames can be selected and input into the AI restoration network, and the current frame can be restored using a dual-reference high-quality frame approach to restore the image details and texture of the current frame at the original resolution.
[0034] During the video restoration phase, by introducing the ROI region information, the ROI region can be protected during the video image restoration process, preventing blur in the ROI region after restoration. Furthermore, it can significantly improve issues such as breathing, smearing, ringing, pattern, and blocking artifacts caused by video encoding compression loss.
[0035] In some embodiments, the second image is restored to obtain a third image, including: using a first filtering parameter to filter the region of interest in the second image, and using a second filtering parameter to filter the non-region of interest outside the region of interest in the second image to obtain the third image, wherein the first filtering parameter is smaller than the second filtering parameter; or, filtering the non-region of interest outside the region of interest in the second image and not filtering the region of interest to obtain the third image.
[0036] During the restoration process, using a lower filtering parameter to filter the ROI area in the image, or not filtering the ROI area in the image, can avoid the filtering effect in the restoration process causing blurring of the ROI area in the image.
[0037] In some embodiments, the first encoded data includes a region of interest and a region of no interest, wherein the region of interest is encoded based on a first code rate, and the region of no interest is encoded based on a second code rate, and the first code rate is greater than the second code rate.
[0038] In a third aspect, an embodiment of the present application provides an image processing method for a system including a first electronic device and a second electronic device, the method including: the first electronic device obtains a first image to be transmitted; the first electronic device determines an area of interest in the first image, and encodes the first image to obtain first encoded data; the first electronic device sends the first encoded data and first area information corresponding to the area of interest to the second electronic device; the second electronic device decodes the first encoded data to obtain a second image; the second electronic device uses the first image area in the second image as the area of interest to restore the second image to obtain a third image, wherein the second image area is the image area corresponding to the first area information.
[0039] After the decoding device determines the ROI region of the first image and encodes the first image to obtain encoded data, it can transmit the encoded data and regional information of the ROI region to the decoding device. After receiving the encoded data and regional information of the ROI region from the encoding device, the decoding device can restore the decoded image based on the regional information of the ROI region to obtain a restored image. During the restoration process of the decoded image, the decoding device accurately determines the ROI region in the decoded image based on the regional information of the ROI region, thereby protecting the ROI region and preventing the ROI region from being blurred due to filtering effects during the restoration process.
[0040] In some embodiments, the second electronic device inputs the first region information and the second image into a second model, and uses the second model to restore the second image based on the first region information to obtain a third image.
[0041] In some embodiments, corresponding to the first image being an image in a video, the second electronic device uses the first image area in the second image as the area of interest to restore the second image to obtain a third image, including: the second electronic device inputs the first area information, the second image and the decoded image corresponding to the fourth image in the video into the second model, and restores the second image based on the first area information and the decoded image corresponding to the fourth image through the second model to obtain the third image.
[0042] In some embodiments, the second electronic device performs restoration processing on the second image to obtain a third image, including: the second electronic device uses a first filtering parameter to filter the area of interest in the second image, and uses a second filtering parameter to filter the non-interest area outside the area of interest in the second image to obtain the third image, wherein the first filtering parameter is less than the second filtering parameter; or, the second electronic device performs filtering processing on the non-interest area outside the area of interest in the second image and does not filter processing on the area of interest to obtain the third image.
[0043] In some embodiments, the first electronic device encodes the first image to obtain first encoded data, including: the first electronic device encodes the area of interest in the first image using a first code rate, and encodes the non-interest area outside the area of interest in the first image using a second code rate to obtain the first encoded data, wherein the first code rate is greater than the second code rate.
[0044] In a fourth aspect, embodiments of the present application provide an electronic device, comprising: a memory for storing instructions executed by one or more processors of the electronic device; and a processor, which, when executing the instructions in the memory, causes the electronic device to perform the method described in any of the first, second, or third aspects of the present application. The beneficial effects achieved in the fourth aspect can be referenced to the beneficial effects of the method provided in any of the first, second, or third aspects, and are not further elaborated here.
[0045] In a fifth aspect, embodiments of the present application provide a computer-readable storage medium having instructions stored thereon. When executed on a computer, the instructions cause the computer to perform the method described in any of the first, second, or third aspects. The beneficial effects achieved in the fifth aspect can be referenced to the beneficial effects of the method described in any of the first, second, or third aspects, and are not further elaborated here.
[0046] In a sixth aspect, embodiments of the present application provide a chip system comprising a processing circuit and a storage medium storing computer program code; the computer program code, when executed by the processing circuit, implements the method described in any of the first, second, or third aspects. The beneficial effects achieved in the sixth aspect can be referenced to the beneficial effects of the method provided in any of the first, second, or third aspects, and are not further elaborated here.
[0047] In a seventh aspect, embodiments of the present application provide a computer program product, including a computer program / instructions. When executed, the computer program / instructions cause a computer to perform the method described in any one of the first, second, or third aspects. The beneficial effects achieved in the seventh aspect can be referenced to the beneficial effects of the method provided in any one of the first, second, or third aspects, and are not further elaborated here. BRIEF DESCRIPTION OF THE DRAWINGS
[0048] FIG1 is a schematic diagram of video or picture data transmission provided by some embodiments;
[0049] FIG2 is an exemplary application scenario of the present application;
[0050] FIG3 is a flowchart illustrating an image processing method according to an embodiment of the present application;
[0051] FIG4A is an example diagram 1 of ROI region recognition of an image provided by an embodiment of the present application;
[0052] FIG4B is an example diagram of image restoration provided by some embodiments;
[0053] FIG4C is an example diagram of image restoration provided by an embodiment of the present application;
[0054] FIG5 is a flowchart illustrating a video transmission method according to an embodiment of the present application;
[0055] FIG6 is a second example diagram of ROI region recognition of an image provided by an embodiment of the present application;
[0056] FIG7 is an example diagram of reference frame coding provided by an embodiment of the present application;
[0057] FIG8 is a comparison diagram of images restored under different conditions provided by an embodiment of the present application;
[0058] 9A to 9C show a flow chart of video transmission;
[0059] FIG10 shows a schematic diagram of video encoding and decoding in a security scenario;
[0060] FIG11 shows a block diagram of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0061] The embodiment of the present application is used to provide an image processing method. The image processing method of the present application is described below with reference to specific embodiments.
[0062] FIG2 is an exemplary application scenario of the present application.
[0063] 2 , the surveillance camera 10 can capture video in real time, encode the captured video, and then transmit the encoded video to the computer 20. After receiving the video transmitted by the surveillance camera 10, the computer 20 can decode the video and further restore the decoded video to obtain high-quality video.
[0064] As mentioned above, after restoring the encoded video, the foreground in the video may become blurred. In order to solve the above technical problems, an embodiment of the present application provides an image processing method. In this image processing method, the encoding device can determine the ROI area (for example, image areas such as people, animals, license plates, etc.) in the image to be transmitted and the regional information of the ROI area (or referred to as "ROI information"). Then, a higher bit rate is used to encode the ROI area in the image, and a lower bit rate is used to encode the non-ROI area in the image to obtain corresponding encoded data. Then, the encoding device can send the encoded data and the regional information of the ROI area to the decoding device. The decoding device decodes the encoded data to obtain a decoded image, and restores the decoded image based on the regional information of the ROI area to obtain a restored image. In the process of restoring the decoded image, the decoding device can accurately determine the ROI area in the decoded image based on the regional information of the ROI area to perform regional protection on the ROI area to avoid the filtering effect during the restoration process causing the ROI area to become blurred.
[0065] For example, during the restoration process, a mask may be created in the ROI region to mask the ROI region so that the ROI region is not filtered, thereby avoiding image blurring caused by filtering in the ROI region.
[0066] For another example, during the restoration process, a smaller filtering parameter can be used to filter the ROI region to preserve the texture and details of the ROI region and avoid excessive filtering of the ROI region, which causes blurring of the ROI region.
[0067] As you can understand, the ROI region is the image region where the ROI object is located, such as the closed image region where the ROI object's outline is located. The ROI object can be a foreground object in the image, such as a person, animal, or license plate. The non-ROI region is the image region where non-ROI objects are located. Non-ROI objects can be background objects, such as the sky, ground, or buildings in the image.
[0068] As you can understand, bitrate refers to the number of bits transmitted per unit time during data transmission. During the encoding process, image regions with higher bitrates retain more image data during encoding, lose less image information, and are easier to restore, resulting in higher decoded image quality. Image regions with lower bitrates retain more image data during encoding, lose more image information, are more difficult to restore, and have lower decoded image quality. Using a higher bitrate for encoding the ROI region can improve image quality in that region. Furthermore, protecting the ROI region during the restoration phase can prevent filtering effects and ensure image quality in that region.
[0069] The technical solution of the present application is described below with reference to specific embodiments.
[0070] FIG3 is a flowchart illustrating an example of an image processing method according to an embodiment of the present application.
[0071] Referring to FIG3 , the method includes the following steps:
[0072] S101: The encoding device obtains an image to be transmitted.
[0073] This application does not limit the specific form of the encoding device and decoding device. The encoding device can be a surveillance camera, a car camera, a mobile phone, a tablet, a computer, a camera, or other electronic device with an encoding function. The encoding device can be a mobile phone, a tablet, a computer, a server, or other electronic device with a decoding function.
[0074] Taking a surveillance camera as an encoding device as an example, the encoding device can take a picture of the target area in real time to obtain an image of the target area to be transmitted.
[0075] S102: The encoding device determines a ROI region and a non-ROI region in the image to be transmitted.
[0076] In some embodiments, the encoding device can input the image to be transmitted into a deep learning network model, perform semantic-level ROI segmentation processing on the image through the deep learning network model, identify the ROI area and non-ROI area in the image, and output the area information and label of the ROI area, as well as the area information and label of the non-ROI area.
[0077] It can be understood that deep learning network models include but are not limited to fully convolutional networks (FCNs), pyramid scene parsing networks (PSPNets), image cascade networks (ICNets), etc.
[0078] For example, referring to FIG4A , the image L1 captured by the encoding device includes a person and a sky. After the encoding device inputs image L1 into the deep learning network model, the deep learning network model can identify the person region and the sky region in image L1 and output regional information and labels for the person region, as well as regional information and labels for the sky region. The label for the person region is "person," and the label for the sky region is "sky."
[0079] S103: The encoding device encodes the ROI region using a first code rate and encodes the non-ROI region using a second code rate to obtain corresponding encoded data, wherein the first code rate is greater than the second code rate.
[0080] As you can understand, the foreground in an image generally has higher image quality requirements than the background. The ROI region typically contains foreground areas such as people, animals, and license plates, while the non-ROI region contains background areas such as the sky, land, and buildings. Therefore, the ROI region generally has higher image quality requirements, requiring high clarity after image decoding and restoration to preserve important image information. Non-ROI regions generally have lower image quality requirements; even if the image is relatively blurry after decoding and restoration, important image information will not be lost.
[0081] During the encoding process, higher bitrate image regions retain more image data and lose less image information, while lower bitrate image regions retain more image data and lose more image information. In some embodiments, the encoding device may use an encoder to encode the ROI region at a higher bitrate and non-ROI regions at a lower bitrate to obtain corresponding encoded data.
[0082] For example, when encoding the image shown in Figure 4A, the encoding device can use a bit rate of 2Mbps (as an example of the first bit rate) to encode the person area in the image, and use a bit rate of 1.2Mbps (as an example of the second bit rate) to encode the sky area in the image to obtain the encoded data of the image.
[0083] It is understood that for an image including multiple ROI regions, different types of ROI regions correspond to different first bit rates. That is, when encoding a bit rate including multiple ROI regions, different bit rates can be used to encode different ROI regions.
[0084] S104: The encoding device sends the encoded data of the image to be transmitted and the region information of the ROI region to the decoding device.
[0085] In some embodiments, after the encoding device encodes the image to be transmitted to obtain corresponding encoded data, the encoding device may send the encoded data and region information of the ROI region to the decoding device.
[0086] S105: The decoding device decodes the received coded data to obtain a decoded image.
[0087] In some embodiments, after the decoding device receives the encoded data sent by the encoding device, the decoding device may use a decoder to decode the encoded data to obtain a decoded image.
[0088] S106: The decoding device restores the decoded image based on the region information of the ROI region to obtain a restored image.
[0089] In some embodiments, the decoding device can input the decoded image and the regional information of the ROI area into a trained AI restoration network, and restore the image based on the regional information of the ROI area through the AI restoration network to obtain a restored image.
[0090] It can be understood that the input of the AI restoration network is the image and the regional information of the ROI area of the image, and the output is the image restored based on the regional information of the ROI area. The resolution of the image output by the AI restoration network is higher than the resolution of the input image.
[0091] The following describes the training process of the AI restoration network.
[0092] In some embodiments, a large number of images containing ROI objects can be collected as a data set for training the AI restoration network, and the data set contains low-resolution and high-resolution image pairs. An image pair is two images with the same picture but different resolutions. The low-resolution image in the image pair is used as the input image, and the high-resolution image is used as the reference image. The image with the low resolution in the image pair and the regional information of the ROI area of the image can be input into the AI restoration network, and the AI restoration network restores the input low-resolution image based on the regional information of the ROI area and outputs a predicted image. Then, based on the resolution difference between the reference image and the predicted image, the model parameters of the AI restoration network are adjusted. In this way, the AI restoration network is continuously trained until the training is stopped after the AI restoration network achieves the expected effect.
[0093] It can be understood that AI restoration networks include but are not limited to super-resolution networks (such as super resolution convolutional network (SRCNN), enhanced super-resolution generative adversarial network (ESRGAN), etc.), deblurring networks (such as deblur generative adversarial network (DeblurGAN)), denoising networks (such as deep denoising convolutional neural network (DnCNN), weighted non-local mean (weighted nuclear norm minimization, WNNM) network, etc.), etc.
[0094] In an embodiment of the present application, during the image restoration process of the AI restoration network, the ROI area can be filtered using smaller filtering parameters based on the regional information of the ROI area, or only the non-ROI area can be filtered without filtering the ROI area, etc., so as to protect the ROI area, thereby ensuring the image quality of the ROI area and avoiding blurring of the ROI area after restoration.
[0095] For example, Figure 4B shows an example of image restoration without ROI protection. Referring to Figure 4B , when restoring image h1, the lack of ROI protection for the face region results in over-filtering of the face edges in image h1, resulting in blurred face edges in the restored image h1'.
[0096] Figure 4C shows an example of image restoration using ROI protection. Referring to Figure 4C , when restoring image h1, the facial region is protected, preventing it from being affected by the filtering during restoration. As a result, the facial region in the restored image h1" is clearer.
[0097] FIG5 is a flowchart illustrating an exemplary video transmission method according to an embodiment of the present application.
[0098] S201: The encoding device obtains the video to be transmitted.
[0099] Taking the encoding device as a surveillance camera as an example, the encoding device can shoot the target area in real time to obtain a video of the target area as the video to be transmitted.
[0100] S202: The encoding device determines the ROI region and the non-ROI region of each frame image in the video to be transmitted.
[0101] In some embodiments, the encoding device can input each frame image in the video to be transmitted into a deep learning network model, perform semantic-level ROI segmentation processing on the image through the deep learning network model, identify the ROI area and non-ROI area in the image, and output the area information and label of the ROI area, as well as the area information and label of the non-ROI area. In some embodiments, the type of display object displayed in the area is used as a label. For example, if the type of display object in the ROI area is a person, the person can be used as the label of the ROI area. If the type of display object in the non-ROI area is the sky, the sky can be used as the label of the non-ROI area.
[0102] For example, referring to FIG6 , when the encoding device inputs image M1 from a video to be transmitted into the deep learning network model, the deep learning network model identifies the human, animal, and sky regions in image M1 and outputs region information and labels for the human, animal, and sky regions, respectively. The human and animal regions are considered ROI regions, while the sky region is considered a non-ROI region.
[0103] S203: The encoding device determines the bit rate of each region based on the importance of each region in the video to be transmitted.
[0104] In some embodiments, the importance of each region can be determined based on the label of each region, and then the rate weight of each region can be determined based on the importance of each region. Finally, the rate of each region can be determined based on the rate weight of each region and the basic rate.
[0105] For example, referring to Figure 6, assuming a base bitrate of 8 Mbps, the importance of the regions in image M1 is ranked from highest to lowest: person region, animal region, and sky region. Based on the importance of each region, the bitrate weight of the person region can be determined to be 0.5, the bitrate weight of the animal region to be 0.35, and the bitrate weight of the sky region to be 0.15. In this case, the bitrate of the person region is: 0.5 * 8 = 4 Mbps. The bitrate of the animal region is: 0.35 * 8 = 2.8 Mbps. The bitrate of the sky region is: 0.15 * 8 = 1.2 Mbps.
[0106] S204: The encoding device encodes each region based on the bit rate of each region to obtain encoded data.
[0107] In some embodiments, after the encoding device determines the bit rate of each region in each frame image in the video to be transmitted, the encoder can be used to encode each region based on the bit rate of each region to obtain corresponding encoded data.
[0108] For example, referring to FIG6 , after the encoding device determines that the bit rate of the human region in image M1 is 4 Mbps, the bit rate of the animal region is 2.8 Mbps, and the bit rate of the sky region is 1.2 Mbps, the encoding device can use bit rates of 4 Mbps, 2.8 Mbps, and 1.2 Mbps to encode the human region, animal region, and sky region in image M1, respectively. In this way, when encoding image M1, the human region and animal region can retain more image data, while the sky region can retain less image data. This achieves both data compression and higher image quality for the human region and animal region after decoding.
[0109] It is understandable that in a stable video scene, high-quality reference frames can improve the image quality of subsequent frames. This is because the encoder can use high-quality reference frames to more accurately predict the content of subsequent frames, thereby reducing the amount of data that needs to be transmitted.
[0110] In some embodiments, referring to FIG. 7 , when encoding an image in a video, the encoder may select an encoded high-quality reference frame as a reference to encode subsequent similar frames.
[0111] It is understood that a high-quality reference frame can be a reference frame with a higher definition. For example, a reference frame with a definition of 1280×720 or a reference frame with a definition of 1920×1080, etc., without limitation. Reference frames are video frames that serve as a reference when encoding a video, such as I-frames (intra-coded frames), P-frames (predictive coded frames), and B-frames (bidirectionally predictive coded frames).
[0112] S205: The encoding device sends the encoded data and the region information of the ROI region in each frame image of the video to be transmitted to the decoding device.
[0113] In some embodiments, after encoding each frame of the video to be transmitted, the encoding device may send the encoded data of each frame of the video to the decoding device in the form of a code stream. Furthermore, the encoding device may send the region information of each region in each frame of the video to be transmitted to the decoding device, so that the decoding device can subsequently restore each frame of the video based on the region information of each region in each frame of the video to be transmitted.
[0114] S206: The decoding device decodes the received coded data of each frame of the video to be transmitted to obtain a decoded video.
[0115] After receiving the encoded data sent by the encoding device, the decoding device can use a decoder to decode the received encoded data to obtain a decoded video.
[0116] S207: The decoding device restores the decoded video to obtain a restored video.
[0117] In some embodiments, when restoring the decoded video, the decoding device can input the regional information of each frame image in the video and the ROI area in the image into the AI restoration network, and restore the image based on the regional information of the ROI area through the AI restoration network to obtain a restored image.
[0118] It is understood that the image restoration processing performed by the AI restoration network may include filtering, enhancement, inpainting, and super-resolution processing. Filtering is used to suppress (reduce) noise in the image. Enhancement is used to enhance the contrast and brightness of the image. Inpainting is used to repair defects and damaged areas in the image to restore the image's integrity. Super-resolution processing is used to increase the resolution of the image.
[0119] During image restoration by the AI restoration network, the ROI region can be identified and protected based on the region's regional information. For example, filtering the ROI region with a smaller filter parameter, or filtering only the non-ROI region while omitting the ROI region, can ensure image quality in the ROI region and prevent blurring after restoration. It's understandable that noise and detail information in an image are often mixed together. When filtering an image to suppress noise, some detail information may be lost. Therefore, when filtering an image, a larger filter parameter removes more detail information from the noise; a smaller filter parameter removes less detail information. Using a larger filter parameter to filter the ROI region can result in a loss of detail, resulting in oversmoothing of edge details in the ROI region and blurring of the ROI edges. Using a smaller filter parameter to filter the ROI region can preserve more detail information and prevent blurring of the ROI edges. Similarly, not filtering the ROI region can also prevent blurring of the ROI region.
[0120] In other embodiments, when restoring the decoded video, the decoding device may also input the reference frame in the video, the current frame (the video frame currently to be encoded), and the regional information of the ROI area of the current frame into the AI restoration network. The AI restoration network then restores the current frame based on the regional information of the reference frame and the ROI area to obtain a restored image of the current frame. During the AI restoration network's restoration of the current frame, the AI restoration network may migrate the information of the reference frame to the current frame to improve the background clarity of the image, and determine the ROI area in the image for regional protection based on the regional information of the ROI area.
[0121] It should be noted that the reference frame input to the AI restoration network can be a single reference frame or multiple reference frames. For example, two high-quality reference frames can be selected and input into the AI restoration network, and the current frame can be restored using a dual-reference high-quality frame approach to restore the image details and texture of the current frame at the original resolution.
[0122] In an embodiment of the present application, during the encoding process of a video, the encoding device can set different bit rates for encoding according to the importance of each region in each frame of the video, thereby controlling the image quality of each region and ensuring that the display object that the user is concerned about has a high degree of clarity after decoding. In addition, during the video restoration stage, by introducing the regional information of the ROI region, the ROI region can be protected during the image restoration process of the video to avoid blurring of the ROI region in the video image after restoration. In addition, it can significantly improve the breathing effect, tailing effect, ringing effect, pattern effect, blocking effect and other problems caused by video encoding compression loss.
[0123] For example, referring to Figure 8, image a1 is obtained by encoding the original image's human figure area at a bitrate of 1.4 Mbps during the encoding phase, decoding using a dual-reference frame (2Ref) approach during the decoding phase, and restoring it using an AI restoration network during the restoration phase. Because foreground protection was not performed during the restoration phase, the foreground edges in the restored image a1 appear blurred.
[0124] Image a2 is obtained by encoding the original image's human figure area at a bitrate of 1.4Mbps, decoding it using a dual-reference frame (2Ref) method during the decoding phase, and restoring it using an AI restoration network with foreground-preserving information based on the ROI region during the restoration phase. Because the AI restoration network protects the image's foreground based on the ROI region's information during the restoration phase, the foreground in the restored image a2 is clearer.
[0125] Image a3 was obtained by encoding the original image at a 2Mbps bitrate and decoding it using a single reference frame. Because the image was decoded using a single reference frame and no AI restoration network was used after decoding, the foreground of image a3 is still blurred despite encoding at a higher bitrate.
[0126] Based on Figure 8, it can be seen that in the video restoration stage, by introducing ROI information, the AI restoration network can protect the foreground of the image when restoring each frame image in the video, which can ensure that the foreground of the restored image has high clarity.
[0127] In some embodiments, the encoding device and decoding device can be linked to reduce computing power and improve performance. For example, the encoding device can dynamically adjust the encoding bit rate based on the performance and available resources of the decoding device, allowing the decoding device to fully utilize its own performance and available resources during decoding to quickly decode the encoded data. At the same time, real-time feedback and control are provided with the encoding device to ensure video quality and smoothness.
[0128] 9A to 9C show a flow chart of video transmission. The technical solution of the present application will be described below with reference to FIG9A to 9C.
[0129] Referring to Figure 9A , after the encoding device acquires video A to be transmitted, it can first perform intelligent analysis on video A to determine the bitrate for each region within each frame of video A. During this process, the encoding device can input video A into a deep learning network model, which then performs ROI segmentation on the image within video A, determining ROI and non-ROI regions within the image and inputting region information and labels for each region. Based on the labels of each region, the importance of each region is determined, and the bitrate for each region is determined based on its importance. This allows the encoding of each region in the image to be tailored to that bitrate during subsequent encoding.
[0130] For example, after the encoding device inputs image M1 from video A into a deep learning network model, the deep learning network model can extract the features of the people, animals, and sky in image M1 through multiple residual blocks, with their importance ranked from high to low as: people, animals, and sky. The encoding device can set the bitrate weights of the people, animals, and sky in image M1 to W1, W2, and W3, respectively, based on the importance of these features. Then, the basic bitrate is quantized based on the bitrate weight of each feature to obtain the bitrate of the area where each feature is located. For example, if the basic bitrate is Q1, the bitrate of the person area can be W1*Q1, the bitrate of the animal area can be W2*Q1, and the bitrate of the sky area can be W3*Q1, where W1+W2+W3=1.
[0131] Referring to Figure 9B , after the encoding device determines the bitrate for each region in each frame of video A, it can encode each region in the image based on the bitrate to obtain corresponding encoded data. The encoding device then sends the encoded data for each frame of the video to the decoding device in the form of a code stream. After receiving the encoded data from the encoding device, the decoding device can decode the encoded data and, after decoding all the encoded data, obtain the decoded video A. Furthermore, the encoding device can also send the regional information of the ROI region of each frame of video A to the decoding device, so that the decoding device can restore the decoded video A based on the regional information of the ROI region.
[0132] 9C , after the decoding device decodes the encoded video A, the decoding device may restore the decoded video through the network P1 to obtain a restored video A′.
[0133] The network P1 includes a current frame extraction branch, a long-term reference frame feature extraction branch, and multiple downsampling modules. The current frame extraction branch includes multiple multi-path residual blocks (MP-ResBlocks). The long-term reference frame feature extraction branch includes multiple multi-path residual blocks and a high-quality implicit feature alignment and fusion layer.
[0134] The multi-path residual block is used to extract image features such as texture and edges. By replacing the ordinary 3x3 convolution with four parallel convolutions with different kernels and then summing the convolution results, the multi-path residual block can significantly improve the ability of ordinary convolution to extract important visual features (such as texture and edges).
[0135] The high-quality implicit feature alignment and fusion layer is used to implicitly align (a method that automatically learns the correspondence and intrinsic connection between different feature information through a deep learning network model) and fuse different feature information of the input.
[0136] The video restoration process is described below with reference to the network P1 shown in FIG9C .
[0137] As shown in Figure 9C , when a decoding device restores a frame in a video as the current frame, it can input the current frame, a high-quality reference frame, and the ROI information received from the encoding device into network P1. The decoding device then extracts feature information S1, S2, and S3 for the current frame using the first-layer multipath residual block, the first three-layer multipath residual block, and the first four-layer multipath residual block in the previous frame extraction branch of network P1. Furthermore, feature information S4 for the reference frame is extracted using the three-layer multipath residual block in the long-term reference frame feature extraction branch of network P1. Feature information S4 is then fused with the ROI information. This fused feature information is then aligned and fused with feature information S2 using a high-quality implicit feature alignment and fusion layer to obtain feature information S5. Feature information S5 and feature information S3 are upsampled by an upsampling module and then concatenated with feature information S1. This concatenation is then upsampled by the next upsampling module. The current frame is then restored based on the upsampled feature information, and the restored image of the current frame is output.
[0138] As you can understand, both the long-term reference frame feature extraction branch and the current frame extraction branch utilize a sequentially downsampled MP-ResBlock structure. A high-quality reference frame and current frame are simultaneously input into Network P1, where they are processed separately by the long-term reference frame feature extraction branch and the current frame extraction branch to extract high-level feature information. This leverages the deep neural network's ability to learn optimal high-level representations of the reference and current frames. Compared to widely used pixel-level information extraction and processing methods, Network P1 is more accurate in extracting, transferring, and fusing information and is more robust to image noise.
[0139] The technical solution of this application can be applied to various video transmission or image transmission scenarios, such as security, autonomous driving, etc. Figure 10 shows the principle diagram of video encoding and decoding in the security scenario. The technical solution of this application is introduced below with reference to Figure 10.
[0140] Referring to Figure 10, when the front-end network camera (IP Camera, IPC) (as an encoding device) is working, the sensor of the front-end IPC can convert the incoming light signal into a corresponding digital image signal. Then, the image signal processing (ISP) of the front-end IPC can perform correction, interpolation, white balance, color correction and other processing on the digital image signal to output a high-quality digital image signal, such as a video. After the image information processor outputs a high-quality video, the front-end IPC can perform intelligent analysis on the video to perform semantic-level ROI region segmentation processing on each frame of the video to determine the ROI region and non-ROI region in the image, and determine the bit rate of each region. Among them, the bit rate of the ROI region is higher than the bit rate of the non-ROI region.
[0141] After the front-end IPC determines the bit rate of each region in each frame of the video, the front-end IPC can use the encoder to encode according to the bit rate of each region to obtain the corresponding encoded data. Then, the front-end IPC transmits the encoded data of each frame of the image to the back-end network digital video recorder (NVR) (as a decoding device) in the form of a code stream. After the back-end NVR receives the encoded data transmitted by the front-end IPC, it can use the decoder to decode the encoded data to obtain the decoded video. Then, when the back-end NVR restores each frame of the decoded video through the AI restoration network deployed by the neural processing unit (NPU), the frame can be used as the current frame, and then the current frame, high-quality reference frame and ROI area information are input into the AI restoration network. The AI restoration network restores the current frame based on the reference frame and the ROI area information to obtain the restored image. In this way, after the back-end NVR restores each frame of the video, the restored video can be obtained.
[0142] In an embodiment of the present application, deploying an AI restoration network on the back-end NVR side can reduce the distortion introduced by the encoding process and improve the clarity of the video. In addition, when using the AI restoration network to restore the video, by introducing high-quality reference frames, with the help of the fact that most of the backgrounds of the monitoring scenes are the same and cross-group of pictures (GOP) references, the background quality (for example, background clarity, tailing effect, breathing effect, etc.) can be significantly improved. In addition, by fusing the ROI information transmitted by the front-end IPC into the AI restoration network, the image quality of the ROI area can be improved while significantly improving the background quality.
[0143] Figure 11 shows a schematic diagram of the structure of an electronic device 1000. Electronic device 100 may be the encoding device and decoding device mentioned in this application. As shown in Figure 11, the electronic device 1000 includes: one or more processors 1032, a communication interface 1033, a memory 1031, a bus system 1034, and one or more programs. The communication interface 1033, the processor 1032, and the memory 1031 are interconnected via a bus system 1034; the bus system 1034 may be a peripheral component interconnect standard bus or an extended industrial standard architecture bus. The bus can be divided into an address bus, a data bus, a control bus, etc. For ease of illustration, only one thick line is used in the figure, but this does not mean that there is only one bus or one type of bus. The one or more programs are stored in the memory 1031, and the one or more programs include instructions. When the instructions are executed by the electronic device 1000, the electronic device 1000 executes the relevant methods in Figure 3 or Figure 5.
[0144] The processor 1032 is used to execute instructions stored in the memory 1031 to encode, decode, and restore the image to be encoded.
[0145] The processor 1032 can run one or more operating systems, such as a real-time operating system and a non-real-time operating system. The processor 1032 can execute tasks in user mode. When the processor 1032 receives an interrupt request while executing a task in user mode, the processor 1032 can interrupt the currently executing task and switch to kernel mode to execute the interrupt handler corresponding to the interrupt request. After the processor 1032 interrupts the currently executing task, it can determine the task corresponding to the interrupt request based on the mapping relationship between the type of interrupt request and the task, and execute the task corresponding to the interrupt request after the interrupt handler is executed.
[0146] The memory 1032 is used to store data such as videos and pictures.
[0147] The present application also provides a computer storage medium storing one or more programs, wherein the one or more programs include instructions that, when executed by the electronic device 1000, enable the electronic device 1000 to execute the relevant method in Figure 3 or Figure 5.
[0148] The present application also provides a computer program product comprising instructions, which, when executed on the electronic device 1000, enables the electronic device 1000 to execute the relevant method in FIG. 3 or FIG. 5 .
[0149] Among them, the electronic device, computer storage medium or computer program product provided in this application is used to execute the corresponding method provided above. Therefore, the beneficial effects that can be achieved can refer to the beneficial effects in the corresponding method provided above, and will not be repeated here.
[0150] It should be understood that in the various embodiments of the present application, the size of the serial numbers of the above-mentioned processes does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present application.
[0151] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0152] Those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.
[0153] In the several embodiments provided in this application, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the device embodiments described above are merely schematic. For example, the division of the units is merely a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.
[0154] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.
[0155] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit.
[0156] In the above embodiments, all or part of the embodiments may be implemented by software, hardware, firmware, or any combination thereof. When implemented using a software program, all or part of the embodiments may be implemented in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of the present application are generated. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions may be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions may be transmitted from one website, computer, database server, or data center to another website, computer, database server, or data center via a wired (e.g., coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) method. The computer-readable storage medium may be any available medium that can be accessed by a computer or a data storage device such as a database server or data center that includes one or more media integrated therein. The available medium may be a magnetic medium (eg, a floppy disk, a hard disk, a magnetic tape), an optical medium, or a semiconductor medium (eg, a solid state disk (SSD)).
[0157] The above description is merely a specific embodiment of the present application, but the scope of protection of the present application is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in this application should be included in the scope of protection of this application. Therefore, the scope of protection of this application should be based on the scope of protection of the claims.
Claims
1. An image processing method, characterized in that, For a first electronic device, the method includes: Obtain a first image to be transmitted; Determine a region of interest in the first image, and encode the first image to obtain first encoded data; Send the first encoded data and first region information corresponding to the region of interest to a second electronic device, where the first encoded data and the first region information are used to instruct the second electronic device to perform a restoration process on a second image obtained by decoding the first encoded data to obtain a third image.
2. The method according to claim 1, characterized in that, The encoding the first image to obtain first encoded data includes: Encode the region of interest in the first image at a first coding rate, and encode a non-region of interest outside the region of interest in the first image at a second coding rate to obtain the first encoded data, where the first coding rate is greater than the second coding rate.
3. The method according to claim 2, wherein The coding rate is determined based on a region type of the region, and the region type includes at least one of the following types: a person region, an animal region, a license plate region, a sky region, a large area region, a building region.
4. The method according to claim 1, wherein The third image is obtained by filtering the region of interest in the second image with a first filtering parameter and filtering a non-region of interest outside the region of interest in the second image with a second filtering parameter; Alternatively, the third image is obtained by filtering the non-region of interest in the second image and not filtering the region of interest.
5. The method according to any one of claims 1 or 4, characterized in that, The resolution of the third image is higher than the resolution of the second image.
6. The method according to claim 1, characterized in that, The determining a region of interest in the first image includes: Input the first image into a first model, and identify the region of interest in the first image through the first model.
7. The method according to claim 6, characterized in that The first model includes at least one of the following: a fully convolutional network, a pyramid scene parsing network, an image cascade network.
8. An image processing method, characterized in that, For a second electronic device, the method includes: Receive the first encoded data of the first image sent by the first electronic device and first region information corresponding to the region of interest in the first image; Decode the first encoded data to obtain a second image; Use a first image region in the second image as the region of interest to perform a restoration process on the second image to obtain a third image, where the second image region is an image region corresponding to the first region information.
9. The method according to claim 8, characterized in that The using a first image region in the second image as the region of interest to perform a restoration process on the second image to obtain a third image includes: Input the first region information and the second image into a second model, and perform a restoration process on the second image based on the first region information through the second model to obtain the third image.
10. The method according to claim 8, wherein Corresponding to the first image being an image in a video, the using a first image region in the second image as the region of interest to perform a restoration process on the second image to obtain a third image includes: Input the first region information, the second image, and the decoded image corresponding to the fourth image in the video into a second model. Based on the first region information and the decoded image corresponding to the fourth image, the second model performs restoration processing on the second image to obtain the third image.
11. The method according to claim 9 or 10, characterized in that The performing restoration processing on the second image to obtain the third image includes: Using a first filtering parameter to perform filtering processing on the region of interest in the second image, and using a second filtering parameter to perform filtering processing on the non-region of interest outside the region of interest in the second image to obtain the third image, where the first filtering parameter is less than the second filtering parameter; Alternatively, perform filtering processing on the non-region of interest outside the region of interest in the second image and do not perform filtering processing on the region of interest to obtain the third image.
12. The method according to any one of claims 8 to 10, characterized in that, The resolution of the third image is higher than the resolution of the second image.
13. The method according to claim 8, characterized in that, The first encoded data includes a region of interest and a non-region of interest, where the region of interest is encoded based on a first coding rate, the non-region of interest is encoded based on a second coding rate, and the first coding rate is greater than the second coding rate.
14. An image processing method for a system including a first electronic device and a second electronic device, characterized in that, The method includes: The first electronic device acquires a first image to be transmitted; The first electronic device determines the region of interest in the first image and encodes the first image to obtain first encoded data; The first electronic device sends the first encoded data and the first region information corresponding to the region of interest to a second electronic device; The second electronic device decodes the first encoded data to obtain a second image; The second electronic device performs restoration processing on the second image using the first image region in the second image as the region of interest to obtain a third image, where the second image region is the image region corresponding to the first region information.
15. The method according to claim 14, wherein The second electronic device performing restoration processing on the second image using the first image region in the second image as the region of interest to obtain the third image includes: The second electronic device inputs the first region information and the second image into a second model. Based on the first region information, the second model performs restoration processing on the second image to obtain the third image.
16. The method according to claim 14, wherein, Corresponding to the first image being an image in a video, the second electronic device performing restoration processing on the second image using the first image region in the second image as the region of interest to obtain the third image includes: The second electronic device inputs the first region information, the second image, and the decoded image corresponding to the fourth image in the video into a second model. Based on the first region information and the decoded image corresponding to the fourth image, the second model performs restoration processing on the second image to obtain the third image.
17. The method according to claim 15 or 16, characterized in that, The second electronic device performing restoration processing on the second image to obtain the third image includes: The second electronic device filters the region of interest in the second image using a first filtering parameter, and filters the non-region of interest outside the region of interest in the second image using a second filtering parameter to obtain the third image, where the first filtering parameter is less than the second filtering parameter; Alternatively, the second electronic device filters the non-region of interest outside the region of interest in the second image and does not filter the region of interest to obtain the third image.
18. The method according to claim 14, wherein The first electronic device encodes the first image to obtain first encoded data, including: The first electronic device encodes the region of interest in the first image using a first coding rate, and encodes the non-region of interest outside the region of interest in the first image using a second coding rate to obtain the first encoded data, where the first coding rate is greater than the second coding rate.
19. An electronic device, characterized in that, Including: A memory for storing instructions executed by one or more processors of the electronic device; A processor, when the processor executes the instructions in the memory, enables the electronic device to execute the image processing method according to any one of claims 1 to 7 or claims 8 to 13.
20. A chip, characterized in that, The chip system includes a processing circuit and a storage medium, and computer program code is stored in the storage medium; when the computer program code is executed by the processing circuit, the image processing method according to any one of claims 1 to 7 or claims 8 to 13 is implemented.
21. A computer-readable storage medium, characterized in that, Instructions are stored on the computer-readable storage medium, and when the instructions are executed on a computer, the computer is enabled to execute the image processing method according to any one of claims 1 to 7 or claims 8 to 13.
22. A computer program product, characterized in that, Including computer program / instructions, when the computer program / instructions are executed, the computer is enabled to execute the image processing method according to any one of claims 1 to 7 or claims 8 to 13.
Citation Information
Patent Citations
Image processing method and device and electronic equipment
CN110572579A
Fast region of interest coding using multi-segment resampling
CN112655210A
Video processing device and method, and computer storage medium
CN114554212A
Video coding method and device, electronic equipment and storage medium
CN115442615A
Encoding method, encoder, storage medium and chip
CN116489364A