Image processing method, electronic equipment and readable storage medium

Differential coding and AI-enhanced decoding methods preserve ROI clarity by applying higher rates to ROI areas, addressing blurring and enhancing image quality in image and video processing.

CN120321405APending Publication Date: 2025-07-15HUAWEI TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410060386.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-01-15
Publication Date
2025-07-15

AI Technical Summary

Technical Problem

In the prior art, the region of interest (ROI) of the image is prone to blur after decoding, especially when excessive noise suppression during filtering leads to a decrease in clarity.

Method used

By using a higher code rate for the region of interest (ROI) during the encoding process, and using a deep learning network to identify the ROI region during decoding, combined with the AI recovery network for area protection, avoiding blur caused by filtering.

Benefits of technology

The image quality of the ROI area is improved, and the blurring of the ROI area after decoding is avoided, which improves the sharpness and resolution of the image.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120321405A_ABST
    Figure CN120321405A_ABST
Patent Text Reader

Abstract

The invention provides an image processing method, electronic equipment and a readable storage medium, and relates to the technical field of image processing. In the image processing method provided by the invention, the coding device can identify the ROI region in the image, codes the ROI region by adopting a relatively high code rate, and codes the non-ROI region by adopting a relatively low code rate to obtain the corresponding coded data. And then, the encoding device sends the encoding data of the image and the region information of the ROI region to a decoding device, and the decoding device decodes the encoding data to obtain a decoded image. Then, the decoding device performs restoration processing on the decoded image on the basis of the received region information of the ROI region. According to the technical scheme, in the process of performing restoration processing on the decoded image, the ROI region in the image can be accurately determined based on the region information of the ROI region, and region protection is performed on the ROI region, so that the phenomenon that the ROI region in the image is blurred due to the filtering effect in the restoration process is avoided.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of image processing technology, and in particular, to an image processing method, an electronic device, and a readable storage medium. Background Art

[0002] Reference Figure 1 , in the process of transmitting data such as videos or pictures, in order to ensure the transmission efficiency, the sender usually encodes the video or picture to be transmitted before transmission and transmits the encoded video or picture to the receiver. Then the receiver decodes the received video or picture and restores the decoded image or video to further improve the image quality of the video or picture.

[0003] In the prior art, after receiving a video or picture, the receiver can use an artificial intelligence (AI) restoration network to restore the decoded video or picture to obtain the restored video or picture. However, the restoration filtering efficiency may cause the foreground of the image (such as regions of interest (ROIs) such as faces, animals, license plates, etc.) to become blurred. For example, when filtering, excessive noise suppression may cause a decrease in the clarity of the foreground of the image. Summary of the Invention

[0004] Some embodiments of this application provide an image processing method, an electronic device, and a readable storage medium. This application will be introduced from multiple aspects below, and the embodiments and beneficial effects of the following multiple aspects can be referred to each other.

[0005] In a first aspect, an embodiment of this application provides an image processing method for a first electronic device. The method includes: obtaining a first image to be transmitted; determining a region of interest in the first image, and encoding the first image to obtain first encoded data; sending the first encoded data and first region information corresponding to the region of interest to a second electronic device, where the first encoded data and the first region information are used to instruct the second electronic device to perform a restoration process on a second image obtained by decoding the first encoded data to obtain a third image.

[0006] The first electronic device may be the encoding device mentioned in this application. The second electronic device may be the decoding device mentioned in this application. The first region information may be the region information of the region of interest mentioned in this application, that is, ROI information.

[0007] After the decoding device determines the ROI region of the first image and encodes the first image to obtain encoded data, it can send the encoded data and the region information of the ROI region to the decoding device. After receiving the encoded data and the region information of the ROI region sent by the encoding device, the decoding device can perform restoration processing on the decoded image based on the region information of the ROI region to obtain a restored image. During the process of performing restoration processing on the decoded image, the decoding device will accurately determine the ROI region in the decoded image according to the region information of the ROI region to perform region protection on the ROI region and avoid the filtering effect during the restoration process from making the ROI region blurred.

[0008] In some embodiments, encoding the first image to obtain the first encoded data includes: encoding the region of interest in the first image using a first coding rate and encoding the non-region of interest outside the region of interest in the first image using a second coding rate to obtain the first encoded data, where the first coding rate is greater than the second coding rate.

[0009] During the encoding process, for an image region with a higher coding rate, more image data is retained during encoding, less image information is lost, it is easier to restore, and the quality of the decoded image is higher; for an image region with a lower coding rate, more image data is retained during encoding, more image information is lost, it is more difficult to restore, and the quality of the decoded image is lower. Encoding the ROI region using a higher coding rate can improve the image quality of the ROI region.

[0010] In some embodiments, the coding rate is determined based on the region type of the region, and the region type includes at least one of the following types: person region, animal region, license plate region, sky region, large area region, building region.

[0011] The importance level of each region can be determined according to the label (type) of each region, and then the coding rate weight of each region can be determined according to the importance level of each region. Then, according to the coding rate weight of each region and the base coding rate, the coding rate of each region can be determined.

[0012] In some embodiments, the third image is obtained by filtering the region of interest in the second image using a first filtering parameter and filtering the non-region of interest outside the region of interest in the second image using a second filtering parameter;

[0013] Alternatively, the third image is obtained by filtering the non-region of interest in the second image and not filtering the region of interest.

[0014] During the restoration process, filtering the ROI region in the image with a lower filtering parameter, or not filtering the ROI region in the image, can avoid blurring of the ROI region in the image caused by the filtering effect during the restoration process.

[0015] In some embodiments, the resolution of the third image is higher than that of the second image.

[0016] After performing restoration processing on the second image, the resolution of the second image can be improved, so that the clarity of the third image obtained after the restoration processing of the second image is higher than that of the second image.

[0017] In some embodiments, determining the region of interest in the first image includes: inputting the first image into a first model, and identifying the region of interest in the first image through the first model.

[0018] The first model can be the deep learning network model mentioned in this application.

[0019] In some embodiments, the first model includes at least one of the following: fully convolutional network, pyramid scene parsing network, image cascade network.

[0020] After inputting the image into the deep learning network model, the deep learning network model can perform semantic-level ROI segmentation processing on the image, identify the ROI region in the image, and improve the accuracy of ROI region recognition.

[0021] In a second aspect, an embodiment of the present application provides an image processing method for a second electronic device. The method includes: receiving the first encoded data of the first image sent by the first electronic device, and the first region information corresponding to the region of interest in the first image; decoding the first encoded data to obtain a second image; using the first image region in the second image as the region of interest to perform restoration processing on the second image to obtain a third image, where the second image region is the image region corresponding to the first region information.

[0022] After the decoding device determines the ROI region of the first image and encodes the first image to obtain encoded data, it can send the encoded data and the region information of the ROI region to the decoding device. After receiving the encoded data and the region information of the ROI region sent by the encoding device, the decoding device can perform restoration processing on the decoded image based on the region information of the ROI region to obtain the restored image. During the process of performing restoration processing on the decoded image, the decoding device will accurately determine the ROI region in the decoded image according to the region information of the ROI region to perform regional protection on the ROI region and avoid blurring of the ROI region caused by the filtering effect during the restoration process.

[0023] In some embodiments, using the first image region in the second image as the region of interest to perform restoration processing on the second image to obtain a third image includes: inputting the first region information and the second image into a second model, and through the second model, based on the first region information, performing restoration processing on the second image to obtain the third image. The resolution of the third image is higher than that of the second image.

[0024] The second model may be the AI restoration network mentioned in this application.

[0025] It can be understood that the restoration processing of the image by the AI restoration network may include filtering processing, enhancement processing, repair processing, super-resolution processing, etc. Filtering processing is used to suppress (reduce) noise in the image. Enhancement processing is used to enhance the contrast and brightness of the image. Repair processing is used to repair defective and damaged areas in the image to restore the integrity of the image. Super-resolution processing is used to increase the resolution of the image.

[0026] Through the restoration processing of the image by the AI restoration network, the resolution of the image can be increased, and during the restoration process, the AI restoration network can determine the ROI region in the image according to the ROI region information, and perform region protection on the ROI region to prevent the ROI region from being blurred after restoration due to filtering processing.

[0027] In some embodiments, corresponding to the first image being an image in a video, using the first image region in the second image as the region of interest to perform restoration processing on the second image to obtain a third image includes: inputting the first region information, the second image, and the decoded image corresponding to the fourth image in the video into the second model, and through the second model, based on the first region information and the decoded image corresponding to the fourth image, performing restoration processing on the second image to obtain the third image.

[0028] The first image may be the current frame mentioned in this application.

[0029] The decoded image corresponding to the fourth image may be the reference frame corresponding to the current frame during the video restoration stage mentioned in this application. The fourth image may be the image before decoding of the reference frame.

[0030] When performing restoration processing on the images in the video, the current frame, the reference frame, and the region information of the ROI region can be input into the AI restoration network, and through the AI restoration network, based on the reference frame and the region information of the ROI region, perform restoration processing on the current frame to obtain the restored image of the current frame. During the process of the AI restoration network performing restoration processing on the current frame, the AI restoration network can transfer the information of the reference frame to the current frame to improve the background clarity of the image, and determine the ROI region in the image based on the region information of the ROI region for region protection.

[0031] It should be noted that the reference frame input to the AI restoration network can be one reference frame or multiple reference frames. For example, two high-quality reference frames can be selected and input into the AI restoration network, and the current frame can be restored by using the dual-reference high-quality frame method to restore the image details and textures of the current frame at the original resolution.

[0032] In the video restoration stage, by introducing the regional information of the ROI region, regional protection can be carried out on the ROI region during the process of image restoration of the video, avoiding blurring of the ROI region in the video image after restoration. Moreover, it can significantly improve problems such as breathing effect, trailing effect, ringing effect, pattern, and blocking effect caused by video coding compression loss.

[0033] In some embodiments, restoring the second image to obtain a third image includes: filtering the region of interest in the second image with a first filtering parameter and filtering the non-region of interest outside the region of interest in the second image with a second filtering parameter to obtain a third image, where the first filtering parameter is less than the second filtering parameter; or, filtering the non-region of interest outside the region of interest in the second image and not filtering the region of interest to obtain a third image.

[0034] During the restoration process, filtering the ROI region in the image with a lower filtering parameter or not filtering the ROI region in the image can avoid blurring of the ROI region in the image caused by the filtering effect during the restoration process.

[0035] In some embodiments, the first encoded data includes a region of interest and a non-region of interest, where the region of interest is encoded based on a first bit rate and the non-region of interest is encoded based on a second bit rate, and the first bit rate is greater than the second bit rate.

[0036] In a third aspect, an embodiment of the present application provides an image processing method for a system including a first electronic device and a second electronic device. The method includes: the first electronic device acquires a first image to be transmitted; the first electronic device determines the region of interest in the first image and encodes the first image to obtain first encoded data; the first electronic device sends the first encoded data and the first region information corresponding to the region of interest to the second electronic device; the second electronic device decodes the first encoded data to obtain a second image; the second electronic device restores the second image using the first image region in the second image as the region of interest to obtain a third image, where the second image region is the image region corresponding to the first region information.

[0037] After the decoding device determines the ROI region of the first image and encodes the first image to obtain encoded data, it can send the encoded data and the region information of the ROI region to the decoding device. After receiving the encoded data and the region information of the ROI region sent by the encoding device, the decoding device can perform restoration processing on the decoded image based on the region information of the ROI region to obtain a restored image. During the process of performing restoration processing on the decoded image, the decoding device will accurately determine the ROI region in the decoded image according to the region information of the ROI region to perform region protection on the ROI region and avoid the filtering effect during the restoration processing from making the ROI region blurred.

[0038] In some embodiments, the second electronic device inputs the first region information and the second image into the second model, and the second model performs restoration processing on the second image based on the first region information to obtain a third image.

[0039] In some embodiments, corresponding to the first image being an image in a video, the second electronic device performs restoration processing on the second image by using the first image region in the second image as the region of interest to obtain a third image, including: the second electronic device inputs the first region information, the second image, and the decoded image corresponding to the fourth image in the video into the second model, and the second model performs restoration processing on the second image based on the first region information and the decoded image corresponding to the fourth image to obtain a third image.

[0040] In some embodiments, the second electronic device performs restoration processing on the second image to obtain a third image, including: the second electronic device uses a first filtering parameter to perform filtering processing on the region of interest in the second image, and uses a second filtering parameter to perform filtering processing on the non-region of interest outside the region of interest in the second image to obtain a third image, where the first filtering parameter is less than the second filtering parameter; or, the second electronic device performs filtering processing on the non-region of interest outside the region of interest in the second image and does not perform filtering processing on the region of interest to obtain a third image.

[0041] In some embodiments, the first electronic device encodes the first image to obtain first encoded data, including: the first electronic device encodes the region of interest in the first image by using a first bit rate, and encodes the non-region of interest outside the region of interest in the first image by using a second bit rate to obtain first encoded data, where the first bit rate is greater than the second bit rate.

[0042] Fourth aspect, an embodiment of the present application provides an electronic device, including: a memory for storing instructions executed by one or more processors of the electronic device; a processor, when the processor executes the instructions in the memory, enables the electronic device to execute the method described in any embodiment of the first aspect, the second aspect, or the third aspect of the present application. The beneficial effects that can be achieved by the fourth aspect can refer to the beneficial effects of the method provided in any embodiment of the first aspect, the second aspect, or the third aspect, which will not be elaborated here.

[0043] Fifth aspect, an embodiment of the present application provides a computer-readable storage medium, on which instructions are stored, and when the instructions are executed on a computer, the computer can execute the method described in any embodiment of the first aspect, the second aspect, or the third aspect. The beneficial effects that can be achieved by the fifth aspect can refer to the beneficial effects of the method provided in any embodiment of the first aspect, the second aspect, or the third aspect, which will not be elaborated here.

[0044] Sixth aspect, an embodiment of the present application provides a chip, the chip system includes a processing circuit and a storage medium, and computer program code is stored in the storage medium; when the computer program code is executed by the processing circuit, it implements the method described in any embodiment of the first aspect, the second aspect, or the third aspect. The beneficial effects that can be achieved by the sixth aspect can refer to the beneficial effects of the method provided in any embodiment of the first aspect, the second aspect, or the third aspect, which will not be elaborated here.

[0045] Seventh aspect, an embodiment of the present application provides a computer program product, including computer programs / instructions, and when the computer programs / instructions are executed, the computer executes the method described in any embodiment of the first aspect, the second aspect, or the third aspect. The beneficial effects that can be achieved by the seventh aspect can refer to the beneficial effects of the method provided in any embodiment of the first aspect, the second aspect, or the third aspect, which will not be elaborated here. Description of the Drawings

[0046] Figure 1 Schematic diagram of data transmission for videos or pictures provided for some embodiments;

[0047] Figure 2 Exemplary application scenario of the present application;

[0048] Figure 3 Flow example diagram of the image processing method provided by an embodiment of the present application;

[0049] Figure 4A Example of ROI region recognition of an image provided by an embodiment of the present application Figure 1 ;

[0050] Figure 4B Example diagram of image restoration provided for some embodiments;

[0051] Figure 4C An example diagram of image restoration provided by an embodiment of the present application;

[0052] Figure 5 A flowchart example of a video transmission method provided by an embodiment of the present application;

[0053] Figure 6 An example of ROI region recognition of an image provided by an embodiment of the present application Figure 2 ;

[0054] Figure 7 An example diagram of coding based on a reference frame provided by an embodiment of the present application;

[0055] Figure 8 A comparison diagram of restored images under different conditions provided by an embodiment of the present application;

[0056] Figures 9A - 9C Shows a flowchart of video transmission;

[0057] Figure 10 Shows a schematic diagram of video encoding and decoding in a security scenario;

[0058] Figure 11 Shows a block diagram of an electronic device provided by an embodiment of the present application. Detailed implementation manners

[0059] The embodiments of the present application are used to provide an image processing method. The image processing method of the present application is introduced below in combination with specific embodiments.

[0060] Figure 2 An exemplary application scenario of the present application.

[0061] Reference Figure 2 , the monitoring camera 10 can collect videos in real time, encode the collected videos, and then transmit the encoded videos to the computer 20. After receiving the videos transmitted by the monitoring camera 10, the computer 20 can decode the videos and further perform restoration processing on the decoded videos to obtain high-quality videos.

[0062] As described above, after the encoded video is restored, the foreground in the video may become blurred. To solve the above technical problems, an embodiment of the present application provides an image processing method. In this image processing method, the encoding device can determine the ROI region (such as image regions of people, animals, license plates, etc.) in the image to be transmitted and the region information of the ROI region (or referred to as "ROI information"). Then, a higher bitrate is used to encode the ROI region in the image, and a lower bitrate is used to encode the non-ROI region in the image to obtain corresponding encoded data. Then, the encoding device can send the encoded data and the region information of the ROI region to the decoding device. The decoding device decodes the encoded data to obtain a decoded image, and performs restoration processing on the decoded image based on the region information of the ROI region to obtain a restored image. During the process of performing restoration processing on the decoded image, the decoding device can accurately determine the ROI region in the decoded image according to the region information of the ROI region to perform regional protection on the ROI region and avoid the filtering effect during the restoration process from causing the ROI region to become blurred.

[0063] For example, during the restoration process, a mask can be created in the ROI region, and the ROI region is covered by the mask so that the ROI region is not filtered, thereby avoiding image blurring of the ROI region caused by the filtering process.

[0064] For another example, during the restoration process, smaller filtering parameters can be used to filter the ROI region to retain the texture and details of the ROI region and avoid over-filtering of the ROI region, resulting in blurring of the ROI region.

[0065] It can be understood that the ROI region is the image region where the ROI object is located, such as the closed image region where the contour of the ROI object is located. The ROI object can be the foreground such as people, animals, license plates, etc. in the image. The non-ROI region is the image region where the non-ROI object is located. The non-ROI object can be the background such as the sky, the earth, buildings, etc. in the image.

[0066] It can be understood that the bitrate is the number of data bits transmitted per unit time during data transmission. During the encoding process, for an image region with a higher bitrate, more image data is retained during encoding, less image information is lost, it is easier to restore, and the quality of the decoded image is higher; for an image region with a lower bitrate, more image data is retained during encoding, more image information is lost, it is more difficult to restore, and the quality of the decoded image is lower. Encoding the ROI region with a higher bitrate can improve the image quality of the ROI region. Moreover, performing regional protection on the ROI region during the restoration stage can also avoid the influence of filtering on the ROI region and ensure the image quality of the ROI region.

[0067] The technical solution of the present application will be described below in conjunction with specific embodiments.

[0068] Figure 3 It is a flowchart example of the image processing method provided by the embodiment of the present application.

[0069] Referring to Figure 3 , the method includes the following steps:

[0070] S101: The encoding device acquires the image to be transmitted.

[0071] The present application does not limit the specific forms of the encoding device and the decoding device. The encoding device may be an electronic device with encoding functions such as a surveillance camera, a vehicle-mounted camera, a mobile phone, a tablet, a computer, a camera, etc. The encoding device may be an electronic device with decoding functions such as a mobile phone, a tablet, a computer, a server, etc.

[0072] Taking a surveillance camera as an example of the encoding device, the encoding device can take pictures of the target area in real time to obtain the image of the target area to be transmitted.

[0073] S102: The encoding device determines the ROI area and the non-ROI area in the image to be transmitted.

[0074] In some embodiments, the encoding device may input the image to be transmitted into a deep learning network model, perform semantic-level ROI segmentation processing on the image through the deep learning network model, identify the ROI area and the non-ROI area in the image, and output the area information and labels of the ROI area, as well as the area information and labels of the non-ROI area.

[0075] It can be understood that the deep learning network model includes but is not limited to a fully convolutional network (FCN), a pyramid scene parsing network (PSPNet), an image cascade network (ICNet), etc.

[0076] Exemplarily, referring to Figure 4A , the image L1 collected by the encoding device includes a person and the sky. After the encoding device inputs the image L1 into the deep learning network model, the deep learning network model can identify the person area and the sky area in the image L1, and output the area information and labels of the person area, as well as the area information and labels of the sky area. Among them, the label of the person area is "person", and the label of the sky area is "sky".

[0077] S103: The encoding device encodes the ROI region using the first bit rate and encodes the non-ROI region using the second bit rate to obtain corresponding encoded data, where the first bit rate is greater than the second bit rate.

[0078] It can be understood that the foreground in an image usually has a higher requirement for image quality than the background. The ROI region is usually the region where the foreground such as people, animals, license plates, etc. is located, and the non-ROI region is the region where the background such as the sky, the earth, buildings, etc. is located. Therefore, the ROI region usually has a higher requirement for image quality and needs to have higher clarity after image decoding and restoration to retain important image information. The non-ROI region usually has a relatively low requirement for image quality, and even if it is relatively blurred after image decoding and restoration, important image information will not be lost.

[0079] Since, during the encoding process, the higher the bit rate of an image region, the more image data is retained during encoding and the less image information is lost; the lower the bit rate of an image region, the more image data is retained during encoding and the more image information is lost. In some embodiments, the encoding device can use an encoder to encode the ROI region with a higher bit rate and encode the non-ROI region with a lower bit rate to obtain corresponding encoded data.

[0080] For example, when the encoding device encodes the Figure 4A image shown, it can encode the person region in the image with a bit rate of 2 Mbps (as an example of the first bit rate) and encode the sky region in the image with a bit rate of 1.2 Mbps (as an example of the second bit rate) to obtain the encoded data of the image.

[0081] It can be understood that for an image including multiple ROI regions, different types of ROI regions correspond to different first bit rates. That is to say, when encoding the bit rates of an image including multiple ROI regions, different bit rates can be used to encode different ROI regions.

[0082] S104: The encoding device sends the encoded data of the image to be transmitted and the region information of the ROI region to the decoding device.

[0083] In some embodiments, after the encoding device encodes the image to be transmitted to obtain corresponding encoded data, the encoding device can send the encoded data and the region information of the ROI region to the decoding device.

[0084] S105: The decoding device decodes the received encoded data to obtain the decoded image.

[0085] In some embodiments, after the decoding device receives the encoded data sent by the encoding device, the decoding device may use a decoder to decode the encoded data to obtain a decoded image.

[0086] S106: The decoding device restores the decoded image based on the region information of the ROI region to obtain a restored image.

[0087] In some embodiments, the decoding device may input the decoded image and the region information of the ROI region into a trained AI restoration network, and the AI restoration network performs restoration processing on the image based on the region information of the ROI region to obtain a restored image.

[0088] It can be understood that the input of the AI restoration network is an image and the region information of the ROI region of the image, and the output is an image obtained by restoring the input image based on the region information of the ROI region. The resolution of the image output by the AI restoration network is higher than the resolution of the input image.

[0089] The training process of the AI restoration network will be described below.

[0090] In some embodiments, a large number of images containing ROI objects can be collected as a dataset for training the AI restoration network. The dataset contains low-resolution and high-resolution image pairs. An image pair is two images with the same picture but different resolutions. The low-resolution image in the image pair is used as the input image, and the high-resolution image is used as the reference image. The low-resolution image in the image pair and the region information of the ROI region of the image can be input into the AI restoration network. The AI restoration network performs restoration processing on the input low-resolution image based on the region information of the ROI region and outputs a predicted image. Then, based on the resolution difference between the reference image and the predicted image, the model parameters of the AI restoration network are adjusted. In this way, the AI restoration network is continuously trained until the AI restoration network reaches the expected effect and then the training stops.

[0091] It can be understood that the AI restoration network includes, but is not limited to, super-resolution networks (such as super resolution convolutional network (SRCNN), enhanced super-resolution generative adversarial network (ESRGAN), etc.), deblurring networks (such as deblur generative adversarial network (DeblurGAN), etc.), denoising networks (such as deep denoising convolutional neural network (DnCNN), weighted nuclear norm minimization (WNNM) network, etc.).

[0092] In the embodiments of this application, during the process of the AI restoration network restoring an image, according to the regional information of the ROI region, smaller filtering parameters can be used to filter the ROI region, or only the non-ROI region can be filtered and the ROI region is not filtered, etc., to protect the ROI region, thereby ensuring the image quality of the ROI region and avoiding blurring of the ROI region after restoration.

[0093] Exemplarily, Figure 4B The example diagram of image restoration without ROI region protection is shown. Refer to Figure 4B , when performing restoration processing on the image h1, since the face region is not protected, excessive filtering is performed on the face edge in the image h1 during restoration, resulting in blurring of the face edge in the restored image h1'.

[0094] Figure 4C The example diagram of image restoration with ROI region protection is shown. Refer to Figure 4C , when performing restoration processing on the image h1, since the face region is protected, the face region is not affected by the filtering effect during restoration, and the face region in the restored image h1” is relatively clear.

[0095] Figure 5 It is a flowchart example diagram of the video transmission method provided by the embodiments of this application.

[0096] S201; The encoding device obtains the video to be transmitted.

[0097] Taking the encoding device as a surveillance camera as an example, the encoding device can capture the target area in real time to obtain the video of the target area as the video to be transmitted.

[0098] S202: The encoding device determines the ROI regions and non-ROI regions of each frame image in the video to be transmitted.

[0099] In some embodiments, the encoding device may input each frame image in the video to be transmitted into a deep learning network model, perform semantic-level ROI segmentation processing on the image through the deep learning network model, identify the ROI regions and non-ROI regions in the image, and output the region information and labels of the ROI regions, as well as the region information and labels of the non-ROI regions. In some embodiments, the type of the display object shown in the region is used as the label. For example, if the type of the display object in the ROI region is a person, the person can be used as the label of the ROI region. If the type of the display object in the non-ROI region is the sky, the sky can be used as the label of the non-ROI region.

[0100] Exemplarily, referring to Figure 6 , after the encoding device inputs the image M1 in the video to be transmitted into the deep learning network model, the deep learning network model identifies the person region, animal region, and sky region in the image M1, and outputs the region information and labels of the person region, animal region, and sky region respectively. The person region and the animal region belong to the ROI regions, and the sky region belongs to the non-ROI region.

[0101] S203: The encoding device determines the bitrate of each region based on the importance level of each region in the video to be transmitted.

[0102] In some embodiments, the importance level of each region can be determined according to the label of each region, then the bitrate weight of each region can be determined according to the importance level of each region, and then the bitrate of each region can be determined according to the bitrate weight of each region and the base bitrate.

[0103] Exemplarily, referring to Figure 6 , assuming the base bitrate is 8 Mbps, the importance levels of each region in the image M1 from high to low are: person region, animal region, sky region. The bitrate weight of the person region can be determined to be 0.5, the bitrate weight of the animal region can be determined to be 0.35, and the bitrate weight of the sky region can be determined to be 0.15 according to the importance level of each region. In this case, the bitrate of the person region is: 0.5 * 8 = 4 Mbps. The bitrate of the animal region is 0.35 * 8 = 2.8 Mbps. The bitrate of the sky region is: 0.15 * 8 = 1.2 Mbps.

[0104] S204: The encoding device encodes each region based on the bitrate of each region to obtain encoded data.

[0105] In some embodiments, after determining the bitrates of the respective regions in each frame image of the video to be transmitted, the encoding device may use an encoder to encode each region based on the bitrate of each region to obtain corresponding encoded data.

[0106] Exemplarily, with reference to Figure 6 , after the encoding device determines that the bitrate of the human region in image M1 is 4 Mbps, the bitrate of the animal region is 2.8 Mbps, and the bitrate of the sky region is 1.2 Mbps, the encoding device may use bitrates of 4 Mbps, 2.8 Mbps, and 1.2 Mbps to encode the human region, animal region, and sky region in image M1 respectively. Thus, when encoding image M1, more image data can be retained in the human region and animal region, and less image data can be retained in the sky region. In this way, it can not only play a role in data compression but also enable the human region and animal region to have a higher image quality after decoding.

[0107] It can be understood that in a stable video scene, a high-quality reference frame can improve the image quality of subsequent frames. This is because the encoder can use the high-quality reference frame to more accurately predict the content of subsequent frames, thereby reducing the amount of data that needs to be transmitted.

[0108] In some embodiments, with reference to Figure 7 , during the process of encoding the images in the video, the encoder may select an already encoded high-quality reference frame as a reference to encode subsequent similar frames.

[0109] It can be understood that a high-quality reference frame can be a reference frame with higher clarity. For example, a reference frame with a clarity of 1280×720, or a reference frame with a clarity of 1920×1080, etc., which is not limited thereto. A reference frame is a video frame that has a reference effect when encoding a video, such as an I-frame (intra-coded frame), a P-frame (predicted-coded frame), a B-frame (bi-directionally predicted-coded frame), etc.

[0110] S205: The encoding device sends the encoded data and the region information of the ROI regions in each frame image of the video to be transmitted to the decoding device.

[0111] In some embodiments, after encoding each frame image of the video to be transmitted, the encoding device may send the encoded data of each frame image to the decoding device in the form of a bitstream. And send the region information of each region in each frame image of the video to be transmitted to the decoding device, so that the subsequent decoding device can perform restoration processing on each frame image in the video based on the region information of each region in each frame image of the video to be transmitted.

[0112] S206: The decoding device decodes the encoded data of each frame image in the video to be transmitted received, and obtains the decoded video.

[0113] After the decoding device receives the encoded data sent by the encoding device, it can use a decoder to decode the received encoded data to obtain the decoded video.

[0114] S207: The decoding device restores the decoded video to obtain the restored video.

[0115] In some embodiments, when the decoding device restores the decoded video, it can input the region information of each frame image in the video and the ROI region in the image into the AI restoration network. The AI restoration network performs restoration processing on the image based on the region information of the ROI region to obtain the restored image.

[0116] It can be understood that the restoration processing of the image by the AI restoration network may include filtering processing, enhancement processing, repair processing, super-resolution processing, etc. Filtering processing is used to suppress (reduce) noise in the image. Enhancement processing is used to enhance the contrast and brightness of the image. Repair processing is used to repair defective and damaged areas in the image to restore the integrity of the image. Super-resolution processing is used to improve the resolution of the image.

[0117] During the process of the AI restoration network restoring the image, the ROI region in the image can be determined for region protection according to the region information of the ROI region. For example, a smaller filtering parameter is used to filter the ROI region, or only the non-ROI region is filtered and the ROI region is not filtered, etc., so as to ensure the image quality of the ROI region and avoid blurring of the ROI region after restoration. It can be understood that the noise and detail information in the image are usually mixed together. When filtering the image to suppress the noise in the image, some detail information of the image may be lost. Therefore, when filtering the image, the larger the filtering parameter, the more detail information of the noise is filtered out; the smaller the filtering parameter, the less detail information is filtered out. When a larger filtering parameter is used to filter the ROI region, more detail information may be lost, resulting in the edge details of the ROI region in the image being too smooth and the edge of the ROI region being blurred. While using a smaller filtering parameter to filter the ROI region can retain more detail information and avoid blurring of the edge of the ROI region. Similarly, not filtering the ROI region can also avoid blurring of the ROI region.

[0118] In some other embodiments, when the decoding device restores the decoded video, it may also input the reference frame, the current frame (the video frame that needs to be encoded currently), and the region information of the ROI region of the current frame into the AI restoration network. Based on the reference frame and the region information of the ROI region, the AI restoration network performs restoration processing on the current frame to obtain the restored image of the current frame. During the process of the AI restoration network performing restoration processing on the current frame, the AI restoration network may transfer the information of the reference frame to the current frame to enhance the clarity of the background of the image, and determine the ROI region in the image based on the region information of the ROI region for region protection.

[0119] It should be noted that the reference frame input to the AI restoration network can be one reference frame or multiple reference frames. For example, two high-quality reference frames can be selected and input into the AI restoration network, and the current frame can be restored in the manner of double-reference high-quality frames to restore the image details and textures of the current frame at the original resolution.

[0120] In the embodiments of the present application, during the process of the encoding device encoding the video, different bitrates can be set for encoding according to the importance degree of each region in each frame of the video, so as to control the image quality of each region and ensure that the display object concerned by the user has a high clarity after decoding. And, in the video restoration stage, by introducing the region information of the ROI region, during the process of image restoration of the video, the ROI region can be protected, avoiding the blurring of the ROI region in the video image after restoration. And, it can significantly improve problems such as breathing effect, trailing effect, ringing effect, pattern effect, blocking effect, etc. caused by video coding compression loss.

[0121] For example, refer to Figure 8 , Image a1 is obtained by encoding the human region in the original image at a bitrate of 1.4 Mbps in the encoding stage, decoding in the manner of double-reference frames (2Ref) in the decoding stage, and using the AI restoration network for restoration in the restoration stage. Since foreground protection is not performed in the restoration stage, the foreground edge in the restored image a1 appears blurred.

[0122] Image a2 is obtained by encoding the human region in the original image at a bitrate of 1.4 Mbps in the encoding stage, decoding in the manner of double-reference frames (2Ref) in the decoding stage, and using the AI restoration network and performing restoration in the foreground protection manner based on the region information of the ROI region in the restoration stage. Since the AI restoration network protects the foreground region of the image based on the region information of the ROI region in the restoration stage, the foreground in the restored image a2 is relatively clear.

[0123] Image a3 is obtained by encoding the original image at a bitrate of 2 Mbps and then decoding it in a single-reference-frame manner. Since it is decoded in a single-reference-frame manner and the AI restoration network is not used for image restoration after decoding, although it is encoded at a relatively high bitrate, the foreground of image a3 still appears blurred.

[0124] Based on Figure 8 It can be seen that in the video restoration stage, by introducing ROI information, when the AI restoration network restores each frame of the video, it can protect the foreground in the image, ensuring that the foreground in the restored image has high clarity.

[0125] In some embodiments, the encoding device and the decoding device can be linked to reduce computing power and improve the effect. For example, the encoding device can dynamically adjust the encoding bitrate according to the performance and available resources of the decoding device, enabling the decoding device to fully utilize its own performance and available resources during decoding to quickly decode the encoded data, and at the same time providing real-time feedback and control to the encoding device to ensure the quality and smoothness of the video.

[0126] Figures 9A - 9C The flowchart of video transmission is shown. The technical solution of the present application will be described below in conjunction with Figures 9A - 9C to illustrate the technical solution of the present application.

[0127] Referring to Figure 9A , after obtaining the video A to be transmitted, the encoding device can first perform intelligent analysis on the video A to determine the bitrate of each region in each frame of the video A. During this process, the encoding device can input the video A into the deep learning network model, perform ROI segmentation processing on the images in the video A through the deep learning network model to determine the ROI regions and non-ROI regions in the image, and input the region information and labels of each region. Then, according to the labels of each region, determine the importance of each region, and determine the bitrate of each region according to the importance of each region, so as to encode each region in the image according to the bitrate of each region during subsequent encoding.

[0128] For example, after the encoding device inputs the image M1 in video A into the deep learning network model, the deep learning network model can extract the features of the people, animals, and sky in the image M1 through multiple residual blocks. The importance levels of these features, from high to low, are: people, animals, sky. The encoding device can set the bitrate weights of the people, animals, and sky in the image M1 to W1, W2, and W3 respectively according to the importance levels of these features. Then, based on the bitrate weights of each feature, the basic bitrate is quantized to obtain the bitrates of the regions where each feature is located. For example, if the basic bitrate is Q1, the bitrate of the people region can be W1*Q1, the bitrate of the animal region can be W2*Q1, and the bitrate of the sky region can be W3*Q1, where W1 + W2 + W3 = 1.

[0129] Reference Figure 9B , after the encoding device determines the bitrates of each region in each frame image of video A, it can encode each region in the image based on the bitrate to obtain the corresponding encoded data. Then, the encoding device sends the encoded data of each frame image in the video to the decoding device in the form of a bitstream. After receiving the encoded data sent by the encoding device, the decoding device can decode the encoded data. After decoding all the encoded data, the decoded video A is obtained. In addition, the encoding device can also send the region information of the ROI region of each frame image in video A to the decoding device, so that the decoding device can perform restoration processing on the decoded video A based on the region information of the ROI region.

[0130] Reference Figure 9C , after the decoding device decodes the encoded video A, the decoding device can perform restoration processing on the decoded video through network P1 to obtain the restored video A'.

[0131] Among them, network P1 includes a current frame extraction branch, a long-term reference frame feature extraction branch, and multiple downsampling modules. The current frame extraction branch includes multiple multi-path resblocks (MP-ResBlocks). The long-term reference frame feature extraction branch includes multiple multi-path resblocks and a high-quality implicit feature alignment and fusion layer.

[0132] The multi-path resblock is used to extract image features, such as texture, edges, etc. By replacing the ordinary 3x3 convolution with 4-way parallel convolutions with different kernels and then summing the convolution results, the multi-path resblock can significantly improve the ability of ordinary convolution to extract important visual features (such as texture, edges, etc.).

[0133] The high-quality implicit feature alignment and fusion layer is used to perform implicit alignment (a method of automatically learning the correspondence and internal relationship between different feature information through a deep learning network model) and fusion on the input different feature information.

[0134] The video restoration process will be introduced below in conjunction with Figure 9C the network P1 shown.

[0135] As Figure 9C shown, when the decoding device restores a certain frame image in the video as the current frame, the decoding device can input the current frame, the high-quality reference frame, and the ROI information transmitted by the encoding device into the network P1. Through the first-layer multi-path residual block, the first three-layer multi-path residual block, and the first four-layer multi-path residual block in the previous frame extraction branch of the network P1, the feature information S1, the feature information S2, and the feature information S3 of the current frame are respectively extracted. In addition, the feature information S4 of the reference frame is extracted through the three-layer multi-path residual block in the long-term reference frame feature extraction branch of the network P1, and the feature information S4 is fused with the ROI information. Then, the fused feature information is aligned and fused with the feature information S2 through the high-quality implicit feature alignment and fusion layer to obtain the feature information S5. The feature information S5 and the feature information S3 are upsampled by the upsampling module and then concatenated with the feature information S1. Then, the concatenated feature information is upsampled by the next upsampling module, and then the current frame is restored according to the feature information obtained after the upsampling process, and the restored image of the current frame is output.

[0136] It can be understood that both the long-term reference frame feature extraction branch and the current frame extraction branch are MP-ResBlock structures with successive downsampling. The high-quality reference frame and the current frame are simultaneously input into the network P1, and the high-abstractness feature-level information is extracted through the long-term reference frame feature extraction branch and the current frame extraction branch respectively, giving full play to the characteristic that the deep neural network can obtain the optimal high-level representation of the reference frame and the current frame through learning. Compared with the widely used pixel-level information extraction and processing, the network P1 is more accurate in information extraction, migration, and fusion, and is more robust to the noise in the image.

[0137] The technical solution of this application can be applied to various scenarios of video transmission or image transmission, such as security, autonomous driving, and other scenarios. Figure 10 shows the schematic diagram of video coding and decoding in the security scenario. The technical solution of this application will be introduced below in conjunction with Figure 10 the following.

[0138] Refer to Figure 10, when the front-end IP camera (IPC) (as an encoding device) is working, the sensor of the front-end IPC can convert the incoming optical signal into a corresponding digital image signal. Then, the image signal processing (ISP) unit of the front-end IPC can perform processing such as correction, interpolation, white balance, and color correction on the digital image signal to output a high-quality digital image signal, such as a video. After the image information processor outputs a high-quality video, the front-end IPC can perform intelligent analysis on the video to perform semantic-level ROI region segmentation processing on each frame of the video to determine the ROI region and non-ROI region in the image, and determine the bitrate of each region. Among them, the bitrate of the ROI region is higher than that of the non-ROI region.

[0139] After the front-end IPC determines the bitrate of each region in each frame of the video, the front-end IPC can use an encoder to encode according to the bitrate of each region to obtain corresponding encoded data. Then, the front-end IPC transmits the encoded data of each frame of the image to the back-end network video recorder (NVR) (as a decoding device) in the form of a bitstream. After the back-end NVR receives the encoded data transmitted by the front-end IPC, it can use a decoder to decode the encoded data to obtain the decoded video. Then, when the back-end NVR restores each frame of the video through the AI restoration network deployed by the neural processing unit (NPU), the frame can be used as the current frame, and then the current frame, high-quality reference frame, and region information of the ROI region are input into the AI restoration network, and the current frame is restored by the AI restoration network based on the reference frame and region information of the ROI region to obtain the restored image. In this way, after the back-end NVR restores each frame of the video, a restored video can be obtained.

[0140] In the embodiment of the present application, deploying the AI restoration network on the back-end NVR side can reduce the distortion introduced in the encoding process and improve the clarity of the video. Moreover, when using the AI restoration network to restore the video, by introducing a high-quality reference frame and relying on the fact that most of the background in the monitoring scene is the same and the reference of the group of pictures (GOP), the background quality (such as background clarity, trailing effect, breathing effect, etc.) is significantly improved. Furthermore, by integrating the ROI information transmitted by the front-end IPC into the AI restoration network, the image quality of the ROI region can be improved while significantly improving the background quality.

[0141] Figure 11The structural schematic diagram of the electronic device 1000 is shown. The electronic device 100 can be the encoding device and decoding device mentioned in this application. As Figure 11 shown, the electronic device 1000 includes: one or more processors 1032, a communication interface 1033, a memory 1031, a bus system 1034, and one or more programs. Among them, the communication interface 1033, the processor 1032, and the memory 1031 are interconnected through the bus system 1034; the bus system 1034 can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of simplicity, only a thick line is used to represent it in the figure, but it does not mean that there is only one bus or one type of bus. Among them, the one or more programs are stored in the memory 1031, and the one or more programs include instructions, and when the instructions are executed by the electronic device 1000, the electronic device 1000 is caused to execute Figure 3 or Figure 5 the relevant methods in

[0142] Among them, the processor 1032 is used to execute the instructions stored in the memory 1031, and perform encoding, decoding, and restoration processing on the image to be encoded.

[0143] The processor 1032 can run one or more operating systems, such as a real-time operating system and a non-real-time operating system. The processor 1032 can execute tasks in the user mode. During the process of the processor 1032 executing tasks in the user mode, when the processor 1032 receives an interrupt request, the processor 1032 can interrupt the currently executing task and switch to the kernel mode to execute the interrupt handler corresponding to the interrupt request. After the processor 1032 interrupts the currently executing task, it can determine the task corresponding to the interrupt request based on the mapping relationship between the type of the interrupt request and the task, and after the execution of the interrupt handler is completed, execute the task corresponding to the interrupt request.

[0144] The memory 1032 is used to store data such as videos and pictures.

[0145] This application also provides a computer storage medium storing one or more programs, and the one or more programs include instructions, and when the instructions are executed by the electronic device 1000, the electronic device 1000 is caused to execute Figure 3 or Figure 5 the relevant methods in

[0146] This application also provides a computer program product containing instructions, and when the computer program product runs on the electronic device 1000, the electronic device 1000 is caused to execute Figure 3 or Figure 5 the relevant methods in

[0147] Among them, the electronic device, computer storage medium, or computer program product provided in this application is all used to execute the corresponding method provided above. Therefore, the beneficial effects it can achieve can refer to the beneficial effects in the corresponding method provided above, and will not be elaborated here.

[0148] It should be understood that in various embodiments of this application, the order numbers of the above processes do not mean the sequence of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of this application.

[0149] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this application.

[0150] Those skilled in the art can clearly understand that for the convenience and simplicity of description, the specific working processes of the systems, devices, and units described above can refer to the corresponding processes in the foregoing method embodiments, and will not be elaborated here.

[0151] In several embodiments provided in this application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are only illustrative. For example, the division of the units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection to each other can be through some interfaces, and the indirect coupling or communication connection of devices or units can be in an electrical, mechanical, or other form.

[0152] The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they can be located in one place, or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0153] In addition, in each embodiment of this application, the functional units can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit.

[0154] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using a software program, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions described in the embodiments of the present application are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from one website, computer, database server, or data center to another website, computer, database server, or data center by wire (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or wirelessly (such as infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a database server or a data center that contains one or more media integrated therein. The available medium can be a magnetic medium (such as a floppy disk, hard disk, magnetic tape), an optical medium, or a semiconductor medium (such as a solid state disk (SSD)), etc.

[0155] As described above, the above are only the specific embodiments of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present application can easily think of changes or substitutions, which should all be covered by the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.

Claims

1. An image processing method, characterized in that For a first electronic device, the method includes: Obtain a first image to be transmitted; Determine a region of interest in the first image, and encode the first image to obtain first encoded data; Send the first encoded data and first region information corresponding to the region of interest to a second electronic device, where the first encoded data and the first region information are used to instruct the second electronic device to perform a restoration process on a second image obtained by decoding the first encoded data to obtain a third image.

2. The method according to claim 1, wherein The encoding the first image to obtain first encoded data includes: Encode the region of interest in the first image using a first code rate, and encode a non-region of interest outside the region of interest in the first image using a second code rate to obtain the first encoded data, where the first code rate is greater than the second code rate.

3. The method according to claim 2, characterized in that, The code rate is determined based on a region type of the region, and the region type includes at least one of the following types: a human region, an animal region, a license plate region, a sky region, a large area region, a building region.

4. The method according to claim 1, wherein The third image is obtained by filtering the region of interest in the second image using a first filtering parameter and filtering a non-region of interest outside the region of interest in the second image using a second filtering parameter; Or, the third image is obtained by filtering the non-region of interest in the second image and not filtering the region of interest.

5. The method according to any one of claims 1 or 4, characterized in that The resolution of the third image is higher than the resolution of the second image.

6. The method according to claim 1, characterized in that The determining a region of interest in the first image includes: Input the first image into a first model, and identify the region of interest in the first image through the first model.

7. The method according to claim 6, characterized in that The first model includes at least one of the following: a fully convolutional network, a pyramid scene parsing network, an image cascade network.

8. An image processing method, characterized in that, For a second electronic device, the method includes: Receive the first encoded data of the first image sent by the first electronic device and first region information corresponding to the region of interest in the first image; Decode the first encoded data to obtain a second image; Use a first image region in the second image as the region of interest to perform a restoration process on the second image to obtain a third image, where the second image region is an image region corresponding to the first region information.

9. The method according to claim 8, wherein The using a first image region in the second image as the region of interest to perform a restoration process on the second image to obtain a third image includes: Input the first region information and the second image into a second model, and perform a restoration process on the second image based on the first region information through the second model to obtain the third image.

10. The method according to claim 8, characterized in that, Corresponding to the first image being an image in a video, the using a first image region in the second image as the region of interest to perform a restoration process on the second image to obtain a third image includes: Input the first region information, the second image, and the decoded image corresponding to the fourth image in the video into a second model. Based on the first region information and the decoded image corresponding to the fourth image, the second model performs restoration processing on the second image to obtain the third image.

11. The method according to claim 9 or 10, characterized in that, The performing restoration processing on the second image to obtain the third image includes: Using a first filtering parameter to perform filtering processing on the region of interest in the second image, and using a second filtering parameter to perform filtering processing on the non-region of interest outside the region of interest in the second image to obtain the third image, where the first filtering parameter is less than the second filtering parameter; Alternatively, perform filtering processing on the non-region of interest outside the region of interest in the second image and do not perform filtering processing on the region of interest to obtain the third image.

12. The method according to any one of claims 8 to 10, characterized in that The resolution of the third image is higher than the resolution of the second image.

13. The method according to claim 8, wherein The first encoded data includes a region of interest and a non-region of interest, where the region of interest is encoded based on a first bit rate, the non-region of interest is encoded based on a second bit rate, and the first bit rate is greater than the second bit rate.

14. An image processing method for a system including a first electronic device and a second electronic device, characterized in that, The method includes: The first electronic device acquires a first image to be transmitted; The first electronic device determines the region of interest in the first image and encodes the first image to obtain first encoded data; The first electronic device sends the first encoded data and the first region information corresponding to the region of interest to a second electronic device; The second electronic device decodes the first encoded data to obtain a second image; The second electronic device performs restoration processing on the second image using the first image region in the second image as the region of interest to obtain a third image, where the second image region is the image region corresponding to the first region information.

15. The method according to claim 14, wherein, The second electronic device performing restoration processing on the second image using the first image region in the second image as the region of interest to obtain the third image includes: The second electronic device inputs the first region information and the second image into a second model. Based on the first region information, the second model performs restoration processing on the second image to obtain the third image.

16. The method according to claim 14, characterized in that, Corresponding to the first image being an image in a video, the second electronic device performing restoration processing on the second image using the first image region in the second image as the region of interest to obtain the third image includes: The second electronic device inputs the first region information, the second image, and the decoded image corresponding to the fourth image in the video into a second model. Based on the first region information and the decoded image corresponding to the fourth image, the second model performs restoration processing on the second image to obtain the third image.

17. The method according to claim 15 or 16, characterized in that, The second electronic device performing restoration processing on the second image to obtain the third image includes: The second electronic device filters the region of interest in the second image using a first filtering parameter, and filters the non-region of interest outside the region of interest in the second image using a second filtering parameter to obtain the third image, where the first filtering parameter is less than the second filtering parameter; Alternatively, the second electronic device filters the non-region of interest outside the region of interest in the second image and does not filter the region of interest to obtain the third image.

18. The method according to claim 14, characterized in that, The first electronic device encodes the first image to obtain first encoded data, including: The first electronic device encodes the region of interest in the first image using a first coding rate, and encodes the non-region of interest outside the region of interest in the first image using a second coding rate to obtain the first encoded data, where the first coding rate is greater than the second coding rate.

19. An electronic device, characterized in that, Including: A memory for storing instructions executed by one or more processors of the electronic device; A processor, when the processor executes the instructions in the memory, enables the electronic device to execute the image processing method according to any one of claims 1 to 7 or claims 8 to 13.

20. A chip, characterized in that, The chip system includes a processing circuit and a storage medium, and computer program code is stored in the storage medium; when the computer program code is executed by the processing circuit, the image processing method according to any one of claims 1 to 7 or claims 8 to 13 is implemented.

21. A computer-readable storage medium, characterized in that, Instructions are stored on the computer-readable storage medium, and when the instructions are executed on a computer, the computer is enabled to execute the image processing method according to any one of claims 1 to 7 or claims 8 to 13.

22. A computing program product, characterized in that, Including a computer program / instructions, and when the computer program / instructions are executed, the computer is enabled to execute the image processing method according to any one of claims 1 to 7 or claims 8 to 13.