An image processing method, device, storage medium and electronic device

By identifying regions of interest and non-regions of interest in an image and determining mask blocks of different proportions for processing, the problems of low image compression efficiency and information loss in existing technologies are solved, achieving efficient image compression and high-quality reconstruction.

CN116366856BActive Publication Date: 2025-12-05BEIJING ZITIAO NETWORK TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202310289177.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-03-22
Publication Date
2025-12-05
Estimated Expiration
2043-03-22

AI Technical Summary

Technical Problem

Existing image compression technologies cannot effectively distinguish the characteristics of different regions of an image, resulting in limited compression efficiency and a high risk of losing important information.

Method used

By identifying regions of interest and non-regions of interest in an image, mask blocks of different proportions are determined, and masking and encoding are performed to reduce the amount of data and improve the compression ratio.

Benefits of technology

While preserving important information, the image compression rate was improved, the amount of data transmitted and bandwidth usage were reduced, and the quality of the reconstructed image was guaranteed.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116366856B_ABST
    Figure CN116366856B_ABST
Patent Text Reader

Abstract

Embodiments of the present disclosure provide an image processing method and device, a storage medium and an electronic device. The method comprises: obtaining a target image, determining a region of interest and a non-region of interest in the target image; determining mask blocks in the region of interest and the non-region of interest respectively, and performing mask processing on the mask blocks to obtain a mask image, wherein the proportion of the number of mask blocks in the region of interest is less than the proportion of the number of mask blocks in the non-region of interest; performing compression processing and encoding processing on the mask image to obtain compressed and encoded data, transmitting the compressed and encoded data to a receiving end, and the receiving end reconstructing an image based on the compressed and encoded data to obtain a reconstructed image. By performing compression processing and encoding processing on the mask image, the image compression rate is improved, and the number of mask blocks in the region of interest is reduced, thereby avoiding loss of important image information in the image compression process.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to image processing technology, and more particularly to an image processing method, apparatus, storage medium, and electronic device. Background Technology

[0002] With the continuous development of image capturing technology, the amount of image data is getting larger and larger. During the transmission of images, it is necessary to compress the images to reduce the data size.

[0003] Current compression techniques generally compress image data by performing signal conversion, quantization, and encoding using a compression codec. However, current compression techniques have limited efficiency in image compression. Furthermore, because different regions of an image have different characteristics—such as regions with detailed textures, rich features, and important features—current compression techniques cannot distinguish between different regions during the compression process. Summary of the Invention

[0004] This disclosure provides an image processing method, apparatus, storage medium, and electronic device to preserve image detail information and improve image compression rate during image compression.

[0005] In a first aspect, embodiments of this disclosure provide an image processing method, including:

[0006] Acquire a target image and determine the region of interest and non-region of interest in the target image;

[0007] Mask blocks are determined in the region of interest and the region of non-interest respectively, and the mask blocks are masked to obtain a mask image, wherein the proportion of mask blocks in the region of interest is less than the proportion of mask blocks in the region of non-interest.

[0008] The masked image is compressed and encoded to obtain compressed encoded data, which is then transmitted to the receiving end. The receiving end performs image reconstruction based on the compressed encoded data to obtain a reconstructed image.

[0009] Secondly, embodiments of this disclosure also provide an image processing method, including:

[0010] The sending end acquires the target image, determines the region of interest and non-region of interest in the target image, determines mask blocks in the region of interest and non-region of interest respectively, and performs masking processing on the mask blocks to obtain a masked image, wherein the proportion of mask blocks in the region of interest is less than the proportion of mask blocks in the non-region of interest;

[0011] The transmitting end performs compression and encoding processing on the mask image to obtain compressed encoded data, and transmits the compressed encoded data to the receiving end.

[0012] The receiving end performs image reconstruction based on the compressed encoded data to obtain a reconstructed image.

[0013] Thirdly, embodiments of this disclosure also provide an image processing apparatus, including:

[0014] A region recognition module is used to acquire a target image and determine the region of interest and non-region of interest in the target image;

[0015] A masking module is used to determine mask blocks in the region of interest and the region of non-interest respectively, and to perform masking processing on the mask blocks to obtain a masked image, wherein the proportion of mask blocks in the region of interest is less than the proportion of mask blocks in the region of non-interest.

[0016] The compression and encoding module is used to compress and encode the mask image to obtain compressed and encoded data, transmit the compressed and encoded data to the receiving end, and the receiving end performs image reconstruction based on the compressed and encoded data to obtain a reconstructed image.

[0017] Fourthly, embodiments of this disclosure also provide an electronic device, the electronic device comprising:

[0018] One or more processors;

[0019] Storage device for storing one or more programs.

[0020] When the one or more programs are executed by the one or more processors, the one or more processors implement the image processing method provided in any embodiment of this disclosure.

[0021] Fifthly, embodiments of this disclosure also provide a storage medium containing computer-executable instructions, which, when executed by a computer processor, are used to perform the image processing method provided in any embodiment of this disclosure.

[0022] This embodiment of the disclosure reduces the amount of image data by determining and masking blocks in the target image. By compressing and encoding the masked image, the image compression rate is improved, reducing the amount of image data transmitted and the bandwidth occupied. Furthermore, regions of interest (ROIs) and non-ROIs are identified in the target image. Masking blocks are determined in the ROIs and non-ROIs based on different data ratios, with the proportion of masking blocks in the ROIs being less than that in the non-ROIs. This ensures that important information within the ROIs is not masked, preventing the loss of important image information during image compression and guaranteeing the reconstruction quality of the compressed and encoded data. Attached Figure Description

[0023] The above and other features, advantages, and aspects of the embodiments of this disclosure will become more apparent from the accompanying drawings and the following detailed description. Throughout the drawings, the same or similar reference numerals denote the same or similar elements. It should be understood that the drawings are schematic, and the originals and elements are not necessarily drawn to scale.

[0024] Figure 1 This is a schematic flowchart of an image processing method provided in an embodiment of the present disclosure;

[0025] Figure 2 This is a schematic diagram of the structure of a region of interest identification model provided in an embodiment of this disclosure;

[0026] Figure 3 This is a schematic flowchart of an image processing method provided in an embodiment of the present disclosure;

[0027] Figure 4 This is a flowchart of an image processing method provided in an embodiment of this disclosure;

[0028] Figure 5 This is a schematic diagram of the structure of an image processing apparatus provided in an embodiment of the present disclosure;

[0029] Figure 6 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this disclosure. Detailed Implementation

[0030] Embodiments of this disclosure will now be described in more detail with reference to the accompanying drawings. While some embodiments of this disclosure are shown in the drawings, it should be understood that this disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of this disclosure. It should be understood that the accompanying drawings and embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of protection of this disclosure.

[0031] It should be understood that the steps described in the method embodiments of this disclosure may be performed in different orders and / or in parallel. Furthermore, the method embodiments may include additional steps and / or omit the steps shown. The scope of this disclosure is not limited in this respect.

[0032] The term "comprising" and its variations as used herein are open-ended inclusions, meaning "including but not limited to". The term "based on" means "at least partially based on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". Definitions of other terms will be given in the description below.

[0033] It should be noted that the concepts of "first" and "second" mentioned in this disclosure are used only to distinguish different devices, modules or units, and are not used to limit the order of functions performed by these devices, modules or units or their interdependencies.

[0034] It should be noted that the terms "a" and "a plurality of" used in this disclosure are illustrative rather than restrictive, and those skilled in the art should understand that, unless otherwise expressly indicated in the context, they should be understood as "one or more".

[0035] The names of messages or information exchanged between multiple devices in the embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of such messages or information.

[0036] It is understood that before using the technical solutions disclosed in the various embodiments of this disclosure, users should be informed of the types, scope of use, and usage scenarios of the personal information involved in this disclosure in an appropriate manner in accordance with relevant laws and regulations, and user authorization should be obtained.

[0037] For example, upon receiving a user's active request, a prompt message is sent to the user to explicitly inform them that the requested operation will require the acquisition and use of the user's personal information. This allows the user to independently choose whether to provide personal information to the software or hardware, such as the electronic device, application, server, or storage medium performing the operations of this disclosed technical solution, based on the prompt message.

[0038] As an optional but non-limiting implementation, in response to a user's active request, sending a prompt message to the user can be done via a pop-up window, where the prompt message can be presented in text format. Furthermore, the pop-up window can also include a selection control allowing the user to choose "agree" or "disagree" to provide personal information to the electronic device.

[0039] It is understood that the above notification and user authorization process are merely illustrative and do not constitute a limitation on the implementation of this disclosure. Other methods that comply with relevant laws and regulations may also be applied to the implementation of this disclosure.

[0040] It is understood that the data involved in this technical solution (including but not limited to the data itself, the acquisition or use of the data) shall comply with the requirements of relevant laws, regulations and related provisions.

[0041] Figure 1 This is a schematic flowchart illustrating an image processing method provided in an embodiment of this disclosure. This embodiment is applicable to situations involving image compression. The method can be executed by an image processing device, which can be implemented in software and / or hardware, optionally through an electronic device, such as a mobile terminal, PC, or server. Figure 1 As shown, the method includes:

[0042] S110. Acquire the target image and determine the region of interest and non-region of interest in the target image.

[0043] S120. Determine mask blocks in the region of interest and the region of non-interest respectively, and perform masking processing on the mask blocks to obtain a mask image, wherein the proportion of mask blocks in the region of interest is less than the proportion of mask blocks in the region of non-interest.

[0044] S130. The mask image is compressed and encoded to obtain compressed encoded data. The compressed encoded data is transmitted to the receiving end. The receiving end performs image reconstruction based on the compressed encoded data to obtain a reconstructed image.

[0045] The target image is the image to be transmitted. This target image can be an independent image, an image frame from a video, or an image frame from live stream data; there are no restrictions on the method of acquisition. The target image can be locally stored, captured in real-time, or imported externally; there are no limitations on how the target image is acquired.

[0046] In this embodiment, by performing masking processing on the target image, the amount of image data to be compressed is reduced, thereby improving the compression ratio. Specifically, a mask block is determined in the target image, and masking processing is performed on the mask block to reduce the amount of image data to be compressed. Here, the mask block is the image block to be masked. The masking processing can be a process of converting the pixel values ​​within the image block into specific pixel values, where the specific pixel value can be 0. Specifically, this can be achieved by multiplying the image block with a mask template, converting the content within the image block into a single pixel value, thus reducing the amount of data within the image block. Simultaneously, the mask block can be an image block of a specific size, such as a rectangular block. The mask region is a local area among all image blocks in the target image. By dividing the region of interest and non-region of interest of the target image into blocks, mask blocks can be filtered among multiple image blocks, improving the granularity of image processing.

[0047] Furthermore, to avoid losing important image information during compression, regions of interest (ROIs) and non-ROIs are identified in the target image, and mask blocks of varying proportions are determined for each ROI. The ROI can be a region containing important information, or it can be a region containing a large amount of image detail. For example, the ROI can be the foreground region of the target image, and the non-ROI can be the background region. Alternatively, the ROI can be the region containing a specific object in the target image, and the non-ROI can be the region outside the ROI. For example, the specific object can be a face, but the specific object is not limited.

[0048] Specifically, this can involve identifying regions of interest (ROIs) in a target image and defining regions outside of ROIs as non-ROIs. ROIs can be one or more regions, and there is no limit to the number of ROIs.

[0049] In some embodiments, the identification of regions of interest (ROIs) can be achieved based on a user's selection operation. Optionally, determining ROIs and non-ROIs in the target image includes: displaying the target image, and in response to a region selection operation, identifying the selected region as a ROI and identifying other regions outside the selected region as non-ROIs.

[0050] The display interface may include a selection control that includes a region of interest. When the selection control is triggered, the system enters the region of interest selection mode, detects the user's selection operation, and determines the region of interest based on the selection operation.

[0051] Optionally, under the selected model, the grid of the target image is displayed. Correspondingly, the region selection operation can be a click operation on the grid; the clicked grid is the selected grid, which can be displayed separately. The image regions corresponding to the selected grids form the region of interest. Different grid division precisions can be selected to achieve different granularities of the region of interest, such as 5×5, 7×7, 10×10, etc. The display interface can include a precision selection control, such as a drop-down menu with multiple precision options, or a precision setting control, allowing users to adjust the grid division precision.

[0052] Optionally, in selection mode, edge recognition can be performed on the target image to obtain edge recognition results. These results include the contours of each object in the target image. Based on the edge recognition results, each object in the target image can be segmented according to its contour to obtain image regions for each object. Correspondingly, the region selection operation can be a click operation on an object image region, determining one or more selected object image regions as regions of interest. Further, a local segmentation of an object image region can be performed. Correspondingly, the region selection operation can be a click operation on a local area within the object image region. Specifically, a segmentation control is provided in the display interface. When an object image region is selected, the segmentation control is triggered, and edge recognition and contour segmentation are performed on the selected object image region to obtain a local region, which is then used to further select the region of interest, ensuring the accuracy of the region of interest selection. For example, if the object image region is the human body, the locally segmented region can include the head region, torso region, limb regions, etc., and the selected region of interest can be the head region. For example, if the object image region is the head region, the locally segmented region includes the face region, hair accessory region, earring region, etc., and the selected region of interest can be the face region.

[0053] Optionally, the region selection operation can be a sliding operation, defining the closed region formed by the sliding trajectory as the region of interest. This sliding trajectory can be a circular trajectory, a rectangular trajectory, or a sliding trajectory along the object's contour; there are no limitations on this. After detecting the sliding trajectory, it can be smoothed to obtain a smooth closed region. Optionally, edge recognition and image segmentation can be performed on the closed region formed by the sliding trajectory, defining the area containing the object within the closed region as the region of interest, removing background information within the closed region, and improving the accuracy of the region of interest.

[0054] In some embodiments, the target image is an image frame in a video. At least one reference image frame can be randomly selected from the video frames. A region of interest (ROI) is selected in the reference image frame, and objects included within the ROI of the reference image frame are identified. Based on these objects, the regions where the objects are located in other image frames of the video are identified as ROIs. For example, the reference image frame can be the first frame or any randomly selected frame. Objects within the ROI can be people, animals, etc., and are not limited thereto. Optionally, the video can be a live video, and the objects can be the broadcaster in the frame. The objects in at least one reference image frame can be the same or different, and can include at least one object. By identifying ROIs in each image frame of the video based on the objects in the reference image frames, the process of identifying ROIs in the video is simplified.

[0055] In some embodiments, the region of interest (ROI) can be automatically identified using a pre-set recognition algorithm. Optionally, determining the ROI and non-ROI regions in the target image includes: inputting the target image into a pre-trained ROI recognition model to obtain a ROI dataset, the ROI dataset including at least one coordinate data; determining the ROI based on the image blocks corresponding to the coordinate data; and determining other regions in the target image besides the ROI as non-ROI regions.

[0056] By using a pre-trained region of interest (ROI) identification model, the system automatically identifies ROIs in target images, simplifying the ROI identification process. This is especially beneficial for compressing large numbers of images, as automatic ROI identification significantly accelerates image compression efficiency. Furthermore, since the ROI identification model is a machine learning model, such as a neural network model, it offers high processing performance, further enhancing the efficiency of ROI identification.

[0057] In this embodiment, the output information of the region of interest (ROI) identification model is a ROI dataset, which includes at least one set of coordinate data. Optionally, the coordinate data is the coordinate data of a feature point corresponding to an image block within the ROI, and multiple image blocks constitute the ROI. Optionally, the image block can be a rectangular region, and correspondingly, the coordinate data can be the coordinate data of any vertex of the rectangular region, or the coordinate data of the center point of the rectangular region.

[0058] Furthermore, based on the coordinate data and the data parameters of the image blocks, the image blocks corresponding to the coordinate data are determined, and regions of interest are formed based on the image blocks corresponding to the coordinate data. Taking a rectangular image block as an example, the data parameters of the image block can be the width and height of the rectangular block. For example, the coordinate data can be the coordinate data of the top-left vertex of the rectangular region. Based on the coordinate data of the top-left vertex and the width and height of the image block, the coordinate data of the other vertices of the rectangular block can be determined to define the rectangular region. Taking a circular image block as an example, the data parameters of the image block can be the radius of the circular block.

[0059] Optionally, before performing region of interest (ROI) identification on the target image, the granularity of ROI identification can be set. This granularity is used to adjust the number of image blocks and data parameters. A smaller granularity results in a larger number of image blocks and smaller data parameters for each block; for example, the height and width of a rectangular block are smaller. By calling different ROI identification models at the set granularity, the target image is processed accordingly to obtain a ROI dataset that meets the granularity requirements. Further, based on the data parameters corresponding to this granularity and the aforementioned ROI dataset, the ROI and non-ROI regions are determined.

[0060] Optionally, the coordinate data of the region of interest (ROI) dataset consists of the coordinates of the feature points corresponding to the ROI. The ROI can be a rectangular or circular region, and the coordinate data can be the vertex coordinates or center point coordinates of a rectangular region, or the center coordinates of a circular region, etc. Correspondingly, the ROI is treated as an image block, and the data parameters of the image block are the overall data parameters of the ROI, such as height and width. Taking a rectangular ROI as an example, the coordinates of each vertex of the ROI can be determined using the coordinate data of the feature points and the data parameters of the ROI, thus further defining the ROI.

[0061] Optionally, the training method for the region of interest (ROI) recognition model includes: acquiring a sample image; dividing the sample image into blocks to obtain multiple image blocks; reconstructing the masked sample image after masking any of the image blocks to obtain a reconstructed image; obtaining importance data of the masked image blocks based on the image quality data of the reconstructed image; determining image blocks within the ROI region based on the importance data of each image block in the sample image, and forming a target dataset of the ROI region based on the image blocks within the ROI region; and iteratively training the ROI recognition model to be trained based on the sample image and the corresponding target dataset of the ROI region to obtain a trained ROI recognition model.

[0062] During training, the segmentation precision for the sample image blocks is the same as the segmentation precision for the target image blocks. Optionally, different segmentation precisions can be set by different recognition granularities to train region-of-interest (ROI) recognition models with different recognition granularities. These ROI recognition models with different recognition granularities can be trained in parallel. The segmentation precision is not limited here and can be set according to training requirements.

[0063] The sample image is divided into blocks, resulting in multiple image blocks. The importance data of each image block is determined by traversing each block. The determination of the importance data of each image block can be achieved by: for any image block, masking the block in the sample image to form a masked sample image; then reconstructing the image from the masked sample image to obtain the reconstructed image. This image reconstruction process can be implemented using a MAE (masked autoencoder) decoder, inputting the masked sample image into the MAE decoder to obtain the reconstructed image. The reconstructed image is then subjected to quality assessment to obtain its image quality data. This quality assessment can be achieved by inputting the reconstructed image into a quality assessment model to obtain the output image quality data; or by calculating the similarity between the reconstructed image and the sample image, using the similarity data as the instruction assessment data.

[0064] Based on this image quality data, the importance data of the image blocks to be masked can be determined. Higher image quality data indicates a smaller impact of the masked image block on the sample image, and thus a lower importance for that image block. Conversely, lower image quality data indicates a greater impact of the masked image block on the sample image, and thus a higher importance for that image block. Correspondingly, the importance data of an image block is negatively correlated with the image quality data of the reconstructed image. The importance data of the image blocks is obtained by substituting the image quality data into a pre-set importance data calculation formula.

[0065] The determination of whether an image block belongs to a region of interest (ROI) is based on the importance data of the image blocks. Optionally, the importance data can be compared with a threshold, and image blocks with importance data greater than the threshold can be identified as ROIs, thus determining the ROIs of the sample image. Optionally, a preset number of image blocks can be set for the ROIs, and the image blocks can be sorted based on the importance data. The preset number of image blocks can then be selected based on the sorting to form the ROIs.

[0066] Furthermore, the target dataset for the region of interest is determined based on the coordinate data of specific points within the image blocks within the region of interest. The above processing procedure is then performed on each sample image to obtain the corresponding target dataset, which serves as the label for the sample image. The initially constructed region of interest recognition model is then trained based on multiple sample images and their corresponding target datasets.

[0067] The training process for a region of interest (ROI) recognition model can be as follows: Input sample images into the ROI recognition model to be trained, obtaining the prediction dataset from the model. Generate a loss function based on the prediction dataset and the corresponding target dataset of the sample images. Adjust the model parameters of the ROI recognition model based on this loss function. Then, perform the next round of iterative training based on the adjusted ROI recognition model. Repeat this training process iteratively until the training termination condition is met, resulting in a trained ROI recognition model. The type of loss function is not limited during the training process and can be set according to requirements.

[0068] Based on the above embodiments, the region of interest identification model can be a CNN model, for example, see [link to example]. Figure 2 , Figure 2 This is a schematic diagram of the structure of a region of interest (ROI) identification model provided in an embodiment of this disclosure. The ROI identification model may include three CNN convolutional layers and one fully connected layer. It is understood that... Figure 2 The region of interest (ROI) identification model described herein is merely an example. In other embodiments, it can be implemented using neural network models with other structures, as long as they have ROI identification functionality; there is no limitation on this. Optionally, the ROI identification model can also be a multi-object detection model used to identify multiple ROIs. For example, the ROI identification model could be a Faster-RCNN model or a YOLO model, etc.

[0069] Based on any of the above methods, the region of interest (ROI) of the target image is identified, and the ROI and non-ROI regions of the target image are determined. Based on the masking strategy, mask blocks are determined in the ROI and non-ROI regions respectively and masking processing is performed. This achieves targeted masking processing of the target image, preserves important blocks in the target image, avoids the loss of important information due to masking processing, improves the compression ratio, avoids the loss of important information, and further ensures that the target image can be reconstructed with high quality after compression.

[0070] In this embodiment, mask blocks are determined separately in the region of interest (ROI) and the non-ROI. The proportions of mask blocks determined in the ROI and the non-ROI are different to ensure that the number of mask blocks in the ROI is small, thus retaining more image blocks within the ROI and avoiding the loss of important information. Optionally, the number of mask blocks in the ROI is less than the number of mask blocks in the non-ROI, or the proportion of mask blocks in the ROI is less than the proportion of mask blocks in the non-ROI.

[0071] Optionally, determining mask blocks within the region of interest and the region of non-interest includes: determining the number of mask blocks based on the mask ratio and the size of the target image; determining a first number of mask blocks within the region of interest and a second number of mask blocks within the region of non-interest based on the data ratio between the region of interest and the region of non-interest and the number of mask blocks, wherein the data ratio between the region of interest and the region of non-interest is less than 1; determining mask blocks within the region of interest based on the first number, and determining mask blocks within the region of non-interest based on the second number.

[0072] The mask rate is the proportion of all masked blocks in the target image, and it is positively correlated with the image compression rate. The number of masked blocks can be determined by the mask rate and the number of image blocks in the target image. The number of image blocks can be positively correlated with the size of the target image; here, the data parameters for each image block are preset. For example, with a mask rate of 50% and 100 image blocks in the target image, the corresponding number of masked blocks is 50, meaning the sum of the number of masked blocks in the region of interest and the non-region of interest is 50.

[0073] A data ratio between the region of interest (ROI) and the region of non-ROI is set, and this ratio is less than 1, meaning the number of masked blocks in the ROI is less than the number of masked blocks in the region of non-ROI. For example, the ROI to non-ROI data ratio could be 2:3. Based on the above data ratio and the number of masked blocks, a first number of masked blocks within the ROI and a second number of masked blocks within the non-ROI are determined. In the example above, the number of masked blocks is 50, and the ROI to non-ROI data ratio can be 2:3. Accordingly, the first number of masked blocks within the ROI is 20, and the second number of masked blocks within the non-ROI is 30. By setting the ROI to non-ROI data ratio to be less than 1, the number of masked blocks in the ROI is reduced, avoiding the problem of a large number of blocks within the ROI being masked, which could mask important information and negatively impact the image quality of the reconstructed image.

[0074] Mask blocks are determined in the region of interest (ROI) and non-ROI based on a first and second set of predetermined quantities. Optionally, the mask blocks can be randomly determined. Optionally, a mask priority is determined for each image block within the ROI, and mask blocks are determined based on this priority. Taking a face region as an example, which includes the eye region, mouth region, nose region, and cheek region, the mask priority of the eye region, mouth region, nose region, and cheek region increases sequentially according to the settings. Correspondingly, the number of mask blocks in the eye region, mouth region, nose region, and cheek region increases sequentially. This is just an example. Different priority attributes are set for different types of ROI. Here, the priority attribute can be negatively correlated with the importance of the image block to image reconstruction, that is, the more important the image block, the lower the mask priority, so as to reduce the probability of important image blocks being masked.

[0075] In some embodiments, the region of interest (ROI) occupies a small proportion of the target image. Setting a fixed data ratio between the ROI and non-ROI cannot be adapted to all target images. Here, the number of mask regions within the ROI can be determined based on the number of image blocks within the ROI, allowing for flexible adjustment of the number of mask regions within the ROI.

[0076] Optionally, determining mask blocks within the region of interest and the region of non-interest respectively includes: determining a first number of mask blocks within the region of interest based on a first data ratio and the number of image blocks within the region of interest, and determining mask blocks within the region of interest based on the first number; determining a second number of mask blocks within the region of non-interest based on a second data ratio and the number of image blocks within the region of non-interest, and determining mask blocks within the region of non-interest based on the second number, wherein the second data ratio is greater than the first data ratio.

[0077] Here, the first data ratio is the ratio of the number of masked regions within the region of interest (ROI) to the number of image blocks within the ROI, and the second data ratio is the ratio of the number of masked regions within the non-ROI to the number of image blocks within the non-ROI. The first data ratio is less than the second data ratio; for example, the first data ratio could be 20%, and the second data ratio could be 50%. When the number of image blocks within the ROI is 20 and the number of image blocks within the non-ROI is 80, the first number of masked blocks within the ROI is 4, and the second number of masked blocks within the non-ROI is 40. This ensures the number of masked blocks in the target image while preserving most of the image blocks within the ROI, preventing important information within the ROI from being masked.

[0078] Similarly, mask blocks can be determined in regions of interest and non-interest, either randomly or by the mask priority of image blocks.

[0079] By performing masking processing on the determined masked blocks, a masked image is obtained, which reduces the amount of image data and facilitates the improvement of the compression ratio in subsequent compression processing.

[0080] The mask image is compressed and encoded to obtain compressed encoded data. Specifically, the compressed encoded data is determined by: compressing the mask image based on a compression encoder to obtain compressed data; and encoding the compressed data based on a mask autoencoder to obtain the compressed encoded data.

[0081] Compression processing can be achieved through a compression encoder, including but not limited to HEVC (High Efficiency Video Coding) encoders, JPEG (Joint Photographic Experts Group) encoders, and WebP encoders. The mask autoencoder can be a pre-trained MAE encoder, which further encodes the compressed mask image to obtain compressed encoded data, which can be the embedding data corresponding to the compressed data, thus achieving further data encoding. The MAE encoder and MAE decoder are pre-trained and are respectively located at the sending and receiving ends. The MAE encoder encodes the compressed mask image, and the MAE decoder performs reverse decoding on the compressed encoded data to reconstruct the masked image blocks for image reconstruction. The MAE encoder and MAE decoder can be synchronously trained based on sample images and configured at the sending and receiving ends, respectively. In some embodiments, an electronic device, such as a mobile phone, can act as both a sending and receiving end, and accordingly, the electronic device can be equipped with both a MAE encoder and a MAE decoder.

[0082] Understandably, the sending end sends compressed encoded data to the receiving end, which is equipped with a compression decoder and a mask autoencoder corresponding to the compression encoder and mask autoencoder, respectively, to perform reverse decoding processing on the compressed encoded data, thereby reconstructing the image and obtaining the reconstructed image.

[0083] Based on the above embodiments, the sending end determines the location data of the non-masked blocks in the target image and sends the location data to the receiving end. Correspondingly, the receiving end receives the compressed encoded data and the location data, decodes the compressed encoded data and the location data using a mask self-decoder to obtain compressed data, and decompresses the compressed data using a compression decoder to obtain the reconstructed image.

[0084] The location data of the non-masked block can be a set of coordinate data of feature points of the non-masked block. These feature points can be the center point or vertices of the non-masked block, etc., without limitation. By sending the location data of the non-masked block to the receiving end, it serves as reference information during the reconstruction process of the compressed encoded data, thereby improving the reconstruction efficiency and quality of the reconstructed image.

[0085] The technical solution of this disclosure reduces the amount of image data by determining and masking blocks in the target image. It also improves the image compression rate and reduces the amount of image data transmitted and the bandwidth occupied by transmission by compressing and encoding the masked image. Furthermore, it identifies regions of interest (ROI) and non-ROI regions in the target image, and determines masking blocks in both regions based on different data ratios. The proportion of masking blocks in the ROI regions is less than that in the non-ROI regions, ensuring that important information within the ROI regions is not masked, thus preventing the loss of important image information during image compression and guaranteeing the reconstruction quality of the compressed and encoded data.

[0086] Figure 3 This is a schematic flowchart illustrating an image processing method provided in an embodiment of the present disclosure. The method specifically includes the following steps:

[0087] S210. The sending end acquires the target image, determines the region of interest and non-region of interest in the target image, determines mask blocks in the region of interest and non-region of interest respectively, and performs masking processing on the mask blocks to obtain a masked image, wherein the proportion of mask blocks in the region of interest is less than the proportion of mask blocks in the non-region of interest.

[0088] S220. The transmitting end performs compression and encoding processing on the mask image to obtain compressed encoded data, and transmits the compressed encoded data to the receiving end.

[0089] S230. The receiving end performs image reconstruction based on the compressed encoded data to obtain a reconstructed image.

[0090] In this embodiment, the transmitting end can be a mobile terminal such as a mobile phone, or an electronic device such as a computer or server. The receiving end can be a mobile terminal such as a mobile phone, or an electronic device such as a computer or server. In some embodiments, an electronic device can be either a receiving end or a transmitting end in different transmission scenarios. Accordingly, the electronic device can be configured with a compression encoder, a compression decoder, a mask autoencoder, and a mask autodecoder, which are used to compress the image before transmission and to reconstruct the image after receiving the compressed encoded data, respectively.

[0091] The technical solution in this embodiment improves the image compression rate and reduces the risk of losing important information by performing masking, compression, and MAE encoding at the sending end before transmission. Reverse decoding at the receiving end improves the quality of the reconstructed image.

[0092] Based on the above embodiments, this disclosure also provides a preferred example of an image processing method, see [link to example]. Figure 4 , Figure 4 This is a flowchart of an image processing method provided in an embodiment of this disclosure.

[0093] The transmitting and receiving ends are respectively equipped with a MAE encoder and a MAE decoder for masking and image reconstruction. The transmitting end performs masking processing on the original image I_0 (i.e., the target image in the above embodiment) to obtain a masked image I_MASK. Specifically, the original image I_0 is input into the ROI masking model (i.e., the region of interest identification model in the above embodiment) to obtain a masked block dataset. Based on this masked block dataset, the region of interest is determined and then masked.

[0094] The region of interest here refers to the Region of Interest (ROI) based on image quality. During the training of the ROI masking model, sample images are collected and labeled. These labels are from a quality-based ROI detection label dataset. Specifically, each sample image is uniformly scaled to 224*224 pixels, divided into 7*7 blocks, and each block is masked. The image is then reconstructed using the MAE decoder, and its image quality is scored (i.e., the importance data of the image blocks). Finally, each block in each image receives a score, and the 4*5 blocks with the highest total scores (this data is only for example) are selected as the label for this sample image.

[0095] A Region of Interest (ROI) masking model is trained based on sample images and labels. This ROI masking model can be a single-object detection model, consisting of three CNN convolutional layers and one fully connected layer. The three CNN layers extract image features from different dimensions, and the extracted features are then adaptively averaged using a final average pooling layer. The final extracted features are "flattened," that is, the multi-dimensional input is reduced to one dimension and passed to the fully connected layer. The fully connected layer outputs two values, representing the x and y coordinates of the top-left corner of the ROI region. The ROI masking model can also be a multi-object detection model, such as Faster R-CNN or YOLO.

[0096] Based on the horizontal and vertical coordinates of the top-left corner of each image block, as well as the width and height of the image block, the coordinates of the four vertices of the image block can be determined. A random sampling approach is adopted, with a random mask ratio of 2:3 between the ROI region and the outside region (non-ROI region). Within the ROI region, the mask ratio is smaller than outside the ROI region, thereby improving the image quality of the decoder-reconstructed image.

[0097] The receiving end receives the image embedding (i.e., compressed encoded data) and the position embedding (i.e., position data), performs decoding processing through the MAE decoder, and performs decompression processing through the HEVC decoder to obtain the output image I_out (i.e., the reconstructed image).

[0098] Figure 5 This is a schematic diagram of the structure of an image processing apparatus provided in an embodiment of the present disclosure, as shown below. Figure 5 As shown, the device includes: a region identification module 310, a mask module 320, and a compression and encoding module 330.

[0099] Region recognition module 310 is used to acquire a target image and determine the region of interest and non-region of interest in the target image;

[0100] The masking module 320 is used to determine mask blocks in the region of interest and the region of non-interest respectively, and to perform masking processing on the mask blocks to obtain a mask image, wherein the proportion of mask blocks in the region of interest is less than the proportion of mask blocks in the region of non-interest.

[0101] The compression and encoding module 330 is used to compress and encode the mask image to obtain compressed and encoded data, transmit the compressed and encoded data to the receiving end, and the receiving end performs image reconstruction based on the compressed and encoded data to obtain a reconstructed image.

[0102] The technical solution provided in this disclosure reduces image data volume by determining and masking blocks in the target image. It also improves image compression rate and reduces image transmission data volume and bandwidth usage by compressing and encoding the masked image. Furthermore, it identifies regions of interest (ROI) and non-ROI regions in the target image and determines masking blocks in both regions based on different data ratios. The proportion of masking blocks in the ROI regions is less than that in the non-ROI regions, ensuring that important information within the ROI regions is not masked, preventing the loss of important image information during image compression, and guaranteeing the reconstruction quality of the compressed and encoded data.

[0103] Based on the above embodiments, optionally, the region identification module 310 is used for:

[0104] The target image is input into a pre-trained region of interest (ROI) recognition model to obtain a ROI dataset, which includes at least one coordinate data.

[0105] The region of interest is determined based on the image blocks corresponding to the coordinate data, and other regions in the target image other than the region of interest are determined as regions of non-interest.

[0106] Optionally, the coordinate data is the coordinate data of a feature point corresponding to an image block;

[0107] The region identification module 310 is also used for:

[0108] Based on the coordinate data and the data parameters of the image block, the image block corresponding to the coordinate data is determined.

[0109] Optionally, based on the above embodiments, the device further includes:

[0110] The model training module is used to acquire sample images and divide the sample images into blocks to obtain multiple image blocks.

[0111] In the case of masking any of the image blocks, the masked sample image is reconstructed to obtain a reconstructed image, and the importance data of the masked image block is obtained based on the image quality data of the reconstructed image.

[0112] Based on the importance data of each image block in the sample image, the image blocks within the region of interest are determined, and a target dataset of the region of interest is formed based on the image blocks within the region of interest.

[0113] Based on the sample images and the target dataset of the corresponding regions of interest, the region of interest recognition model to be trained is iteratively trained to obtain a trained region of interest recognition model.

[0114] Based on the above embodiments, optionally, the region identification module 310 is used for:

[0115] The target image is displayed, and in response to a region selection operation, the selected region is identified as the region of interest, and other regions outside the selected region are identified as regions of non-interest.

[0116] Based on the above embodiments, optionally, the mask module 320 is used for:

[0117] The number of mask blocks is determined based on the mask rate and the size of the target image;

[0118] Based on the data ratio between the region of interest and the region of non-interest and the number of mask blocks, a first number of mask blocks within the region of interest and a second number of mask blocks within the region of non-interest are determined, respectively, wherein the data ratio between the region of interest and the region of non-interest is less than 1;

[0119] A mask block is determined in the region of interest based on the first quantity, and a mask block is determined in the region of non-interest based on the second quantity.

[0120] Based on the above embodiments, optionally, the mask module 320 is used for:

[0121] A first number of mask blocks within the region of interest is determined based on a first data ratio and the number of image blocks within the region of interest, and mask blocks are then determined within the region of interest based on the first number.

[0122] A second number of mask blocks in the non-interest area is determined based on a second data ratio and the number of image blocks in the non-interest area, and mask blocks are determined in the non-interest area based on the second number, wherein the second data ratio is greater than the first data ratio.

[0123] Based on the above embodiments, optionally, the compression and encoding module 330 is used for:

[0124] The mask image is compressed using a compression encoder to obtain compressed data;

[0125] The compressed data is encoded using a mask autoencoder to obtain the compressed encoded data.

[0126] Optionally, based on the above embodiments, the device further includes:

[0127] The location data transmission module is used to determine the location data of the non-masked block in the target image and send the location data to the receiving end.

[0128] Based on the above embodiments, optionally, the receiving end receives the compressed encoded data and the location data, decodes the compressed encoded data and the location data based on the mask self-decoder to obtain compressed data, and decompresses the compressed data based on the compression decoder to obtain the reconstructed image.

[0129] The image processing apparatus provided in this disclosure can execute the image processing method provided in any embodiment of this disclosure, and has the corresponding functional modules and beneficial effects for executing the method.

[0130] It is worth noting that the various units and modules included in the above-mentioned device are only divided according to functional logic, but are not limited to the above division, as long as the corresponding functions can be realized; in addition, the specific names of each functional unit are only for easy differentiation and are not used to limit the protection scope of the embodiments of this disclosure.

[0131] Figure 6 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this disclosure. Reference is made below. Figure 6 It illustrates an electronic device suitable for implementing embodiments of the present disclosure (e.g., Figure 6 The diagram below shows the structure of the terminal device or server 500. The terminal device in this embodiment may include, but is not limited to, mobile terminals such as mobile phones, laptops, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), and vehicle terminals (e.g., vehicle navigation terminals), as well as fixed terminals such as digital TVs and desktop computers. Figure 6 The electronic device shown is merely an example and should not be construed as limiting the functionality and scope of the embodiments disclosed herein.

[0132] like Figure 6 As shown, electronic device 500 may include a processing unit (e.g., central processing unit, graphics processor, etc.) 501, which can perform various appropriate actions and processes according to a program stored in read-only memory (ROM) 502 or a program loaded from storage device 508 into random access memory (RAM) 503. The RAM 503 also stores various programs and data required for the operation of electronic device 500. The processing unit 501, ROM 502, and RAM 503 are interconnected via bus 504. An edit / output (I / O) interface 505 is also connected to bus 504.

[0133] Typically, the following devices can be connected to I / O interface 505: input devices 506 including, for example, touchscreens, touchpads, keyboards, mice, cameras, microphones, accelerometers, gyroscopes, etc.; output devices 507 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices 508 including, for example, magnetic tapes, hard disks, etc.; and communication devices 509. Communication device 509 allows electronic device 500 to communicate wirelessly or wiredly with other devices to exchange data. Although Figure 6 An electronic device 500 with various devices is shown; however, it should be understood that it is not required to implement or possess all of the devices shown. More or fewer devices may be implemented or possessed alternatively.

[0134] In particular, according to embodiments of this disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of this disclosure include a computer program product comprising a computer program carried on a non-transitory computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device 509, or installed from a storage device 508, or installed from a ROM 502. When the computer program is executed by the processing device 501, it performs the functions defined in the methods of embodiments of this disclosure.

[0135] The names of messages or information exchanged between multiple devices in the embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of such messages or information.

[0136] The electronic device provided in this embodiment and the image processing method provided in the above embodiments belong to the same inventive concept. Technical details not described in detail in this embodiment can be found in the above embodiments, and this embodiment has the same beneficial effects as the above embodiments.

[0137] This disclosure provides a computer storage medium storing a computer program that, when executed by a processor, implements the image processing method provided in the above embodiments.

[0138] It should be noted that the computer-readable medium described in this disclosure can be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. A computer-readable storage medium can be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this disclosure, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in connection with an instruction execution system, apparatus, or device. In this disclosure, a computer-readable signal medium can include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium can be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to: wires, optical fibers, RF (radio frequency), etc., or any suitable combination thereof.

[0139] In some implementations, clients and servers can communicate using any currently known or future-developed network protocol such as HTTP (Hypertext Transfer Protocol) and can interconnect with digital data communication (e.g., communication networks) of any form or medium. Examples of communication networks include local area networks (“LANs”), wide area networks (“WANs”), the Internet (e.g., the Internet of Things), and peer-to-peer networks (e.g., ad hoc peer-to-peer networks), as well as any currently known or future-developed networks.

[0140] The aforementioned computer-readable medium may be included in the aforementioned electronic device; or it may exist independently and not assembled into the electronic device.

[0141] The aforementioned computer-readable medium carries one or more programs that, when executed by the electronic device, cause the electronic device to:

[0142] The aforementioned computer-readable medium carries one or more programs that, when executed by the electronic device, cause the electronic device to: acquire a target image; determine a region of interest (ROI) and a non-ROI in the target image; determine mask blocks in the ROI and the non-ROI respectively, and perform masking processing on the mask blocks to obtain a masked image, wherein the proportion of mask blocks in the ROI is less than the proportion of mask blocks in the non-ROI; perform compression and encoding processing on the masked image to obtain compressed encoded data; transmit the compressed encoded data to a receiving end; and the receiving end performs image reconstruction based on the compressed encoded data to obtain a reconstructed image.

[0143] Computer program code for performing the operations of this disclosure can be written in one or more programming languages ​​or a combination thereof, including but not limited to object-oriented programming languages ​​such as Java, Smalltalk, and C++, as well as conventional procedural programming languages ​​such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0144] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.

[0145] The units described in the embodiments of this disclosure can be implemented in software or in hardware. The name of a unit does not necessarily limit the unit itself; for example, the first acquisition unit can also be described as "a unit that acquires at least two Internet Protocol addresses".

[0146] The functions described above in this document can be performed, at least in part, by one or more hardware logic components. For example, exemplary types of hardware logic components that can be used, without limitation, include: Field Programmable Gate Arrays (FPGAs), Application-Specific Integrated Circuits (ASICs), Application Standard Products (ASSPs), System-on-Chip (SoCs), Complex Programmable Logic Devices (CPLDs), and so on.

[0147] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0148] [In the detailed implementation section, after the entire text ends, please repeat all the content that you wish to protect in the form of claims in the following form:]

[0149] According to one or more embodiments of this disclosure, [Example 1] provides an image processing method, including:

[0150] Acquire a target image and determine the region of interest and non-region of interest in the target image;

[0151] Mask blocks are determined in the region of interest and the region of non-interest respectively, and the mask blocks are masked to obtain a mask image, wherein the proportion of mask blocks in the region of interest is less than the proportion of mask blocks in the region of non-interest.

[0152] The masked image is compressed and encoded to obtain compressed encoded data, which is then transmitted to the receiving end. The receiving end performs image reconstruction based on the compressed encoded data to obtain a reconstructed image.

[0153] According to one or more embodiments of this disclosure, Example 2 provides the image processing method of Example 1, which further includes:

[0154] The step of determining the region of interest (ROI) and non-ROI regions in the target image includes: inputting the target image into a pre-trained ROI recognition model to obtain a ROI dataset, the ROI dataset including at least one coordinate data; determining the ROI based on the image blocks corresponding to the coordinate data; and determining other regions in the target image besides the ROI as non-ROI regions.

[0155] According to one or more embodiments of this disclosure, [Example 3] provides the image processing method of Example 1, which further includes: the coordinate data is the coordinate data of a feature point corresponding to an image block;

[0156] Determining the region of interest based on the image block corresponding to the coordinate data includes: determining the image block corresponding to the coordinate data based on the coordinate data and the data parameters of the image block.

[0157] According to one or more embodiments of this disclosure, Example 4 provides the image processing method of Example 1, which further includes:

[0158] The training method for the region of interest (ROI) recognition model includes: acquiring a sample image; dividing the sample image into blocks to obtain multiple image blocks; reconstructing the masked sample image after masking any of the image blocks to obtain a reconstructed image; obtaining importance data of the masked image blocks based on the image quality data of the reconstructed image; determining image blocks within the ROI region based on the importance data of each image block in the sample image; and forming a target dataset of the ROI region based on the image blocks within the ROI region; and iteratively training the ROI recognition model to be trained based on the sample image and the corresponding target dataset of the ROI region to obtain a trained ROI recognition model.

[0159] According to one or more embodiments of this disclosure, Example 5 provides the image processing method of Example 1, which further includes:

[0160] The step of determining the region of interest and non-region of interest in the target image includes: displaying the target image, and in response to a region selection operation, determining the selected region as the region of interest and determining other regions outside the selected region as non-regions of interest.

[0161] According to one or more embodiments of this disclosure, Example Six provides the image processing method of Example One, which further includes:

[0162] The step of determining mask blocks in the region of interest and the region of non-interest includes: determining the number of mask blocks based on the mask ratio and the size of the target image; determining a first number of mask blocks in the region of interest and a second number of mask blocks in the region of non-interest based on the data ratio of the region of interest and the region of non-interest and the number of mask blocks, wherein the data ratio of the region of interest and the region of non-interest is less than 1; determining mask blocks in the region of interest based on the first number, and determining mask blocks in the region of non-interest based on the second number.

[0163] According to one or more embodiments of this disclosure, Example 7 provides the image processing method of Example 1, which further includes:

[0164] Determining mask blocks in the region of interest and the region of non-interest respectively includes: determining a first number of mask blocks in the region of interest based on a first data ratio and the number of image blocks in the region of interest, and determining mask blocks in the region of interest based on the first number; determining a second number of mask blocks in the region of non-interest based on a second data ratio and the number of image blocks in the region of non-interest, and determining mask blocks in the region of non-interest based on the second number, wherein the second data ratio is greater than the first data ratio.

[0165] According to one or more embodiments of this disclosure, Example 8 provides an image processing method of Example 1, which further includes:

[0166] The step of compressing and encoding the mask image to obtain compressed encoded data includes: compressing the mask image based on a compression encoder to obtain compressed data; and encoding the compressed data based on a mask autoencoder to obtain the compressed encoded data.

[0167] According to one or more embodiments of this disclosure, Example 9 provides the image processing method of Example 1, which further includes:

[0168] The method further includes: determining the location data of non-masked blocks in the target image, and sending the location data to the receiving end.

[0169] According to one or more embodiments of this disclosure, Example 10 provides an image processing method of Example 1, which further includes:

[0170] The receiving end receives the compressed encoded data and the location data, decodes the compressed encoded data and the location data based on the mask self-decoder to obtain compressed data, and decompresses the compressed data based on the compression decoder to obtain the reconstructed image.

[0171] According to one or more embodiments of this disclosure, [Example 11] provides an image processing method, including:

[0172] The sending end acquires the target image, determines the region of interest and non-region of interest in the target image, determines mask blocks in the region of interest and non-region of interest respectively, and performs masking processing on the mask blocks to obtain a masked image, wherein the proportion of mask blocks in the region of interest is less than the proportion of mask blocks in the non-region of interest;

[0173] The transmitting end performs compression and encoding processing on the mask image to obtain compressed encoded data, and transmits the compressed encoded data to the receiving end.

[0174] The receiving end performs image reconstruction based on the compressed encoded data to obtain a reconstructed image.

[0175] According to one or more embodiments of this disclosure, [Example Twelve] provides an image processing apparatus, including:

[0176] A region recognition module is used to acquire a target image and determine the region of interest and non-region of interest in the target image;

[0177] A masking module is used to determine mask blocks in the region of interest and the region of non-interest respectively, and to perform masking processing on the mask blocks to obtain a masked image, wherein the proportion of mask blocks in the region of interest is less than the proportion of mask blocks in the region of non-interest.

[0178] The compression and encoding module is used to compress and encode the mask image to obtain compressed and encoded data, transmit the compressed and encoded data to the receiving end, and the receiving end performs image reconstruction based on the compressed and encoded data to obtain a reconstructed image.

[0179] The above description is merely a preferred embodiment of this disclosure and an explanation of the technical principles employed. Those skilled in the art should understand that the scope of this disclosure is not limited to technical solutions formed by specific combinations of the above-described technical features, but should also cover other technical solutions formed by arbitrary combinations of the above-described technical features or their equivalents without departing from the above-described concept. For example, technical solutions formed by substituting the above features with (but not limited to) technical features disclosed in this disclosure that have similar functions.

[0180] Furthermore, while the operations are described in a specific order, this should not be construed as requiring these operations to be performed in the specific order shown or in a sequential order. In certain environments, multitasking and parallel processing may be advantageous. Similarly, while several specific implementation details are included in the above discussion, these should not be construed as limiting the scope of this disclosure. Certain features described in the context of individual embodiments may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented individually or in any suitable sub-combination in multiple embodiments.

[0181] Although the subject matter has been described using language specific to structural features and / or methodological logic, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or actions described above. Rather, the specific features and actions described above are merely illustrative examples of implementing the claims.

Claims

1. An image processing method, characterized by, The method comprises: acquiring a target image, determining a region of interest and a non-region of interest in the target image; determining mask blocks in the region of interest and the non-region of interest respectively, and performing mask processing on the mask blocks to obtain a mask image, wherein the proportion of the number of mask blocks in the region of interest is less than the proportion of the number of mask blocks in the non-region of interest, and the mask processing is to convert the content in the mask blocks into a single pixel value; performing compression processing and encoding processing on the mask image to obtain compressed and encoded data, and transmitting the compressed and encoded data to a receiving end, and the receiving end reconstructs an image based on the compressed and encoded data to obtain a reconstructed image.

2. The method of claim 1, wherein, The method comprises: inputting the target image into a pre-trained region of interest identification model to obtain a region of interest data set, the region of interest data set comprising at least one coordinate data; determining a region of interest based on the image block corresponding to the coordinate data, and determining other regions in the target image except the region of interest as non-region of interest.

3. The method of claim 2, wherein, The coordinate data is coordinate data of a feature point corresponding to an image block; The method comprises: determining the image block corresponding to the coordinate data based on the coordinate data and the data parameters of the image block.

4. The method of claim 2, wherein, The training method of the region of interest identification model comprises: acquiring a sample image, and dividing the sample image into a plurality of image blocks; reconstructing the sample image after mask processing to obtain a reconstructed image, and obtaining importance data of the masked image block based on the image quality data of the reconstructed image; determining the image block in the region of interest based on the importance data of each image block in the sample image, and forming a target data set of the region of interest based on the image block in the region of interest; iteratively training a region of interest identification model to be trained based on the sample image and the target data set of the corresponding region of interest, to obtain a trained region of interest identification model.

5. The method of claim 1, wherein, The method comprises: displaying the target image, and determining a selected region as a region of interest and other regions outside the selected region as non-region of interest in response to a region selection operation.

6. The method of claim 1, wherein, The method comprises: determining the number of mask blocks according to the mask rate and the size of the target image; determining a first number of mask blocks in the region of interest and a second number of mask blocks in the non-region of interest based on the data proportion of the region of interest and the non-region of interest and the number of mask blocks, wherein the data proportion of the region of interest and the non-region of interest is less than 1; determining mask blocks in the region of interest based on the first number, and determining mask blocks in the non-region of interest based on the second number.

7. The method of claim 1, wherein, The determining the mask blocks in the region of interest and the region not of interest respectively comprises: determining a first number of mask blocks in the region of interest based on a first data proportion and a number of image blocks in the region of interest, and determining mask blocks in the region of interest based on the first number; determining a second number of mask blocks in the region not of interest based on a second data proportion and a number of image blocks in the region not of interest, and determining mask blocks in the region not of interest based on the second number, the second data proportion being greater than the first data proportion.

8. The method of claim 1, wherein, The compressing and encoding the mask image to obtain compressed and encoded data comprises: compressing the mask image based on a compression encoder to obtain compressed data; encoding the compressed data based on a mask autoencoder to obtain the compressed and encoded data.

9. The method of claim 1, wherein, The method further comprises: determining position data of non-mask blocks in the target image, and sending the position data to the receiving end.

10. The method of claim 9, wherein, The receiving end receives the compressed and encoded data and the position data, decodes the compressed and encoded data and the position data based on a mask auto-decoder to obtain compressed data, and decompresses the compressed data based on a compression decoder to obtain a reconstructed image.

11. An image processing method, characterized by, Comprise: The sending end obtains a target image, determines a region of interest and a region not of interest in the target image, determines mask blocks in the region of interest and the region not of interest respectively, and performs mask processing on the mask blocks to obtain a mask image, wherein the proportion of the number of mask blocks in the region of interest is less than the proportion of the number of mask blocks in the region not of interest, and the mask processing is converting the content in the mask blocks into a single pixel value; The sending end compresses and encodes the mask image to obtain compressed and encoded data, and transmits the compressed and encoded data to the receiving end; The receiving end reconstructs an image based on the compressed and encoded data to obtain a reconstructed image.

12. An image processing apparatus characterized by comprising: Comprise: The region identification module is used to obtain a target image, and determine a region of interest and a region not of interest in the target image; The mask module is used to determine mask blocks in the region of interest and the region not of interest respectively, and perform mask processing on the mask blocks to obtain a mask image, wherein the proportion of the number of mask blocks in the region of interest is less than the proportion of the number of mask blocks in the region not of interest, and the mask processing is converting the content in the mask blocks into a single pixel value; The compression and encoding module is used to compress and encode the mask image to obtain compressed and encoded data, and transmit the compressed and encoded data to the receiving end, and the receiving end reconstructs an image based on the compressed and encoded data to obtain a reconstructed image.

13. An electronic device, comprising: The electronic device comprises: one or more processors; a storage device for storing one or more programs, when the one or more programs are executed by the one or more processors, the one or more processors implement the image processing method as claimed in any one of claims 1-10.

14. A storage medium containing computer-executable instructions for performing the image processing method of any one of claims 1-10 when executed by a computer processor.

Citation Information

Patent Citations

  • Method and system of video coding using an image data correction mask

    CN109076246A

  • Model training and feature extraction method and device, electronic equipment and medium

    CN114842457A

  • Method and system of region-based image coding with dynamic streaming of code blocks

    WO2000049571A2