System and method for encoding image
By identifying and checking critical image regions and applying non-generative encoding where necessary, the method ensures accurate and reliable image compression, addressing the issue of hallucinations in machine learning-based encoding.
Patent Information
- Application Number
- JP2024178595
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-10-19
- Filing Date
- 2024-10-11
- Publication Date
- 2025-05-09
AI Technical Summary
Machine learning-based image compression methods, such as GANs, can introduce 'hallucinations' or false features in encoded images, which may compromise the probative value of the data, especially in legal or forensic contexts where accurate representation is crucial.
A method and system that identifies regions of interest in an image, performs a quality check by comparing reference points between the original and decoded images, and uses a non-generative model for encoding if the difference exceeds a threshold, ensuring accurate representation of critical areas while maintaining low bit rates.
Minimizes the risk of hallucinations in critical image areas while achieving low bit rates by selectively using generative and non-generative encoding methods, ensuring the integrity of the encoded image data.
Smart Images

Figure 2025072311000001_ABST
Abstract
Description
[Technical field]
[0001] The present disclosure relates to a method for encoding images in a video, in particular to a method for encoding images and performing quality checks thereof using artificial intelligence. The present disclosure further relates to an image processing system for performing the disclosed method. [Background technology]
[0002] To reduce the cost of storage and compression, image and video compression is widely used. Image and video compression generally works by applying algorithms to encode image / video data during compression and to decode image / video data during playback or when the data is received. Traditional compression often involves several mathematical methods, such as color space conversion, spatial compression, and temporal compression.
[0003] More recently, machine learning based methods have been developed, using for example convolutional neural networks and generative adversarial networks (GANs). Artificial intelligence (AI) based image codecs, where trained neural networks perform the encoding at the transmitter and the decoding at the receiver, can provide high quality images / videos at significantly reduced bit rates. With this technique, a plain representation of an image can be encoded by a network to output a description and / or some features, which can be decoded by a network trained to reconstruct the encoded image.
[0004] Besides significantly reduced bitrates, GANs can also produce better images compared to older codecs, in that they can create visually appealing images even when there is very little information.
[0005] However, there are also known limitations and drawbacks associated with compression using machine learning. "Hallucination" in the context of GAN coding refers to false features or details created by GANs in the encoding / decoding process. There may be several reasons for such behavior. Hallucination occurs when content is created that was not present in the original data. In many applications, this is not a problem since the differences are not significant or relevant for some uses of the image data. However, if the compressed data is to be used as evidence and / or in court or forensic investigations, it may be argued that the data generated by artificial intelligence may be of lower probative value. For example, the restored image may be contaminated with elements of the training data. Summary of the Invention
[0006] The present disclosure relates to a method for encoding images in a video, in particular to a method for encoding images using artificial intelligence and preferably performing a quality check on the encoded images before they are transmitted or stored, where the final encoded image data may include both image data encoded using AI and image data encoded using conventional codecs.
[0007] A first aspect of the present invention is a computer-implemented method for encoding one or more images in a video, the method comprising: collecting an original image from an image sensor of a video camera system; encoding an original image using the generative image model, thereby obtaining a first encoded image; decoding the first encoded image, thereby obtaining a first decoded image; Identifying or obtaining one or more regions of interest in the original image; For each region of interest, - performing an encoding quality check by comparing some reference points in the region of interest of the original image with corresponding reference points in the region of interest of the first decoded image, thereby obtaining a level of difference; - encoding the region of interest using a non-generative image model if the level of difference is greater than a threshold, thereby obtaining a non-generative encoded image area; Providing final encoded image data including: a) a non-generated encoded image area for the region of interest having a difference level greater than a threshold; and b) the first encoded image for at least a remaining portion of the original image. The present invention relates to a computer-implemented method,
[0008] It may be noted that the process of generating the final encoded image data, i.e. the image data to be transmitted or stored, itself includes a step of internally checking the quality of both the encoding and the decoding, as well as the one or more regions of interest, and then providing the final encoded image data. The one or more regions of interest may be, for example, a face or part of a face, or a person, or a particular part or detail on a building or vehicle, but in principle may be any type of item or detail for a given application.
[0009] The method may be said to be based on a discrimination between parts of the original image that are considered critical for an application or use and parts of the original image that are less important, or at least for which it may be acceptable for the application that the restored image may contain a certain level of hallucination. Thus, the method includes a step of identifying or obtaining one or more regions of interest in the original image. As explained above, the original image is both encoded to obtain a first encoded image and decoded to obtain a corresponding first decoded image. Then, for the identified one or more regions of interest, a quality check is performed in which a number of reference points in the original image are compared to corresponding reference points in the first decoded image to obtain a level of difference for each region of interest. As will be appreciated by those skilled in the art, there are more than one way of determining the "level of difference" for a region of interest. This will be explained in more detail below. Also, as will be appreciated by those skilled in the art, the expression "performing an encoding quality check" may also be seen, to some extent, as an examination of the corresponding decoding, since it is the first decoded image that is compared to the original image. A common practice in AI-generated encoding is to train an encoder network and a decoder network in tandem. Once the process of evaluating the encoding has been performed, the method will make a decision as to whether the quality is good enough to use image data generated using a generative image model, or whether encoding using a non-generative image model will have to be performed for the area. As a final step, final encoded image data is provided that includes a) non-generative encoded image areas for regions of interest having a level of difference greater than a threshold, and b) a first encoded image for at least a remainder of the original image. The final encoded image data may include a first encoded image for only the remainder of the original image, or for more than the remainder of the original image, such as for the entire original image.
[0010] One advantage of providing a flexible encoded image is that by using this setup it is possible to achieve low bit rates of AI encoding for a maximized portion of the image while minimizing the risk of introducing unacceptable hallucination for key points or areas in the image.
[0011] The present disclosure further relates to Processing Circuit An image processing system comprising: collecting an original image from an image sensor of a video camera system; encoding an original image using the generative image model, thereby obtaining a first encoded image; decoding the first encoded image, thereby obtaining a first decoded image; Identifying or obtaining one or more regions of interest in the original image; For each region of interest, - performing an encoding quality check by comparing some reference points in the region of interest of the original image with corresponding reference points in the region of interest of the first decoded image, thereby obtaining a level of difference; - encoding the region of interest using a non-generative image model if the level of difference is greater than a threshold, thereby obtaining a non-generative encoded image area; Providing final encoded image data including: a) a non-generated encoded image area for the region of interest having a difference level greater than a threshold; and b) the first encoded image for at least a remaining portion of the original image. The present invention relates to an image processing system configured to perform the above.
[0012] Those skilled in the art will recognize that the disclosed methods for encoding one or more images in a video may be implemented using any embodiment of the disclosed image processing system, and vice versa.
[0013] Various embodiments are described below with reference to the drawings, which are example embodiments and are intended to illustrate some of the features of the disclosed systems and methods for encoding images in video. [Brief description of the drawings]
[0014] [Figure 1] 1 is a flowchart of a method according to one embodiment of a method of the present disclosure for encoding one or more images in a video. [Diagram 2] FIG. 2 illustrates an example process for encoding an image according to one embodiment of the method of the present disclosure. [Diagram 3] FIG. 2 shows an example of a comparison of several reference points in a region of interest of an original image with corresponding reference points in a region of interest of a decoded image. [Figure 4] A diagram showing one embodiment of an image processing system of the present disclosure in which the method of the present disclosure for encoding one or more images in a video is performed at a transmitter side and the final encoded image data is received at a receiver side. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0015] The present disclosure relates to a method for encoding one or more images in a video. FIG. 1 shows a flowchart of a method 100 according to one embodiment of the method of the present disclosure for encoding one or more images in a video. The method may include collecting one or more original images from an image sensor of a video camera system (101). The collected images may then be encoded using a generative image model to obtain a first encoded image (102). The method may further include decoding the first encoded image to obtain a first decoded image (103). The decoding may be performed as part of the encoding process at the transmitter side. The method may further include identifying or obtaining one or more regions of interest in the original image (104). For each of the regions of interest, an encoding quality check may be performed by comparing a number of reference points in the region of interest of the original image with corresponding reference points in the region of interest of the first decoded image, thereby obtaining a level of difference for each region of interest. If the level of difference is greater than a threshold for the region of interest, the region of interest may be encoded using a non-generative image model (105). The method may further include providing final encoded image data including a) a non-generated encoded image area for the region of interest having a level of difference greater than a threshold, and b) the first encoded image for at least a remaining portion of the original image (106).
[0016] An "image model" may be said to include a functional definition (e.g., functional specification, pseudocode, formula) of an encoder and a functional definition of a decoder adapted to convert images to and from a common image data format. The image data may be digital data. The generative image model may include one or more of a machine learning based encoder, an AI based encoder, an artificial neural network encoder, a generative adversarial network (GAN) encoder, a variational autoencoder (VAE), a convolutional neural network (CNN) encoder, and a recurrent neural network (RNN) encoder. For example, if a GAN encoder is used, the first encoded image may be referred to as a GAN encoded image, and the first decoded image may be referred to as a GAN decoded image.
[0017] In contrast, using a non-generative image model may generally refer to the opposite, i.e., a more conventional way of encoding an image or video, which does not create information that was not in the original image or video, using any of machine learning based encoders, AI based encoders, artificial neural network encoders, generative adversarial network (GAN) encoders, variational autoencoder (VAE) encoders, convolutional neural network (CNN) encoders, and recurrent neural network (RNN) encoders. For example, encoding an image using a non-generative image model may be viewed as encoding the image without inserting information derived from images other than the one being encoded. If the image is a frame of a video sequence, encoding the image using a non-generative image model may be viewed as encoding the image without inserting information derived from images outside the video sequence. Alternatively, the process of encoding an image using a non-generative image model may be viewed as encoding the image without processing the image with functions that depend on information derived from images other than the image, or, if the image is a frame in a video sequence, without processing the image with functions that depend on information derived from images outside the video sequence. Examples of non-generative image models include, but are not limited to, transform coding, a combination of predictive coding (e.g., inter-frame temporal predictive coding) and transform coding (so-called hybrid coding), ITU H.26x, particularly versions of JPEG such as H.264, H.265 and H.266, AOMedia Video 1 (AV1), JPEG2000, etc. At least H.26x and AV1 may be described as image models with hybrid coding.
[0018] FIG. 2 shows an example of a process 200 for encoding an image according to an embodiment of the method of the present disclosure. In the upper left corner, an original image 300 may be used as a starting point. The original image 300 includes a house 301a, a tree 301b, and a rabbit 301c, which represents an object of interest. By analogy, the rabbit may be a human, a human face, or any other item of interest in an application. Starting from the original image 300, two parallel tasks are then performed. In this regard, it should be understood that the system may comprise standard hardware blocks for performing various image processing tasks on the original image 300. For example, the system may comprise an image processing pipeline (IPP) for applying different filters, such as noise filtering, segmentation, etc. In this regard, the "original" image is not necessarily to be seen as an unprocessed image for the method of the present disclosure, but as a starting point and input. It should be understood that a camera has an image sensor that captures light, and the image sensor creates a signal that is processed to create an original image. The tasks of the method of the present disclosure do not need to be performed in a particular order and may be performed in any suitable order. Preferably, the entire original image 300 is encoded using the generative image model 201. The encoded image is then decoded using the generative image model decoder 202. A first decoded image is thereby obtained. In FIG. 2, it can be seen that the first decoded image is used in the encoding quality checker 204. The original image 300 may also be used in a process 203 that identifies one or more regions of interest 302 in the original image 300. In the step of performing the encoding quality check 204, the identified region of interest 302a of the original image is compared with the region of interest 302b of the first decoded image. The steps of the method may be ordered such that the identification of one or more regions of interest in the original image is performed before the first decoded image is sent for quality check. In the example of Fig. 2, it can be seen that the rabbit's "face" has been identified as the region of interest. In the quality check, some reference points in the region of interest of the original image are then compared with corresponding reference points in the region of interest of the first decoded image. This is further explained and illustrated in Fig. 3.If the difference between the original image 302a and the first decoded image 302b is too large in this respect (the level of difference exceeds a threshold), the region of interest is in fact coded using a non-generative image model, which may be a conventional codec. In a step 205 of providing / generating final coded image data 303, the final coded image data is prepared such that it includes non-generative coded image areas for the region of interest having a level of difference larger than the threshold and the first coded image for at least the remaining part of the original image, i.e. potentially including a mixture of image data that has been generated using a generative image model and a non-generative image model.
[0019] Identifying one or more regions of interest in the original image may include detecting an object in the original image. Detecting the object may include applying a machine learning model, such as a neural network, trained to detect the object. Alternatively, or in combination, identifying one or more regions of interest in the original image may include defining one or more regions of interest based on the location of the detected object and / or other context information, such as what the object is doing and / or what is happening in the scene. Identifying one or more regions of interest in the original image is not limited to a particular manner of identifying a region or object and in fact includes several options. This may involve, for example, an additional neural network trained to perform such a task. As a non-limiting example, a semantic segmentation algorithm may detect in which region(s) in an image grass is present and determine that for an application, a car on the grass is an object of interest, but a car on a road is not considered to be of interest.
[0020] Object detection is a computer vision technique that involves locating an object in an image frame and identifying the object. Those skilled in the art will generally be familiar with such techniques and know how to implement them. Convolutional Neural Networks (CNNs) or other machine learning based methods are popular because they are generally very accurate and fast, but there are several other object detection techniques that do not rely on CNNs or machine learning.
[0021] One specific, non-limiting example of a more old-fashioned class of object detection algorithms is the Viola-Jones detection framework, in which an image frame is scanned with a sliding window and each region is classified as containing or not containing an object. The method uses Haar features and cascade classifiers to detect objects.
[0022] When an object is detected, a set of signatures may be created to describe the visual appearance of the detected object. Image data from a single image frame or a video sequence may be used to create signatures for the detected object. Various image and / or video analysis algorithms may be used to extract and create signatures from the image data. Examples of such image or video analysis algorithms are various algorithms for extracting features in a face, such as those in Turk, Matthew A., and Alex P. Pentland. "Face recognition using eigenfaces." Computer Vision and Pattern Recognition, 1991. Proceedings CVPR'91., IEEE Computer Society Conference on. IEEE, 1991; gait features, such as those in Lee, Lily, and W. Eric L. Grimson. "Gait analysis for recognition and classification." Automatic Face and Gesture Recognition. 2002. Proceedings. Fifth IEEE International Conference on. IEEE, 2002; or color, such as those in US 8,472,714 by Brogren et al.
[0023] The database may include some objects and some distinguishing characteristics. In the method of the present disclosure, identifying one or more regions of interest in the original image may include matching the distinguishing characteristics with distinguishing characteristics in the database to classify the object as a type of object, for example, a car, a person, or any other item. In one embodiment, detecting an object in the image frame includes comparing the image frame with a reference image in the database to match a feature corresponding to the object in the image frame with a feature corresponding to the object in the reference image.
[0024] Classification of objects can be achieved by neural networks. Classification neural networks are often used in applications such as character recognition, monitoring, surveillance, image analysis, natural language processing, etc. There are many neural network algorithms / techniques that can be used to classify objects, e.g., convolutional neural networks, recurrent neural networks, etc.
[0025] FIG. 3 shows an example of a comparison of some reference points 304 in a region of interest 302a of an original image with corresponding reference points 304 in a region of interest 302b of a decoded image. There are several options and ways to implement the coding quality check of the region of interest. The coding quality check is not limited to the example shown in FIG. 3. One option is to identify some key reference points for a given application. For example, in FIG. 3 it can be seen that some characteristic points of a "face" have been identified. Thus, an embodiment of the disclosed method of coding one or more images in a video may include identifying characteristics for a given type of content in the region of interest and selecting reference points in the identified characteristics. The identified characteristics may be any type item or feature, for example, nose and / or eyes or ears of a face, and / or characteristics of a person, or characteristics of a vehicle, such as a license plate of a car, and / or characteristics of a building, and / or characteristics of an animal. The actual comparison may be, for example, a pixel-by-pixel comparison for a selection of pixels.
[0026] In the illustrative example of Fig. 3, several reference points have been identified. The identification of the reference points may be done by applying any machine learning model trained to recognize several characteristics of the image. A person skilled in the art will generally be able to implement or use such a model.
[0027] Another option, which may be used independently or in combination with the option of identifying specific reference points, is to divide the region of interest into a grid, for example as a grid with cells 305 according to the example shown in FIG. 3. The reference points may then be placed in the region of interest to form a grid. The grid may be a grid with uniformly sized and placed cells, or a grid with non-uniform cell sizes. For example, the method may include using a function to create a grid with larger cells in areas that may be categorized as background and smaller cells in areas that may be categorized as foreground or that contain certain items or characteristics. For example, there are known methods that can distinguish between background and foreground in an image. Such methods would be readily available to one skilled in the art. In general, the reference points may be selected such that the distance between the reference points in the grid does not exceed a predetermined limit. In the example in FIG. 3, the cells do not need to be uniformly sized. For example, there may be smaller cells in areas with more detail and / or more local variation and / or based on local content.
[0028] As mentioned above, the selection of the reference point may be a combination of identifying characteristics for a given type of content in the region of interest and dividing the region of interest into a grid of cells. In such a combination, it is possible to have cells with smaller sizes in sub-regions close to the identified characteristics and cells with larger sizes in the remainder of the region of interest.
[0029] The disclosed method of encoding one or more images in a video may include performing an encoding quality check by comparing some reference points in a region of interest of an original image with corresponding reference points in a region of interest of a first decoded image, thereby obtaining a level of difference. The "level of difference" is introduced as a way of determining whether the encoding region of interest is close enough to the original image. The level of difference may be defined and measured in several ways. For example, pixels within a certain distance from a point, or pixels in a certain cell of the above-mentioned grid, may be compared between the region of interest of the original image and the region of interest of the first decoded image. This may be done pixel by pixel, or by grouping and averaging pixels. As will be understood by those skilled in the art, different techniques and approaches may be used. According to one embodiment of the disclosed method, a predefined measure of difference or dissimilarity may be used. One example is to calculate or extract the average pixel value for pixels within a certain distance from a point, or pixels in a certain cell. If the difference between the values is larger than a certain threshold, the quality of encoding the region of interest using the generative image model may be considered too low for that region of interest. Therefore, that particular region of interest may then be coded using a non-generative image model. Another method would be a pixel-to-pixel comparison. A sum of pixel differences by accumulating pixel differences, or a sum of squared differences, may be further alternatives. A further option may be a comparison of histograms of the cells of the grid. In general, the step of obtaining the level of differences is not limited to the specific examples provided in this disclosure.
[0030] Moreover, in one embodiment of the disclosed method of encoding one or more images in a video, at least one of the regions of interest of the original image is composed of several image parameters, such as color and luminance, and the step of performing the encoding quality check is performed only for a subset of the image parameters, such as only for luminance. In some applications, some image parameters may be irrelevant or of little importance. As an example, it may be important that the shape or other details of a face or person in the encoded image do not deviate too much from the original image, but other parameters, such as color, may not be as significant.
[0031] The disclosed method of encoding one or more images in a video may include generating a summed encoding quality score for each region of interest. This information may be added to the final encoded image data for each region of interest. By adding the encoding quality scores for the regions of interest, the receiver will receive not only a certain minimum quality encoding, but also more detailed information on how similar it is to the original image. For example, if the level of similarity is determined to be very high, the compressed data may ultimately be admissible as evidence and / or in court or forensic investigation.
[0032] The final encoded image data may include a) non-generative encoded image areas for regions of interest having a level of difference greater than a threshold, and b) a first encoded image for at least a remaining portion of the original image. Since the first encoded image, i.e., an image encoded using a generative image model, will generally be more compressed, it is generally desirable to have as much of the first encoded image as possible in the final encoded image data. However, in areas where the difference between the first encoded image and the original image is too significant, the final encoded image may include a mixture of image data that has been generated using generative and non-generative image models. Thus, the step of decoding the final encoded image data may include decoding the non-generative encoded image areas without relying on information derived from images other than the non-generative encoded image areas, and decoding the first encoded image for at least the remaining portion using a machine learning model, such as a generative adversarial network. The final encoded image data may include the first encoded image for only the remaining portion of the original image, or for more than the remaining portion of the original image, such as for the entire original image. If the final encoded image data includes the first encoded image for the entire image and non-generated encoded image areas for regions of interest having a difference level greater than a threshold, it may be the task of the receiver to select the data to be used in further processing. In that case, the receiver will generally take into account the information specifying the regions of interest.
[0033] The present disclosure further relates to a computer program having instructions that, when executed by a computing device or computing system, cause the computing device or computing system to perform any embodiment of the disclosed method of encoding one or more images in a video. The computer program may be stored in any suitable type of storage medium, such as a non-transitory storage medium.
[0034] The present disclosure further relates to Processing Circuit An image processing system comprising: collecting an original image from an image sensor of a video camera system; encoding an original image using the generative image model, thereby obtaining a first encoded image; decoding the first encoded image, thereby obtaining a first decoded image; Identifying or obtaining one or more regions of interest in the original image; For each region of interest, - performing an encoding quality check by comparing some reference points in the region of interest of the original image with corresponding reference points in the region of interest of the first decoded image, thereby obtaining a level of difference; - encoding the region of interest using a non-generative image model if the level of difference is greater than a threshold, thereby obtaining a non-generative encoded image area; Providing final encoded image data including: a) a non-generated encoded image area for a region of interest having a difference level greater than a threshold value and the first encoded image; and b) the first encoded image for at least a remaining portion of the original image. The present invention relates to an image processing system configured to perform the above.
[0035] The processing circuit, or a further processing circuit, may be configured to decode the final encoded image data at the receiver side. It may therefore be noted that encoding may include both encoding and decoding steps to verify and / or improve the final encoded image data before it is transmitted or stored. Figure 4 shows an example of such an embodiment. An image processing system 500 comprises a processing circuit 501 configured to collect images from a camera system 502. The image processing system 500 in this setup may be seen as a transmitter side of the system, which may further comprise a receiver side system 400, which has a processing circuit 401 configured to receive and decode the transmitted compressed final encoded image data. Both the steps of encoding the original image using the generative image model and decoding the first encoded image as part of the encoding quality check may be performed at the transmitter side prior to communicating the final encoded image data to the receiver side.
[0036] The image processing system may further comprise peripheral components, such as one or more memories that may be used to store instructions that may be executed by any of the processors. The one or more memories may include random access memory (RAM) and / or read only memory (ROM), or any suitable type of memory. The system may further comprise any of internal and external network interfaces, input and / or output ports, modules for wirelessly sending and receiving data, and communication interfaces that allow software and / or data to be transferred between the system and external devices. The software and / or data transferred via the communication interface may be in any suitable form of electrical, optical, or RF signals. As will be appreciated by those skilled in the art, the processing circuit may be a single processor in a multi-core / multi-processor system.
Claims
1. 1. A computer-implemented method for encoding one or more images in a video, the method comprising: collecting an original image from an image sensor of a video camera system; encoding the original image using a generative image model, thereby obtaining a first encoded image; decoding the first encoded image, thereby obtaining a first decoded image; identifying or obtaining one or more regions of interest in the original image; For each of the regions of interest, - identifying characteristics for a given type of content in said region of interest and selecting a number of reference points in said identified characteristics; and / or Dividing the region of interest into a grid of cells, wherein a number of reference points are positioned in the region of interest to form the grid of cells; performing an encoding quality check by comparing said reference points in said region of interest of said original image with corresponding reference points in said region of interest of said first decoded image, thereby obtaining a level of difference; - if said level of difference is greater than a threshold, encoding said region of interest using a non-generative image model, thereby obtaining a non-generative encoded image area; providing final encoded image data including: a) the non-generated encoded image areas for the region of interest having a difference level greater than the threshold; and b) the first encoded image for at least a remaining portion of the original image.
4. A computer-implemented method comprising:
2. 2. The method of claim 1 , wherein the generative image model comprises a generative adversarial network (GAN) and / or a variational autoencoder (VAE) encoder, the first encoded image is a GAN-encoded image and / or a VAE-encoded image, and the first decoded image is a GAN-decoded image and / or a VAE-decoded image.
3. The method of claim 1 , wherein the reference points are selected such that the distance between the reference points in the grid does not exceed a predetermined limit.
4. The method of claim 1 , wherein the identified features include a nose and / or an eye or an ear on a face, and / or a feature of a person, and / or the identified features include a feature of a vehicle, such as a license plate on a car, and / or a feature of a building, and / or a feature of an animal.
5. 2. The method of claim 1 , wherein the step of identifying one or more regions of interest in the original image comprises detecting an object in the original image, preferably wherein the step of detecting an object comprises applying a machine learning model, such as a neural network, trained to detect the object.
6. The method of claim 1 , wherein the step of identifying one or more regions of interest in the original image comprises defining the one or more regions of interest based on a location of a detected object.
7. 2. The method of claim 1, wherein at least one of the regions of interest of the original image is composed of several image parameters, such as color and luminance, and the step of performing an encoding quality check is performed for only a subset of the image parameters, such as only for the luminance.
8. 2. The method of claim 1 , wherein the step of performing an encoding quality check comprises generating a summed encoding quality score for each region of interest, and preferably further comprises the step of adding the summed encoding score for each region of interest to the final encoded image data.
9. 2. The method of claim 1, wherein both the steps of encoding the original image using a generative image model and the step of decoding the first encoded image are performed at a transmitter side prior to communicating the final encoded image data to a receiver side, preferably further comprising the step of receiving and decoding the final encoded image data at a receiver side.
10. 10. The method of claim 9, wherein the step of decoding the final encoded image data includes decoding the non-generated encoded image areas without relying on information derived from images other than the non-generated encoded image areas, and decoding the first encoded image for at least the remaining portion using a machine learning model, such as a generative adversarial network.
11. A computer program having instructions which, when executed by a computing device or computing system, cause the computing device or computing system to perform the method for encoding one or more images in a video according to claim 1.
12. Processing Circuit The image processing system includes: collecting an original image from an image sensor of a video camera system; encoding the original image using a generative image model, thereby obtaining a first encoded image; decoding the first encoded image, thereby obtaining a first decoded image; identifying or obtaining one or more regions of interest in the original image; For each of the regions of interest, - identifying characteristics for a given type of content in said region of interest and selecting a number of reference points in said identified characteristics; and / or - dividing said region of interest into a grid of cells, a number of reference points being placed in said region of interest so as to form said grid of cells; performing an encoding quality check by comparing said reference points in said region of interest of said original image with corresponding reference points in said region of interest of said first decoded image, thereby obtaining a level of difference; - if said level of difference is greater than a threshold, encoding said region of interest using a non-generative image model, thereby obtaining a non-generative encoded image area; providing final encoded image data including: a) the non-generated encoded image areas for the region of interest having a difference level greater than the threshold; and b) the first encoded image for at least a remaining portion of the original image. An image processing system configured to:
13. 13. The image processing system of claim 12, wherein the processing circuitry or further processing circuitry is configured to decode the final encoded image data at a receiver side.