Discovery of Semantic Target Regions in Images

By employing an automatic object detection system and neural network model to identify and prioritize semantic target regions within images, the method enhances image processing efficiency and quality, addressing the limitations of current systems that fail to distinguish between semantically important and unimportant image regions.

JP7700180B2Active Publication Date: 2025-06-30AVID TECHNOLOGY INC
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2023120395
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2022-08-10
Filing Date
2023-07-25
Publication Date
2025-06-30
Estimated Expiration
2043-07-25

AI Technical Summary

Technical Problem

Current image processing systems lack the ability to distinguish between regions of an image based on their semantic importance, leading to inefficient processing of images and videos.

Method used

A method that uses an automatic object detection system and a trained neural network model to identify semantic target regions within an image by re-dividing the image into sub-images, generating image embeddings, and determining the similarity between sub-image and source image embeddings to assign semantic interest levels.

Benefits of technology

This approach enables more efficient image processing tasks such as video compression, automatic panning, image cropping, and color correction by prioritizing semantically important regions, resulting in improved image quality and processing efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007700180000001
    Figure 0007700180000001
  • Figure 0007700180000002
    Figure 0007700180000002
  • Figure 0007700180000003
    Figure 0007700180000003
Patent Text Reader

Abstract

To provide a method and a system that do not distinguish between regions of an image based on their semantic importance.SOLUTION: The method automatically detects an object, extracts a sub-image indicating an extent of each object from an image, and determines image embedding for each sub-image as well as an entire image by using an image encoder implemented as a trained multi-modal neural network. Similarity between image embedding of the sub-image and image embedding of the entire image is used as a measure of semantic importance of the object depicted in the sub-image. An object of high semantic importance includes a semantic object region in the image. Knowledge of such a region is used to improve and enhance efficiency of a downstream image processing task such as image compression, panning, scanning, and contrast enhancement.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

Prior Art

[0001]

[0001] Images processed by a computing system contain a specific region that includes the most important items within that image, while the remaining regions generally fill in the frame without significantly adding to the semantic content of the image. For example, it is most likely that an image of an important person includes a portion that represents the person's face and some or all of their body, but the surrounding region may indicate the location where the image was captured. In other examples, if the image represents a scene from a sports game, the most interesting region portrays the main action captured by the image, such as the player who scored a goal in soccer or the player who served the ball in tennis.

[0002]

[0002] When performing image processing tasks on such images, current systems do not make distinctions between regions of the image based on their semantic importance. Instead, they manipulate the technical characteristics of the image, such as contrast, sharpness, and saturation, to generate an overall best result across the entire image. However, if an image processing system could access information that identifies the target regions of the image being processed, it would be able to execute image processing tasks more efficiently.

Summary of the Invention

Problems to be Solved by the Invention

[0003]

[0003] Therefore, it would be advantageous if a computer system could be made to discover important objects within an image or video stream and perform tasks such as video compression, automatic panning and scanning image cropping, and automatic color correction more efficiently.

Means for Solving the Problems

[0004]

[0004] Generally, in the first aspect, a method for determining a semantic target region in a source image includes receiving the source image, detecting a plurality of objects in the source image using an automatic object detection system, re-dividing the source image into a plurality of sub-images, where each sub-image includes a portion that includes one of the plurality of detected objects in the source image, generating an image embedding for the source image using a trained neural network model, generating an image embedding for each sub-image in the plurality of sub-images, determining a similarity between the image embedding of each sub-image and the image embedding of the source image for each sub-image in the plurality of sub-images, assigning a semantic interest to the detected object included in the sub-image according to the determined similarity between the image embedding of the sub-image corresponding to the detected object and the image embedding of the source image, and outputting an indication of the semantic interest assigned to the detected object included in the sub-image.

[0005]

[0005] Various embodiments include one or more of the following features. The automatic object detection system is a trained neural network model. The trained neural network model used to generate the image embedding is a multimodal neural network. For each of the plurality of detected objects, the method includes generating an object mask for the detected object and generating an object mask image of the source image, wherein each detected object among the plurality of detected objects is replaced in the source image by the shaded silhouette of the object mask generated for that object, and applying a visual indication to each shaded silhouette, the visual indication indicating the semantic interest level assigned to the detected object corresponding to the object mask. The indication of the semantic interest level assigned to each detected object among the plurality of detected objects is used to enhance the image processing of the source image. The image processing includes image compression, and the process of enhancing the image processing of the source image includes varying the number of bits assigned when compressing each sub-image among the plurality of sub-images according to the semantic interest level assigned to the detected object corresponding to the sub-image. The source image is one frame in a video stream. The image processing includes cropping a portion of the source image to achieve a desired aspect ratio of the source image, and the process of enhancing the image processing of the source image includes preferentially retaining objects with a higher assigned semantic interest level within the cropped region of the source image. The objects preferentially retained within the cropped image include the objects assigned the highest semantic interest level.This method further includes the steps of selecting a subset of detected objects from a plurality of detected objects, where the selected subset of objects includes a set of objects assigned a high semantic interest level; determining the centroid of this subset of detected objects within the source image; and cropping the source image such that the centroid of the subset of detected objects within the source image is located at the center of the cropped image. The source image is one frame of a video stream. The image processing includes contrast enhancement, and the contrast enhancement includes a boosting process that increases the contrast in regions of the source image that contain detected objects assigned a high semantic interest level. The source image is one frame within a video stream.

[0006]

[0006] Generally, in other aspects, a computer program product comprises a non-transitory computer-readable medium encoded with computer-readable instructions that, when processed by a processing device, instruct the processing device to execute a method for determining a semantic target region in a source image. The method includes receiving the source image, detecting a plurality of objects in the source image using an automatic object detection system, re-dividing the source image into a plurality of sub-images, each sub-image including a portion of the source image that includes one of the detected plurality of objects, generating an image embedding for the source image using a trained neural network model, generating an image embedding for each sub-image among the plurality of sub-images, determining a similarity between the image embedding of each sub-image and the image embedding of the source image among the plurality of sub-images, assigning a semantic interest degree to the detected object included in the sub-image according to the determined similarity between the image embedding of the sub-image corresponding to the detected object and the image embedding of the source image, and outputting an indication of the semantic interest degree assigned to the detected object included in the sub-image.

[0007]

[0007] Generally, in yet another aspect, the system includes a memory storing computer-readable instructions and a processor connected to the memory. When the processor executes the computer-readable instructions, the system is caused to execute a method for determining a semantic target region in a source image. The method includes receiving a source image, detecting a plurality of objects in the source image using an automatic object detection system, re-dividing the source image into a plurality of sub-images, each sub-image including a portion of the source image that includes one of the detected plurality of objects, generating an image embedding for the source image using a trained neural network model, generating an image embedding for each sub-image in the plurality of sub-images, determining a similarity between the image embedding of each sub-image and the image embedding of the source image for each sub-image in the plurality of sub-images, assigning a semantic interest level to the detected object included in the sub-image according to the determined similarity between the image embedding of the sub-image corresponding to the detected object and the image embedding of the source image, and outputting an indication of the semantic interest level assigned to the detected object included in the sub-image.

Brief Description of the Drawings

[0008]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

DETAILED DESCRIPTION OF THE INVENTION

[0009]

[0017] Between the late 2010s and early 2020s, artificial intelligence (AI) and machine learning (ML) computer systems have rapidly evolved. A common form of machine learning computer system uses a neural network model. Some such neural network models have been developed and trained to detect objects within an image. The methods described herein use such an object detection system in combination with an ML-based image encoder to identify semantically interesting regions within the image. As used herein, a region of an image is considered semantically interesting if it includes anything that the person creating and sharing the image might consider significant about the image. For example, if a video contains an interview of two people, the semantically interesting parts of the scene would be the two people, not the plants in the background. Also, in the case of a cooking show, it would be the pan on the stove that the chef is actually using. In certain cases, this region may also represent the content that warranted taking the image.

[0010]

[0018] Figure 1 is a high-level flowchart showing various steps involved in discovering a semantic target region within an image. The source image 102 for which the semantic target region is to be discovered is processed by a system that performs object detection 104 within this image. Object detection may be performed by a trained ML model or by an algorithmic method such as edge detection or contrast analysis. Examples of suitable object detection systems include the following: Mask DINO by Feng Li.et al., arXiv:2206.02777; Swin Transformer by Ze Liu et al., arXiv:2103.14060; Generic RoI Extractor by Leonardo Rossi et al., arXiv:2004.13665[cs.CV]; and Mask R-CNN by Kaiming He et al., arXiv:1703.06870[cs.CV]. Such systems typically run at 10 frames per second and are used to detect objects within a video stream. When new frames are received approximately 30 times per second, only specific images are analyzed. The images to be analyzed can be arbitrarily selected, such as every 10 frames, every 15 frames, every 20 frames, or can be selected based on a change in the scene when the overall composition of the image changes. The interval between the frames to be analyzed can be automatically determined based on the processing time required to perform object recognition on a given computing platform. The object detection system is effective in identifying various objects it is trained to find within the image. The output from the object detection system includes, for each detected object, the class of the detected object, the probability score that the object belongs to that class, a bounding box defined by the leftmost, top, rightmost, and bottom coordinates of the object mask, and the mask. The mask consists of a binary image of the size of the bounding box indicating which pixels belong to that object.In the method described herein, object classes and scores are discarded, and the object mask is used to create a "cutout" image of object silhouettes, as shown at 118 in FIG. 1. When the object detection system encounters an unknown object, it can identify the presence of the object, but it cannot determine what the object is. In this situation, the system outputs low probability scores across all possible objects.

[0011]

[0019] The system that executes object detection 104 outputs a sub-image 106 of the source image depicting the objects detected by this system. For example, in an image showing multiple people, this system can detect each of the people in the frame. This is shown in FIG. 2. In FIG. 2, an image 202 of world leaders is supplied to the object detection system, and the image portion containing the detected objects is surrounded by boxes in image 204. The object detection system finds all of these people, and the corresponding object masks shown in the object mask image 206 spread across the scene, including these people in the background. This example illustrates how the object detection system itself cannot distinguish between the main objects in the image and the objects of low importance, i.e., peripheral objects.

[0012]

[0020] To generate an object embedding, an image encoder is deployed. As used herein, an image-encoder refers to a multimodal neural network trained to encode images and text into a compatible vector space. Such a vector space is referred to as latent spaces, in which points in the space have coordinates such that "similar" points are closer to each other within the space. The definition of similarity is not explicit. The determination of the embedding is done by training a neural network model on a large set of images whose similarity to each other is known. By encoding an image, a vector representing the semantic embedding of that image is generated in this vector space. Latent spaces and semantic embeddings are well known in the fields of machine learning and neural networks. Examples of image encoders include the following. The Contrastive Language-Image Pretraining system (CLIP), available from OpenAI, San Francisco, CA. "Learning Transferable Visual Models From Natural Language Supervision" by Radford, A. et al., arXiv:2103.00020v1, which is hereby incorporated by reference in its entirety. The Language Interpretability Tool (LiT) from Google Research, Mountain View, CA."The Language Interpretability Tool: Extensible, Interactive Visualizations and Analysis for NLP Models" by Tenney, A. et al., described in arXiv:2008.05122v1, is hereby incorporated by reference in its entirety. Vision - Object - Semantics Aligned Pre - training (Oscar) for language tasks from Microsoft Corp., Redmond, Washington. "Oscar: Object - Semantics Aligned Pre - training for Vision - Language Tasks" by Li, X. et al., described in arXiv:2004.06165, is hereby incorporated by reference in its entirety.

[0013]

[0021] In the method described herein, the entire source image is encoded by an image encoder (108) to generate an image embedding for the entire image (110). This embedding can be expected to appear in latent space coordinates close to other "similar" images, i.e., images depicting similar content or having a similar intent. In a trained neural network image encoder model, images that are close to each other in the embedding space reflect the similarity of the text captions for images similar to them in the training data set. Commonly used similarity measures assign 100% similarity to two identical images, 10% similarity to two completely dissimilar images, and more than 70% similarity to images considered semantically similar. The similarity between two images can be determined as a multi-dimensional distance metric or as a cosine metric, as described in G. Salton and C. Buckley, “Term-weighting Approaches in Automatic Text Retrieval”, Information Processing and Management, 1;24(5):513-23, Jan. 1988. By citing this document here, the entire content thereof is hereby incorporated into this application.

[0014]

[0022] Each member of a set of sub-images 106 depicting the detected object is also encoded (112) by an image encoder, and an image embedding 114 is generated for each of the sub-images 106. In the next step, a similarity determination 116 is performed, in which the embedding 110 of the entire image is compared with each of the object embeddings 114 to determine their relative similarity. The degree of similarity of each object embedding to the embedding of the entire image is used as a measure of the semantic interest of the object represented in the corresponding image portion. In various embodiments, this similarity is represented as a set of semantic interest weights, and the image portion having the embedding most similar to that of the embedding of the entire image has the highest weight. In the object embedding latent space, the degree of similarity between two embeddings corresponds to the multi-dimensional distance between these embeddings. As used herein, an object that is semantically important can also be referred to as an object having high semantic saliency, and these terms are used interchangeably herein. An object having a high semantic importance, i.e., a saliency score, is important for the entire scene and includes the semantic target region of the image.

[0015]

[0023] The semantic weight of each of the detected objects can be represented graphically by generating an object mask image 118, and the area corresponding to each of the object masks is color-coded according to the semantic weight assigned to the corresponding object. In various embodiments, as shown in image 118, the mask is assigned a shade of grey, and the shade assigned to the object mask indicates the semantic importance of the corresponding object. In the example of the image of world leaders shown in FIG. 2, the results of the method described are shown as a grey-scale-shaded object mask image 208, and the object with the highest semantic importance is considered to be President Biden, shown by the object mask 214 with the lightest shade, the second most important is Prime Minister Justin Trudeau of Canada, with the next lighter object mask 212, and German Chancellor Olaf Scholz has the third lightest object mask 210. In certain embodiments, the object mask image can also use text such as semantic weights, or icons indicating the semantic importance determined for each of the objects.

[0016]

[0024] Once object masks for a given source image are generated and ordered according to the saliency of their corresponding objects, the masks and their respective semantic scores are associated with the source image. The semantic scores for each of the object masks of the source image and their association with their corresponding objects can be implemented by including this data within the source image metadata. The source image metadata can include, for each mask, the mask location, the average pixel value, the pixel value standard deviation, and a low-resolution map. In one embodiment, the data is stored in a semantic database and the semantic information is keyed to the source imagery. In other embodiments, the images corresponding to each source image are segmented and stored, with each segment corresponding to an object mask and its semantic score.

[0017]

[0025] The determination of the semantic target area can be used to optimize various image processing tasks, such as video compression, image format conversion, and color correction, as described below. To facilitate the optimization process, the source image 102 and the object mask image 118 with the semantic importance of each mask tagged, or the object mask image 118 graded according to the semantic importance, are input into a system that executes image processing 120.

[0018]

[0026] Figures 3, 4, and 5 show examples of how the system described above operates. In each case, these figures show a source image, an image with frames superimposed to show the portion of the image containing the objects detected by the object detection system, an image showing an object mask corresponding to the spatial extent of the detected objects, and an object mask graded according to the semantic interest level determined for each of the detected objects, where the brighter the gradation, the higher the importance of the object. In the example shown in Figure 3, the original image 302 shows a scene where participants in a cross-country motorcycle race are crowded, and the leading participant partially obscures the participants behind. In such a crowded scene, the object detection system has identified many of the participants, as shown in image 304 with an overlay of the detected objects, and thus supplies many candidate sub-images that make up the target candidate region to the image encoder. The corresponding object mask is shown in white in image 306 and is graded according to their semantic interest level in image 308. As can be seen in image 308, the system has determined that the motorcycle of the person leading the race has the highest interest level, followed by a competing motorcycle very close to the right of the leader. This system has succeeded in distinguishing the riders from their motorcycles, and the importance assigned to the motorcycles is higher. In Figure 4, the source image 402 of a dog talent show includes the dog owner, her dog, and various people in the background. As shown in images 404 and 406, each of these people has been identified by the object detection system. The gradation according to the semantic importance in image 408 shows that the system has succeeded in identifying that the most important aspect of this image is the dog. In Figure 5, the source image 502 features a modern sculpture. The shape of this modern sculpture may be completely different from any of the objects that were characteristic in the set of images used to train the image encoder.Nevertheless, as shown in the detected object image 504, object mask image 506, and the image of the object mask with gradation according to the semantic importance 508, the system identified this modern sculpture as having the highest semantic importance.

[0019]

[0027] Next, examples will be described of how the semantically target regions automatically determined within the image can be used to improve and / or make more efficient various aspects of video editing. These include, for example, improving composition by zooming and panning and scanning, enhancing clarity by, for example, color correction and color emphasis, and improving the quality of the entire image by using, for example, adaptive compression.

[0020]

[0028] In one application, the described method is used to improve the efficiency and quality of image compression, including both video compression and still image compression. In video compression, by default, the encoder treats all microblocks in the image equally, and one microblock typically consists of a 16×16 pixel block. By using the determination of the semantic target area described above, the encoder can allocate compressed bits according to the weighting of the semantic target area. This has the effect of improving the quality of the semantically important areas within the image by sacrificing an increase in artifacts in areas of low importance across the scene and reducing compression artifacts in important areas. For example, in an image containing the face of one or more persons, the system can assign a high semantic weight to the facial expression. Then, by appropriately configuring the video compression system, the facial expression can be more faithfully compressed and decompressed, and the compression artifacts in the most important face areas can be reduced or eliminated. One measure of the quality of a compressed image is the bits per pixel (BPP). Here, the number of bits is the total number of bits in the compressed image, including the chrominance component, and the number of pixels is the number of samples in the luminance component. An image with a BPP value of 0.25 to 0.5 is considered to have moderate quality, an image with BPP = 0.5 to 0.75 is considered to have very good quality, an image with BPP = 0.75 to 1.5 is considered to have excellent quality, and an image with BPP = 1.5 to 2.0 can, in most cases, not be distinguished from the original image. For example, in an image having a semantic target area that includes 10% of the image pixels, to achieve an overall compressed image with BPP = 0.75, the BPP for the semantic target area can be set to 2.0 and the rest of the image can be set to 0.61 so that the average is 0.75 BPP.

[0021]

[0029] In a typical workflow, video compression is performed by a video editing application such as Avid® Media Composer®, a product of Avid Technology, Inc. in Burlington, Massachusetts. FIG. 6 is a schematic screenshot of a user interface dialog box 602 for media creation settings within a video editing application. This figure shows an option 604 for selecting optimization of compression based on semantic target regions when importing video into the video editing application. A similar dialog box may also be made available to the user when the capture tab 606 is selected to perform media capture, and when the mixdown and transcode tab 608 is selected to perform mixing down and transcoding of media.

[0022]

[0030] In other applications, automatic determination of semantic target regions is used when determining how to crop an image when changing the aspect ratio. This process is commonly referred to as panning and scan. Such a change in aspect ratio is often required when importing an image or video clip into a media editing application such as a non-linear video editing system and editing it for output onto a platform having a display with an aspect ratio different from that of the source image or video. For example, a video edited with a 16×9 aspect ratio (landscape) may be exported with a 9×16 aspect ratio (portrait) for playback on a particular platform such as a smartphone. In this case, it is necessary to crop the material on one or both of the left and right edges. The choice of what to crop is guided by the determination of the semantic target region to ensure that the most important objects in the image are not lost during the cropping process.

[0023]

[0031] In various embodiments, the system can track, for example, the object with the largest semantic weight, ensure that it is retained within the cropped image, or ensure that it is placed as close as possible to the center of the image along the horizontal dimension, so that as the most important object moves horizontally across the image, it can be chased. In other embodiments, the system attempts to maximize the retention of several of the most semantically prominent objects, such as the top two, three, or four objects. In various other embodiments, the system uses a threshold of semantic importance, for example, 70%, tracks all objects that meet or exceed this threshold, and centers the cropped image on the centroid of the horizontal positions of these objects. If not all of the objects selected according to semantic importance can be retained within the cropped image, the system can attempt to retain a subset of these objects, prioritizing the objects with the highest semantic salience. The system can pan the image to optimize the positioning of important objects. The amount of cropping is determined by the resolution of the original image and the resolution of the target display. For example, in a photo of a man playing frisbee with a dog, the system would likely attempt to keep the man, the dog, and the frisbee within the cropped frame. However, if all three prominent objects cannot be shown, only the dog and the frisbee can be chased and the man cropped out.

[0024]

[0032] FIG. 7 is a schematic screen shot of a user interface dialog box 702 for importing setting values inside a video editing application. This figure shows an option 704 for selecting an optimization to crop an image into a format having different aspect ratios so as to include a meaningful target area within the cropped image. This process is typically performed when it is necessary to convert a video or still image of one format into a second format for output to a device having a display screen with an aspect ratio different from that of the source image.

[0025]

[0033] Using automatic determination of semantic target areas, color enhancement can also be automated, either partially or fully. The system analyzes pixel values in the identified target areas and calculates the mean and standard deviation of the RGB values. When performing color correction, curves in a lookup table are used to increase the slope of the contrast for values within the semantic target area. Color enhancement can be performed on videos and still images. When enhancing a video, to prevent abrupt changes that may be unpleasant for the viewer, the change in the enhancement parameters applied is smoothed between consecutive frames. This smoothing can be specified, for example, with respect to the percentage or absolute change in the parameter values for the slope of the contrast between consecutive frames or between consecutive groups of frames. In various embodiments, the system automatically increases the contrast in image areas deemed important by the determination of the semantic target area. FIG. 8 is a schematic screenshot of a user interface dialog box 802 for color correction mode settings within a video editing application. This interface includes an option box 804 that enables the user to increase the contrast in the semantic target area within the edited image when performing color correction on the image. The contrast enhancement may be applied only to objects that meet or exceed a threshold of semantic importance. The threshold may be set absolutely or may be defined by a limit on the number of objects processed in a given frame. For example, the system can automatically enhance the contrast of the top 2 to 4 objects ordered by saliency. As a result, details in important parts of the image become clearer.

[0026]

[0034] Figure 9 is a high-level block diagram of a system that performs the determination of semantic target regions in an image and uses this determination to improve image processing. The media editing application may be a non-linear video editing application and is hosted on a media editing system 902. The system 902 may be a general-purpose computer such as a stand-alone workstation or a personal computer, or may be computing resources located within the cloud. The media editing application comprises a plurality of software modules for performing various functions involved in editing still images and video images. For clarity, Figure 9 shows only the modules of the media editing application that are directly involved when implementing the methods described herein. The source image 904 can be stored locally or in the cloud and is input into an image processing module 906. The source image may be a still image or a sequence of images that make up a video stream. A command processor module 908 processes commands executed by the media editing application. The commands may be received from a user or may be automatically generated by a script. When receiving a command 910 to determine a semantic target region, the image processing module 906 sends the source image 904 to an object detection module 912. The object detection module can be implemented by a trained neural network module as described above. The object detection system detects objects and returns an object mask image that defines the spatial extent of the objects detected within the source image. The image processing system then divides the original image into a set of sub-images. Each sub-image contains a portion of the original image that depicts one of the detected objects. The sub-images and the entire source image are then sent to an image encoding module 914. The image encoding module 914 may be a trained neural network model as discussed above.In various embodiments, object detection module 912 and image encoding module 914 are implemented on one system trained to detect objects and determine their semantic interest levels. The image encoding module returns an embedding not only for the entire image, but also for each of the sub-images. The image processing system then performs a similarity determination and orders the semantic importance of each of the various objects according to the closeness of these embeddings to the "closeness" of the entire image. The image processing system can then store the source image tagged with semantic target area information. The tagged source image can also be made available to other modules within the media editing application that can perform the image processing tasks described above. In one embodiment, the various image processing tasks as described above can also be accelerated by special purpose hardware such as a GPU (not shown), or the image processing can be performed by a cloud-based service. The processed image 916 is output by the image processing module 906.

[0027]

[0035] The various components of the systems described herein can be implemented as a computer program using a general purpose computer system. Such a computer system typically includes a main unit connected to both an output device for displaying information to an operator and an input device for receiving input from the operator. The main unit generally includes a processor connected to a memory system via an interconnect mechanism. The input and output devices are also connected to the processor and the memory system via an interconnect mechanism.

[0028]

[0036] One or more output devices can also be connected to the computer system. Examples of output devices include various stereoscopic displays including liquid crystal displays (LCDs), plasma displays, OLED displays, displays that require viewing glasses and those that do not, cathode ray tubes, video projection systems and other video output devices, loudspeakers, headphones and other audio output devices, printers, devices for communicating on low or high bandwidth networks including network interface devices, cable modems, and storage devices such as solid state media including disks, tapes, flash memories but are not limited thereto. One or more input devices can also be connected to the computer system. Examples of input devices include, but are not limited to, keyboards, keypads, trackballs, mice, pens / styluses and tablets, touchscreens, cameras, communication devices, and data input devices. The present invention is not limited to specific input or output devices used in combination with the computer system or to the devices described herein.

[0029]

[0037] A computer system may be a general-purpose computer system, and a general-purpose computer system can be programmed using a computer programming language, a scripting language, or even an assembly language. Also, the computer system may be specially programmed special-purpose hardware. In a general-purpose computer system, the processor is typically a commercially available processor. Also, a general-purpose computer typically has an operating system. The operating system controls the execution of other computer programs and provides for scheduling, debugging, input / output control, accounting, compilation, storage allocation, data management and memory management, as well as communication control and related services. The computer system can be connected to a local network and / or a wide area network such as the Internet. The connected network can transfer to and from the computer system program instructions executed on the computer, media data such as video data, still image data, or audio data, metadata, report and approval information about media composition, media annotation, and other data.

[0030]

[0038] A memory system typically includes a computer-readable medium. This medium may be volatile or non-volatile, writable or non-writable, and / or rewritable or non-rewritable. A memory system typically stores data in binary form. Such data can define an application program executed by a microprocessor or information stored on a disk and processed by an application program. The present invention is not limited to a particular memory system. Time-based media such as video and audio can be stored on and input from magnetic, optical, or solid-state drives. These drives can also include an array of local disks or network attached disks.

[0031]

[0039] A system as described in this specification can be implemented in software, hardware, firmware, or a combination of these three. Various elements of this system can be implemented as one or more computer program products, either individually or in combination. Within a computer program product, computer program instructions are stored on a non-transitory computer-readable medium for execution by a computer, or transferred to a computer system through a local area network or a wide area network to which it is connected. The various steps of the process can be executed by a computer by executing such computer program instructions. The computer system can be a microprocessor computer system, or can include a plurality of computers connected on a computer network, or can be implemented in the cloud. The components described in this specification can be separate modules of a computer program or separate computer programs. These computer programs may also be operable on separate computers. The data generated by these components can be stored in a memory system or transmitted between computer systems by various communication media such as carrier signals.

[0032]

[0040] Although the embodiments have been described above, it should be apparent to those skilled in the art that the above description is merely illustrative and not limiting, and is presented only as an example. A number of modifications and other embodiments are within the scope of those skilled in the art and are considered to fall within the scope of the invention.

Claims

1. A method for determining a semantic target region in a source image, comprising: receiving the source image; detecting a plurality of objects in the source image using an automatic object detection system; re-dividing the source image into a plurality of sub-images, each sub-image including a portion of the source image that contains one of the detected plurality of objects; using a trained neural network model to generate an image embedding for the source image in a latent space, and generating an image embedding for each of the plurality of sub-images in the latent space; for each of the plurality of sub-images, determining a similarity between the image embedding of the sub-image and the image embedding of the source image, wherein the similarity between the image embedding of the sub-image and the image embedding of the source image corresponds to a multi-dimensional distance between the image embedding of the sub-image and the image embedding of the source image in the latent space; assigning a semantic interest degree to the detected object included in the sub-image according to the determined similarity between the image embedding of the sub-image corresponding to the detected object and the image embedding of the source image; outputting an indication of the semantic interest degree assigned to the detected object included in the sub-image; wherein the trained neural network model is implemented as an image encoder trained to encode an image into the latent space, and the image embedding for the source image and the image embeddings for each of the plurality of sub-images are generated, and the semantic interest degree for the detected object is weighted based on the magnitude of the determined similarity.

2. The method according to claim 1, wherein the automatic object detection system is a trained neural network model.

3. The method according to claim 1, wherein the trained neural network model used to generate the image embedding is a multimodal neural network.

4. The method according to claim 1, further comprising: generating an object mask for each detected object among the plurality of detected objects; generating an object mask image of the source image, wherein each detected object among the plurality of detected objects is replaced with a grayscale silhouette of the object mask generated for the object in the source image; applying a visual indication to each grayscale silhouette, wherein the visual indication indicates the semantic interest level assigned to the detected object corresponding to the object mask; A method comprising:

5. In the method according to claim 1, the indication of the semantic interest level assigned to each detected object among the plurality of detected objects is used to improve the image processing of the source image.

6. In the method according to claim 5, the image processing includes image compression, and the step of improving the image processing of the source image includes changing the number of bits assigned when compressing each sub-image among the plurality of sub-images according to the semantic interest level assigned to the detected object corresponding to the sub-image.

7. In the method according to claim 6, the source image is a frame of a video stream.

8. In the method according to claim 5, the image processing includes cropping a portion of the source image to achieve a desired aspect ratio of the source image, and the step of improving the image processing of the source image includes preferentially retaining objects with a higher assigned semantic interest level within the cropped portion of the source image.

9. In the method according to claim 8, the objects preferentially retained within the cropped image include the objects assigned the highest semantic interest level.

10. The method according to claim 8, further comprising: Selecting a subset of the detected objects from the plurality of detected objects, the selected subset of objects including a set of objects assigned a high semantic interest level; Locating the centroid of the subset of the detected objects within the source image; Cropping the source image such that the centroid of the subset of the detected objects within the source image is located at the center of the cropped image; A method comprising the above steps. **Claim 11** The method according to claim 8, wherein the source image is a frame of a video stream. **Claim 12** The method according to claim 5, wherein the image processing includes contrast enhancement, and the contrast enhancement includes enhancing the contrast in a region of the source image that includes detected objects assigned a high semantic interest level. **Claim 13** The method according to claim 12, wherein the source image is a frame of a video stream. **Claim 14** A computer program comprising: A non-transitory computer-readable medium encoded with computer-readable instructions that, when processed by a processing device, cause the processing device to execute a method for determining a semantic target region within a source image, the method including: Receiving the source image; Detecting a plurality of objects within the source image using an automatic object detection system; Redividing the source image into a plurality of sub-images, each sub-image including a portion of the source image that includes one of the plurality of detected objects; Using a trained neural network model, Generating an image embedding for the source image within a latent space, Generating an image embedding for each of the plurality of sub-images within the latent space; For each of the plurality of sub-images, A step of determining a similarity between the image embedding of the sub-image and the image embedding of the source image, wherein the similarity between the image embedding of the sub-image and the image embedding of the source image corresponds to a multi-dimensional distance between the image embedding of the sub-image and the image embedding of the source image in the latent space; A step of assigning a semantic interest degree to the detected object included in the sub-image according to the determined similarity between the image embedding of the sub-image corresponding to the detected object and the image embedding of the source image; A step of outputting an indication of the semantic interest degree assigned to the detected object included in the sub-image; including; The trained neural network model is implemented as an image encoder trained to encode an image into the latent space, and the image embedding for the source image and the image embedding for each of the plurality of sub-images are generated. A computer program in which the semantic interest degree for the detected object is weighted based on the magnitude of the determined similarity.

15. A system, A memory storing computer-readable instructions, A processor connected to the memory, wherein when the processor executes the computer-readable instructions, the system is caused to execute a method of determining a semantically target area in a source image. comprising, the method comprising: Receiving the source image; Detecting a plurality of objects in the source image using an automatic object detection system; A step of re-dividing the source image into a plurality of sub-images, each sub-image including a portion of the source image that includes one of the plurality of detected objects; Using a trained neural network model, Generating an image embedding for the source image in the latent space, Generating an image embedding for each sub-image in the plurality of sub-images in the latent space; For each sub-image in the plurality of sub-images, A step of determining a similarity between the image embedding of the sub-image and the image embedding of the source image, wherein the similarity between the image embedding of the sub-image and the image embedding of the source image corresponds to a multi-dimensional distance between the image embedding of the sub-image and the image embedding of the source image in the latent space; A step of assigning a semantic interest degree to the detected object included in the sub-image according to the determined similarity between the image embedding of the sub-image corresponding to the detected object and the image embedding of the source image; A step of outputting an indication of the semantic interest degree assigned to the detected object included in the sub-image; comprising; The trained neural network model is implemented as an image encoder trained to encode an image into the latent space, and the image embedding for the source image and the image embedding for each of the plurality of sub-images are generated; A system in which the semantic interest degree for the detected object is weighted based on the magnitude of the determined similarity.

Citation Information

Patent Citations

  • Sensor device and signal processing method

    JP2020068521A