Generating images with small objects for training pruned super-resolution networks

By detecting and cropping small objects in video frames, and adding simulated text to cropped images to generate training images, the problem that traditional pruning super-resolution networks are difficult to reconstruct small objects and text areas is solved, achieving clearer and more accurate image reconstruction.

CN119998827APending Publication Date: 2025-05-13SAMSUNG ELECTRONICS CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202380069854.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-09-06
Filing Date
2023-09-26
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

Traditional pruning super-resolution networks are difficult to learn and reconstruct small objects and text areas in the image when reconstructing images containing small objects, resulting in blurred and noisy results in reconstruction.

Method used

By detecting small objects in the input video frame, the crop object area generates a cropped image and overlays the simulated text on the cropped image to generate a training image for training a pruning convolutional neural network.

Benefits of technology

Improves the performance of pruned super-resolution networks when reconstructing images containing small objects and text areas, reduces noise and blur in the reconstruction results, and improves image quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119998827A_ABST
    Figure CN119998827A_ABST
Patent Text Reader

Abstract

An embodiment provides a method that includes detecting at least one object displayed within at least one input frame of an input video. The method further includes cropping at least one cropped image including the at least one object from the at least one input frame. The method further includes generating at least one training image by overlaying simulated text on the at least one cropped image. The method further includes providing the at least one training image to a pruned convolutional neural network (CNN). The pruned CNN learns from the at least one training image to reconstruct objects and text regions during image super-resolution.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates generally to image super-resolution (SR), and in particular to methods and systems for generating images with small objects for training a pruned SR network. Background Art

[0002] Image super-resolution (SR) is the process of recovering a high-resolution (HR) image from a low-resolution (LR) image. Summary of the invention

[0003] Technical Solution

[0004] In an embodiment, a method may include detecting at least one object displayed within at least one input frame of an input video. The method may also include cropping at least one cropped image including the at least one object from the at least one input frame. The method may also include generating at least one training image by overlaying simulated text on the at least one cropped image. The method may also include providing the at least one training image to a pruned convolutional neural network (CNN). The pruned CNN may learn from the at least one training image to reconstruct objects and text regions during image super-resolution.

[0005] In an embodiment, a system may include at least one processor and a processor-readable memory device storing instructions, wherein the instructions, when executed by the at least one processor, cause the at least one processor to perform operations. The operations may include detecting at least one object displayed within at least one input frame of an input video. The operations may also include cropping at least one cropped image including the at least one object from the at least one input frame. The operations may also include generating at least one training image by overlaying simulated text on the at least one cropped image. The operations may also include providing the at least one training image to a pruned CNN. The pruned CNN may learn from the at least one training image to reconstruct objects and text regions during image super-resolution.

[0006] In an embodiment, a processor-readable medium includes a program, wherein the program, when executed by a processor, may cause the processor to perform a method. The method may include detecting at least one object displayed within at least one input frame of an input video. The method may also include cropping at least one cropped image including the at least one object from the at least one input frame. The method may also include generating at least one training image by overlaying simulated text on the at least one cropped image. The method may also include providing the at least one training image to a pruned CNN. The pruned CNN may learn from the at least one training image to reconstruct objects and text regions during image super-resolution. BRIEF DESCRIPTION OF THE DRAWINGS

[0007] For a more complete understanding of the nature and advantages of the present disclosure and the preferred use embodiments, reference should be made to the following detailed description read in conjunction with the accompanying drawings, wherein:

[0008] Figure 1 shows an example of a high-resolution (HR) image reconstructed from a low-resolution (LR) image using a conventional pruned super-resolution (SR) network;

[0009] Figure 2 An example computing architecture for generating training data for training a pruned SR network to learn features corresponding to small objects according to an embodiment of the present disclosure is shown;

[0010] Figure 3 An example training system for generating training data for training a pruned SR network to learn features corresponding to small objects according to an embodiment of the present disclosure is shown;

[0011] Figure 4 An example workflow of a static graph generator of a training system according to an embodiment of the present disclosure is shown;

[0012] Figure 5 An example workflow for training a deep learning model utilized by a deep learning YouOnly Look Once (YOLO) based object detector of a training system is shown according to an embodiment of the present disclosure;

[0013] Figure 6 An example workflow of training a text overlayer of a system according to an embodiment of the present disclosure is shown;

[0014] Figure 7 shows an example input frame, an example probabilistic static map, and an example static detection-based cropped image according to an embodiment of the present disclosure;

[0015] Figure 8 shows an example input frame, an example output frame with one or more bounding boxes, and an example YOLO-based cropped image according to an embodiment of the present disclosure;

[0016] Fig. 9 shows an example cropped image and an example training image with added or overlaid text according to an embodiment of the present disclosure;

[0017] Fig.10 An example showing the visual difference between an LR image and an HR image reconstructed from the LR image using a pruned SR network trained using training images with added or overlaid text according to an embodiment of the present disclosure;

[0018] Fig.11a An example showing the visual difference between an HR image reconstructed by a conventional pruned SR network and an HR image reconstructed by a pruned SR network trained using training images with added or overlaid text according to an embodiment of the present disclosure;

[0019] Fig.11b The embodiment according to the present disclosure is shown Fig.11a The first set of close-up views of the HR images in;

[0020] Fig.11c The embodiment according to the present disclosure is shown Fig.11a A second set of close-up views of the HR images in;

[0021] Fig.11d The embodiment according to the present disclosure is shown Fig.11a The third set of close-up views of the HR images in;

[0022] Fig.12 is a flowchart of an example process for generating training data for training a pruned SR network to learn features corresponding to small objects according to an embodiment of the present disclosure; and

[0023] Fig.13 is a high-level block diagram illustrating an information processing system including a computer system that can be used to implement the disclosed embodiments. DETAILED DESCRIPTION

[0024] The following description is for the purpose of illustrating the general principles of one or more embodiments and is not meant to limit the inventive concept claimed herein. In addition, the specific features described herein may be used in combination with other described features in each of various possible combinations and arrangements. Unless otherwise specifically defined herein, all terms should be given their broadest possible interpretation, including the meaning implied in the specification and the meaning understood by those skilled in the art and / or the meaning defined in dictionaries, papers, etc.

[0025] The present disclosure relates generally to image super-resolution (SR), and in particular to methods and systems for generating images with small objects for training a pruned SR network. Embodiments In an embodiment, a method may include detecting at least one object displayed within at least one input frame of an input video. The method may also include cropping at least one cropped image including at least one object from the at least one input frame. The method may also include generating at least one training image by overlaying simulated text on at least one cropped image. The method may also include providing at least one training image to a pruned convolutional neural network (CNN). The pruned CNN may learn from at least one training image to reconstruct objects and text regions during image super-resolution.

[0026] In an embodiment, a system may include at least one processor and a processor readable memory device storing instructions that, when executed by the at least one processor, may cause the at least one processor to perform operations. The operations may include detecting at least one object displayed within at least one input frame of an input video. The operations may also include cropping at least one cropped image including the at least one object from the at least one input frame. The operations may also include generating at least one training image by overlaying simulated text on the at least one cropped image. The operations may also include providing the at least one training image to a pruned CNN. The pruned CNN may learn from the at least one training image to reconstruct objects and text regions during image super-resolution.

[0027] In an embodiment, a processor-readable medium includes a program that, when executed by a processor, may cause the processor to perform a method. The method may include detecting at least one object displayed within at least one input frame of an input video. The method may also include cropping at least one cropped image including the at least one object from the at least one input frame. The method may also include generating at least one training image by overlaying simulated text on the at least one cropped image. The method may also include providing the at least one training image to a pruned CNN. The pruned CNN may learn from the at least one training image to reconstruct objects and text regions during image super-resolution.

[0028] Traditional single image SR networks based on deep learning (such as CNN) are ubiquitous in the field of SR. Since these networks contain a large number of parameters, they are able to learn complex features ranging from low-level features (e.g., edges and corners) to high-level features (e.g., global features). However, these networks require a large amount of memory for storage on consumer electronic devices (e.g., TV hardware), which in turn becomes a bottleneck in hardware implementation. In order to design a hardware-friendly SR network, the network must be pruned to contain fewer parameters. However, the resulting pruned SR network clearly loses the ability to learn complex features ranging from low-level features to high-level features.

[0029] The performance of a deep neural network depends on the type of training samples included in the training data used to train the network. The better the training samples included in the training, the higher the performance of the network. For example, if a real-world image contains small objects (e.g., small objects such as logos, text, numbers, icons, or maps) that occupy less than five percent (5%) of the area / region of the image, and the image is part of a training pair used to train a pruned SR network, the resulting pruned SR network has a very low probability of learning or seeing the small object because of the size of the small object relative to the entire area / region of the image. Instead, the pruned SR network only learns to reconstruct the background and non-text areas / regions of the image, and fails to learn features corresponding to small objects (including text areas / regions) of the same image. Therefore, if a low-resolution (LR) image contains one or more small objects, the pruned SR network cannot reconstruct a true high-resolution (HR) image without noise from the LR image.

[0030] One or more embodiments provide a framework for improving the performance of pruned SR networks using optimal and well-managed training data sets. The framework provides training samples containing a complete set of features ranging from large objects to small objects and from non-text regions / regions to text regions / regions as training data sets. In an embodiment, one or more small objects of interest (such as icons, maps, logos, numbers and / or text) are extracted from real-world images to generate training images, which are then used as training pairs to train pruned SR networks. Using these training pairs, the pruned SR network learns the background and non-text regions / regions of the image and the small objects (including text regions / regions) of the same image, thereby enabling the pruned SR network to reconstruct a true HR image without noise from the corresponding LR image.

[0031] In an embodiment, static detection is used to extract small objects of interest from real-world images.

[0032] In an embodiment, a deep learning based You Only Look Once (YOLO) object detection algorithm is used to extract small objects of interest from real-world images.

[0033] In an embodiment, random text (such as words, sentences, and paragraphs) is overlaid on one or more training images to improve the ability of the pruned SR network to reconstruct true HR images without noise from the LR images.

[0034] Figure 1 An example of an HR image 10 reconstructed from an LR image using a conventional pruned SR network is shown. Figure 1As shown, the HR image 10 includes a background 15 and a plurality of small objects 11 (such as icons, logos, text, and maps). Each small object 11 occupies less than five percent (5%) of the area / region of the HR image 10. Figure 1 As shown, the conventional pruned SR network reconstructs the background 15 with high probability, but fails to truly reconstruct the small object 11 without noise. Compared to the background 15, the small object 11 displayed in the HR image 10 is blurred and / or has other visual artifacts, and may have a lower degree of color clarity and contrast.

[0035] Figure 2 An example computing architecture 100 for generating training data for training a pruned SR network to learn features corresponding to small objects according to an embodiment of the present disclosure is shown. The computing architecture 100 includes at least one training server 110, which includes resources such as one or more processor units 120 and one or more storage units 130. One or more applications 140 can execute / operate on the training server 110 using the resources of the training server 110.

[0036] In an embodiment, one or more applications 140 executed / operated on the training server 110 are configured to perform off-device (i.e., offline) training. In an embodiment, the off-device training includes: (1) generating training data including paired LR training samples and HR training samples (i.e., training pairs), and (2) training a pruned SR network based on the training data. As described in detail later herein, the pruned SR network is trained to learn features corresponding to background and non-text regions / areas of an image and features corresponding to small objects (including text regions / areas) of the same image. The resulting trained pruned SR network can be deployed for noise-reduced and / or artifact-reduced reconstruction of an HR image based on the LR image, wherein the LR image includes one or more small objects.

[0037] In an embodiment, the computing architecture 100 includes at least one electronic device 200, which includes resources, such as one or more processor units 210 and one or more storage units 220. One or more applications 260 may execute / operate on the electronic device 200 using the resources of the electronic device 200. In an embodiment, the one or more applications 260 may include one or more software mobile applications loaded onto or downloaded to the electronic device 200, such as a camera application, a social media application, a video streaming application, etc.

[0038] Examples of electronic device 200 include, but are not limited to, televisions (TVs) (e.g., smart TVs), mobile electronic devices (e.g., optimal frame rate tablets, smart phones, laptops, etc.), wearable devices (e.g., smart watches, smart bands, head-mounted displays, smart glasses, etc.), desktop computers, game consoles, cameras, media playback devices (e.g., DVD players), set-top boxes, Internet of Things (IoT) devices, cable boxes, satellite receivers, etc.

[0039] In an embodiment, the electronic device 200 includes one or more input / output (I / O) units 230 integrated in or coupled to the electronic device 200. In an embodiment, the one or more I / O units 230 include, but are not limited to, a physical user interface (PUI) and / or a graphical user interface (GUI), such as a remote control, a keyboard, a keypad, a touch interface, a touch screen, a knob, a button, a display screen, etc. In an embodiment, a user may utilize at least one I / O unit 230 to configure one or more parameters, provide user input, etc.

[0040] In an embodiment, the electronic device 200 includes one or more sensor units 240 integrated in or coupled to the electronic device 200. In an embodiment, the one or more sensor units 240 include, but are not limited to, an RGB color sensor, an IR sensor, an illumination sensor, a color temperature sensor, a camera, a microphone, a GPS, a motion sensor, and the like.

[0041] In an embodiment, the electronic device 200 includes a communication unit 250 configured to exchange data with at least one training server 110 via a communication network / connection 50 (e.g., a wireless connection (such as a Wi-Fi connection or a cellular data connection), a wired connection, or a combination of both a wireless connection and a wired connection). The communication unit 250 may include any suitable communication circuitry operable to connect to a communication network and exchange communication operations and media between the electronic device 200 and other devices connected to the same communication network 50. The communication unit 250 may be operable to use any suitable communication protocol (such as, for example, Wi-Fi (e.g., IEEE 802.11 protocol), High frequency systems (e.g., 900 MHz, 2.4 GHz and 5.6 GHz communication systems), infrared, GSM, GSM plus EDGE, CDMA, quad-band and other cellular protocols, VOIP, TCP-IP or any other suitable protocol) interface with the communication network.

[0042] In an embodiment, a trained pruned SR network (e.g., from a training server 110) is loaded or downloaded onto the electronic device 200 so that the pruned SR network can perform on-device (i.e., performed on the electronic device 200) noise-reduced and / or artifact-reduced reconstruction of an HR image from an LR image.

[0043] Figure 3 An example training system 300 for generating training data for training a pruned SR network to learn features corresponding to small objects according to an embodiment of the present disclosure is shown. In an embodiment, on the training server 110 ( Figure 2 ) executes / operates on one or more applications 140 ( Figure 2 ) includes a training system 300.

[0044] In an embodiment, the training system 300 includes a static map generator 310 configured to: (1) receive a sequence of n input frames 305 as input, (2) generate a probabilistic static map 315 based on the n input frames 305 using static detection, and (3) provide the probabilistic static map 315 as output. The probabilistic static map 315 is noise-free and contains only static small objects displayed within the n input frames 305. Examples of small objects include, but are not limited to, icons, text, numbers, maps, logos, etc. For example, in an embodiment, the n input frames 305 represent a history of n image / video frames of an input video, wherein the n image / video frames are previous image / video frames of the input video before the current image / video frame. In an embodiment, the n input frames 305 are received from a remote server hosting one or more online services (e.g., a video streaming service, a game streaming service, etc.). As described in detail later herein, in order to improve the detection of static small objects displayed within the n input frames 305, the static map generator 310 uses edge detection to implement static detection.

[0045] Static detection may fail to detect small objects that appear in the input video for only a brief period of time (e.g., in less than twenty (20) input frames 305 of the input video). In contrast, YOLO is an algorithm for object detection that only looks at an image once to predict what object appears within the image and where the object is.

[0046] In an embodiment, the training system 300 includes a deep learning YOLO-based object detector 320 configured to: (1) receive only one input frame 305 as input, and (2) use a YOLO-based object detection algorithm to detect and locate one or more small objects displayed within the input frame 305, and (3) provide an output frame 325 including the one or more small objects and one or more bounding boxes (BBs) thereof as output. As described in detail later herein, in an embodiment, the YOLO-based object detection algorithm includes a trained deep learning model.

[0047] The training system 300 implements static detection (via a static image generator 310) and YOLO-based object detection (via a YOLO-based object detector 320) to detect small objects that appear in an input video for only a brief period of time (e.g., in fewer than twenty (20) input frames 305 of the input video).

[0048] In an embodiment, the training system 300 includes a static detection-based cropped image generator 330 configured to: (1) receive as input a probabilistic static map 315 including one or more detected static small objects (e.g., from the static map generator 310), (2) obtain the center of each detected static small object based on the probabilistic static map 315, (3) generate a list of each pixel position including each center of each detected static small object, (4) randomly sample high-probability pixel positions (i.e., high-probability regions / areas of the probabilistic static map 315) from the list, and (5) generate a static detection-based cropped image 335 of size m×n including at least one of the detected static small objects, wherein the randomly sampled high-probability pixel positions are used as the centers of the static detection-based cropped image 335. The static detection-based cropped image generator 330 extracts features for at least one of the detected static small objects and provides a static detection-based cropped image 335 including the extracted features.

[0049] In an embodiment, the training system 300 includes a YOLO-based cropped image generator 340 configured to: (1) receive as input an output frame 325 including one or more detected small objects and one or more bounding boxes thereof (e.g., from a deep learning YOLO-based object detector 320), (2) generate a list including each bounding box corresponding to each detected small object, (3) randomly select a bounding box from the list, and (4) generate a YOLO-based cropped image 345 of size m×n including at least one of the detected small objects, wherein the randomly selected bounding box corresponds to the detected small object included in the YOLO-based cropped image 345. The YOLO-based cropped image generator 340 extracts features for at least one of the detected small objects and provides a YOLO-based cropped image 345 including the extracted features.

[0050] In an embodiment, the static detection based cropped image generator 330 and the YOLO based cropped image generator 340 are implemented as separate components of the training system 300. In another embodiment, the static detection based cropped image generator 330 and the YOLO based cropped image generator 340 are implemented as one component of the training system 300.

[0051] Real-world images may contain a variety of text information, such as author information, date and time of image generation or production, subtitles, frame number, etc. However, the traditional pruned SR network is trained using images that do not contain such text information, resulting in the inability of the trained pruned SR network to reconstruct HR images with text-specific features from LR images.

[0052] In an embodiment, the training system 300 includes a text overlay 350, wherein the text overlay 350 is configured to: (1) receive a cropped image 335 or 345 (e.g., from a static detection-based cropped image generator 330 or a YOLO-based cropped image generator 340), (2) generate simulated text using a text database 360 ​​including characters, words, sentences, paragraphs, names, and / or numbers, (3) add or overlay the simulated text on the cropped image 335 or 345 to obtain a training image 355, and (4) provide the training image 355 as output. As described in detail later herein, the simulated text may include characters, words, sentences, paragraphs, names, and / or numbers.

[0053] The text coverer 350 provides training images 355 with added or covered text to be used as training data when training the pruned SR network to learn text and symbols with correct grammar, thereby increasing the probability that the pruned SR network truly reconstructs HR images with characters, words, sentences, paragraphs, names and / or numbers from the LR images. For example, in an embodiment, the training server 110 includes a pruned SR network 370 (e.g., a pruned CNN), wherein the pruned SR network 370 is trained based on training pairs including the training image 355 (e.g., from the text coverer 350). As another example, in an embodiment, the training pair including the training image 355 (e.g., from the text coverer 350) is provided to the electronic device 200 or a remote server for training the pruned SR network deployed on the electronic device 200 or the remote server.

[0054] Figure 4 The static image generator 310 ( Figure 3 ) example workflow. For each input frame 305 ( Figure 3 ), the static image generator 310 converts the RGB image of the input frame 305 into a grayscale image I 图像 .

[0055] Edge detection involves using one or more matrices to calculate one or more regions of different pixel intensities in an image. Regions of extreme differences in pixel intensity generally indicate the edges of an object. Sobel edge detection is a widely used edge detection algorithm in image processing.

[0056] Next, the static image generator 310 applies Sobel edge detection to the grayscale image I 图像 , resulting in an edge map (i.e., edge image) that focuses on one or more regions of interest (i.e., one or more small objects) by removing one or more flat areas. If a small object is only displayed for a brief amount of time in one or more frames, Sobel edge detection helps to quickly detect such changes.

[0057] Let K x and K y Represent the x-direction kernel and the y-direction kernel respectively. Each kernel K x , K y is a 3×3 matrix including differently (or symmetrically) weighted indices. In an embodiment, the kernel K is expressed according to equations (1) to (2) provided below: x and K y :

[0058] as well as

[0059]

[0060] As part of Sobel edge detection, the static image generator 310 uses kernel convolution. Specifically, the static image generator 310 performs kernel convolution by using the x-direction kernel K x For grayscale image I 图像 Convolution is performed to transform the grayscale image I in the x direction 图像 Processing is performed to obtain the grayscale image I shown 图像 The static image generator 310 uses the y-direction kernel K y For grayscale image I 图像 Convolution is performed to transform the grayscale image I in the y direction 图像 Perform separate processing to obtain the grayscale image I 图像 The static image generator 310 calculates the square root of the sum of the square of the first edge map and the square of the second edge map according to the equation (3) provided below to generate a gradient magnitude image I 梯度幅值图像 :

[0061]

[0062] Among them, * represents the convolution operation, I 图像 *K x represents the first edge graph, and I 图像 *K y represents the second edge graph.

[0063] The static image generator 310 then generates a grayscale image I 图像 The weighted Gaussian filter is applied to the unit image to generate a weighted Gaussian image I 加权高斯 (i.e., convolution is performed using a weighted Gaussian filter). Since most small objects are usually displayed on or around the edge (i.e., boundary) of an image, the weighted Gaussian image I 加权高斯 Give more importance to the edges (i.e., boundaries) and less importance to the center - Weighted Gaussian Image I 加权高斯 The edges (i.e., boundaries) of the unit images are emphasized while the centers of the unit images are not emphasized. In an embodiment, the static image generator 310 generates the weighted Gaussian image I according to the equation (4) provided below: 加权高斯 :

[0064]

[0065] in, And γ=4.

[0066] The static image generator 310 uses the gradient magnitude image I 梯度幅值图像For the weighted Gaussian image I 加权高斯 The convolution is performed to generate a noise-free gradient magnitude image that reduces noise by minimizing the effects of one or more non-static objects.

[0067] Next, the static image generator 310 applies non-maximum suppression (edge ​​thinning method) to the noise-free gradient magnitude image, thereby obtaining a non-maximum suppressed and noise-free gradient magnitude image. Non-maximum suppression involves determining for each pixel of the image whether the pixel is a local maximum near the gradient of the pixel. If the pixel is a local maximum, the pixel is maintained as an edge pixel; otherwise, the pixel is suppressed. Therefore, the real edge in the image is more accurately represented by non-maximum suppression.

[0068] Some edges in the image may be brighter than other edges in the image, so that the image has brighter edges and lighter edges. More obvious edges can be seen in the brighter edges, but noise or edges can also be seen in the lighter edges. If the histogram of the image does not show obvious valleys, many background pixels of the image have the same grayscale as the grayscale of the object pixels of the image, and vice versa. In such a case, in order to determine whether a given edge corresponds to a true edge, next, the static image generator 310 applies hysteresis thresholding to the non-maximum suppressed and noise-free gradient magnitude image, thereby obtaining a noise-free binary image (i.e., a detection map).

[0069] Let High(θ 2 ) and Low(θ 1 ) represents two different thresholds used by the static map generator 310 when hysteresis thresholding. The static map generator 310 will have a value equal to or higher than the threshold High (θ 2 ) is classified as a true edge. The static graph generator 310 classifies any edge with a strength equal to or lower than the threshold Low(θ 1 ) is classified as a non-real edge. 2 ) and Low(θ 1 ) is connected to a real edge, the static graph generator 310 classifies the edge as a real edge; otherwise, the edge is deleted. 2 ) and Low(θ 1 ) Improved probability static graph 315 ( Figure 3 )’s quality.

[0070] Next, the static image generator 310 applies temporal averaging and temporal filtering to the noise-free binary image to respectively enhance the presence of static small objects and reduce any noise (if present). Temporal averaging enables the static image generator 310 to track small objects that disappear (i.e., leave) and appear (i.e., enter). In an embodiment, the average binary image obtained by temporal averaging and temporal filtering is expressed according to equation (5) provided below:

[0071] Average binary image = α × average binary image + (1-α) × binary image (5).

[0072] Next, the static map generator 310 applies probabilistic thresholding to the average binary image to generate a more robust, noise-free, and only static small object-containing probabilistic static map 315. Let av(i, j) generally represent a pixel of the average binary image, and let static(i, j) generally represent a pixel of the probabilistic static map 315. For each pixel av(i, j) of the average binary image, if the probability that the pixel av(i, j) has a non-zero pixel intensity is greater than n, then the pixel av(i, j) is included in the probabilistic static map 315 (i.e., static(i, j) = av(i, j)); otherwise, the pixel is set to zero in the probabilistic static map 315 (i.e., static(i, j) = 0).

[0073] Figure 5 The training of the object detector 320 ( Figure 3 ) is an example workflow for a deep learning model utilized by . Before training begins, the object detector 320 configures a YOLO-based object detection algorithm (e.g., learning rate, optimizer, number of layers) and initializes model parameters using a deep learning model pre-trained on a large-scale object detection, segmentation, and captioning dataset (such as the Common Objects in Context (COCO) dataset). The deep learning model is trained for several cycles until the loss function converges to an optimal value. At each iteration of training, it is determined whether the loss function is still decreasing after n cycles. If the loss function is still decreasing, an image containing the object and its bounding box is randomly selected from the dataset, the loss function is minimized using the randomly selected image, and the model parameters are updated. If the loss function is not still decreasing (i.e., the loss function has converged to an optimal value), training ends. The resulting trained deep learning model is then deployed by the object detector 320 to detect and localize the input frame 305 ( Figure 3 ) based on the YOLO object detection algorithm for small objects in .

[0074] Figure 6 The text overlay 350 ( Figure 3) example workflow. For each cropped image 335 ( Figure 3 ) or 345( Figure 3 ), the text overlay 350 uses a random number generator to extract the random number from the text database 360 ​​( Figure 3 ) randomly selects text R 1 If the text R 1 The length is less than the first threshold α 1 , the text overlay 350 will cover the text R 1 is overlaid on the cropped image 335 or 345. If the text R 1 The length is greater than the second threshold α 2 , then the text overlay 350 generates a number L having numbers randomly selected from [0,9]. number (i.e., 1≤L number ≤10), and the number L number is overlaid on the cropped image 335 or 345. If the text R 1 The length is greater than the first threshold α 1 But less than the second threshold α 2 , the text overlay 350 will cover the text R 1 Convert to uppercase and write the resulting uppercase text R 1 Overlaid on the cropped image 335 or 345 .

[0075] Figure 7 An example input frame 305, an example probabilistic static map 315, and an example static detection-based cropped image 335 according to an embodiment of the present disclosure are shown. In an embodiment, a static map generator 310 receives an input frame 305 and generates a probabilistic static map 315 by applying static detection to the input frame 305. A static detection-based cropped image generator 330 receives the probabilistic static map 315 and generates a static detection-based cropped image 335 based on the probabilistic static map 315 (i.e., high probability pixel locations of detected static small objects).

[0076] Figure 8 An example input frame 305, an example output frame 325 with one or more bounding boxes, and an example YOLO-based cropped image 345 are shown according to an embodiment of the present disclosure. In an embodiment, a deep learning YOLO-based object detector 320 receives the input frame 305 and generates the output frame 325 by applying a YOLO-based object detection algorithm. A YOLO-based cropped image generator 340 receives the output frame 325 and generates a YOLO-based cropped image 345 based on the output frame 325 (i.e., the bounding boxes of the detected small objects).

[0077] Fig. 9An example cropped image and an example training image 355 with added or overlaid text according to an embodiment of the present disclosure are shown. In an embodiment, the text overlay 350 receives a cropped image (e.g., Figure 3 ), and generating a training image 355 by adding or overlaying text on the cropped image.

[0078] Fig.10 4 shows an example of the visual difference between an LR image 401 and an HR image 402 reconstructed from the LR image 401 using a pruned SR network using a training image 355 ( Figure 3 ) is trained. In an embodiment, the pruned SR network receives the LR image 401, rescales the LR image 401 to the size of the HR image 402, and reconstructs the HR image 402. Compared with the LR image 401, the HR image 402 has a higher image quality, such as Fig.10 For example, the HR image 402 appears brighter, has less blur and / or other visual artifacts, and may have a higher degree of color clarity and contrast.

[0079] Fig.11a An example of the visual difference between an HR image 410 reconstructed by a conventional pruned SR network and an HR image 411 reconstructed by a pruned SR network using a training image 355 ( Figure 3 ) is trained. Compared with HR image 410, HR image 411 has higher image quality, such as Fig.11a For example, the HR image 411 appears brighter, has less blur and / or other visual artifacts, and may have a higher degree of color clarity and contrast.

[0080] Fig.11b The embodiment according to the present disclosure is shown Fig.11a 4. A first set of close-up views of HR images 410 and 411 in FIG. Some small objects (such as icons and logos) displayed within HR image 410 appear blurry and / or have other visual artifacts, such as Fig.11b In contrast, the same small object displayed within HR image 411 appears brighter, has less blur and / or other visual artifacts, and may have a higher degree of color clarity and contrast, as shown in FIG. Fig.11b shown.

[0081] Fig.11c The embodiment according to the present disclosure is shown Fig.11aA second set of close-up views of HR images 410 and 411 in FIG. The text displayed in HR image 410 has noise, such as Fig.11c In contrast, the same text displayed in HR image 411 is noise-free, as shown in Fig.11c shown.

[0082] Fig.11d The embodiment according to the present disclosure is shown Fig.11a 4. A third set of close-up views of HR images 410 and 411 in FIG. Some small objects (such as icons, logos, and text) displayed within HR image 410 are noisy and appear blurry and / or have other visual artifacts, such as Fig.11d In contrast, the same small object shown in HR image 411 is noise-free and appears brighter, has less blur and / or other visual artifacts, and may have a higher degree of color clarity and contrast, as shown in FIG. Fig.11d shown.

[0083] Fig.12 5 is a flowchart of an example process 500 for generating training data for training a pruned SR network to learn features corresponding to small objects according to an embodiment of the present disclosure. Processing block 501 includes (e.g., via Figure 3 The static image generator 310 in Figure 3 The deep learning YOLO-based object detector 320 in the input video detects at least one input frame (e.g., Figure 3 At least one object (e.g., Figure 1 11). Processing block 502 includes (e.g., via Figure 3 The static detection based cropped image generator 330 or Figure 3 The YOLO-based cropped image generator 340 in the embodiment of the present invention crops at least one cropped image (e.g., Figure 3 Processing block 503 includes (e.g., via Figure 3 The text overlay 350 in the example generates at least one training image (e.g., Figure 3 Processing block 504 includes providing at least one training image to a pruned convolutional neural network (CNN) (e.g., Figure 3 The pruned SR network 370 in FIG. 1 ), wherein the pruned CNN learns from at least one training image to reconstruct objects and text regions during image super-resolution.

[0084] In an embodiment, processing blocks 501 - 504 may be performed by one or more components of training system 300 .

[0085] Fig.13 is a high-level block diagram showing an information processing system including a computer system 900 for implementing the disclosed embodiments. System 300 may be incorporated into computer system 900. Computer system 900 includes one or more processors 910, and may also include an electronic display device 920 (for displaying video, graphics, text, and other data), a main memory 930 (e.g., random access memory (RAM)), a storage device 940 (e.g., a hard drive), a removable storage device 950 (e.g., a removable storage drive, a removable memory module, a tape drive, an optical drive, a computer-readable medium having computer software and / or data stored therein, a viewer interface device 960 (e.g., a keyboard, a touch screen, a keypad, a pointing device), and a communication interface 970 (e.g., a modem, a network interface (such as an Ethernet card), a communication port, or a PCMCIA slot and card). The communication interface 970 allows software and data to be transferred between the computer system and external devices. System 900 also includes a communication infrastructure 980 (e.g., a communication bus, a crossbar, or a network) to which the aforementioned devices / modules 910 to 970 are connected.

[0086] The information transmitted via the communication interface 970 may be in the form of signals, such as electronic, electromagnetic, optical, or other signals that can be received by the communication interface 970 via a communication link that carries the signals and may be implemented using wire or cable, optical fiber, telephone line, cellular telephone link, radio frequency (RF) link, and / or other communication channels. The computer program instructions representing the block diagrams and / or flow charts herein may be loaded onto a computer, a programmable data processing device, or a processing apparatus so that a series of operations performed thereon generate a computer-implemented process. In an embodiment, the process 500 ( Fig.12 ) may be stored as program instructions in the memory 930, the storage device 940 and / or the removable storage device 950 for execution by the processor 910.

[0087] In an embodiment, a method may include detecting at least one object displayed within at least one input frame of an input video.

[0088] In an embodiment, the method may further include: cropping at least one cropped image including at least one object from the at least one input frame.

[0089] In an embodiment, the method may further include generating at least one training image by overlaying simulated text on at least one cropped image.

[0090] In an embodiment, the method may further include providing at least one training image to a pruned convolutional neural network (CNN), wherein the pruned CNN learns from the at least one training image to reconstruct objects and text regions during image super-resolution.

[0091] In an embodiment, the step of detecting at least one object displayed within at least one input frame of the input video may include: applying Sobel edge detection to a predetermined number of input frames of the input video to generate a probabilistic static map, wherein the probabilistic static map is noise-free and contains only one or more static objects detected within the predetermined number of input frames.

[0092] In an embodiment, the step of cropping at least one cropped image including at least one object from at least one input frame may include: determining a center of each of one or more detected static objects based on a probabilistic static map; generating a list including each pixel position of each determined center; randomly sampling pixel positions from the list; and generating a cropped image including at least one of the one or more detected static objects, wherein the randomly sampled pixel position is the center of the cropped image.

[0093] In an embodiment, the step of detecting at least one object displayed within at least one input frame of an input video may include: training a deep learning model for object detection based on You Only Look Once (YOLO); using the deep learning model to detect and locate one or more objects within at least one input frame of the input video; and providing an output frame including one or more objects and one or more bounding boxes corresponding to the one or more objects.

[0094] In an embodiment, the step of cropping at least one cropped image including at least one object from at least one input frame may include: generating a list including each of one or more bounding boxes; randomly selecting a bounding box from the list; and generating a cropped image including at least one of the one or more objects, wherein the randomly selected bounding box corresponds to the object included in the cropped image.

[0095] In an embodiment, each object may occupy less than five percent of the entire area of ​​at least one input frame.

[0096] In an embodiment, each object may be one of an icon, a map, a logo, a number, or text.

[0097] In an embodiment, the simulated text may include at least one of a character, a word, a sentence, a paragraph, a name, or a number.

[0098] In an embodiment, a system may include: at least one processor; and a processor-readable memory device storing instructions, wherein the instructions, when executed by the at least one processor, cause the at least one processor to perform operations including: detecting at least one object displayed within at least one input frame of an input video.

[0099] In an embodiment, the operation may further include: cropping at least one cropped image including at least one object from the at least one input frame.

[0100] In an embodiment, the operation may further include generating at least one training image by overlaying simulated text on at least one cropped image.

[0101] In an embodiment, the operation may further include: providing at least one training image to a pruned convolutional neural network (CNN), wherein the pruned CNN learns from the at least one training image to reconstruct objects and text regions during image super-resolution.

[0102] In an embodiment, the step of detecting at least one object displayed within at least one input frame of the input video may include: applying Sobel edge detection to a predetermined number of input frames of the input video to generate a probabilistic static map, wherein the probabilistic static map is noise-free and contains only one or more static objects detected within the predetermined number of input frames.

[0103] In an embodiment, the step of cropping at least one cropped image including at least one object from at least one input frame may include: determining a center of each of one or more detected static objects based on a probabilistic static map; generating a list including each pixel position of each determined center; randomly sampling pixel positions from the list; and generating a cropped image including at least one of the one or more detected static objects, wherein the randomly sampled pixel position is the center of the cropped image.

[0104] In an embodiment, the step of detecting at least one object displayed within at least one input frame of an input video may include: training a deep learning model for object detection based on You Only Look Once (YOLO); using the deep learning model to detect and locate one or more objects within at least one input frame of the input video; and providing an output frame including one or more objects and one or more bounding boxes corresponding to the one or more objects.

[0105] In an embodiment, the step of cropping at least one cropped image including at least one object from at least one input frame may include: generating a list including each of one or more bounding boxes; randomly selecting a bounding box from the list; and generating a cropped image including at least one of the one or more objects, wherein the randomly selected bounding box corresponds to the object included in the cropped image.

[0106] In an embodiment, each object may occupy less than five percent of the entire area of ​​at least one input frame.

[0107] In an embodiment, each object may be one of an icon, a map, a logo, a number, or text.

[0108] In an embodiment, the simulated text may include at least one of a character, a word, a sentence, a paragraph, a name, or a number.

[0109] In an embodiment, a processor-readable medium includes a program, wherein the program, when executed by a processor, causes the processor to perform a method, the method comprising: detecting at least one object displayed within at least one input frame of an input video.

[0110] In an embodiment, the method may further include: cropping at least one cropped image including at least one object from the at least one input frame.

[0111] In an embodiment, the method may further include generating at least one training image by overlaying simulated text on at least one cropped image.

[0112] In an embodiment, the method may further include providing at least one training image to a pruned convolutional neural network (CNN), wherein the pruned CNN learns from the at least one training image to reconstruct objects and text regions during image super-resolution.

[0113] In an embodiment, the step of detecting at least one object displayed within at least one input frame of the input video may include: applying Sobel edge detection to a predetermined number of input frames of the input video to generate a probabilistic static map, wherein the probabilistic static map is noise-free and contains only one or more static objects detected within the predetermined number of input frames.

[0114] In an embodiment, the step of detecting at least one object displayed within at least one input frame of an input video may include: training a deep learning model for object detection based on You Only Look Once (YOLO); using the deep learning model to detect and locate one or more objects within at least one input frame of the input video; and providing an output frame including one or more objects and one or more bounding boxes corresponding to the one or more objects.

[0115] In an embodiment, each object may occupy less than five percent of the entire area of ​​at least one input frame.

[0116] Embodiments have been described with reference to flowchart illustrations and / or blocks of methods, devices (systems) and computer program products. Each block of such diagrams / illustrations or their combination may be implemented by computer program instructions. The computer program instructions produce a machine when provided to a processor so that instructions executed by the processor create a device for implementing the functions / operations specified in the flowchart and / or blocks. Each block in the flowchart / block may represent a hardware module and / or a software module or logic. In alternative implementations, the functions mentioned in the blocks may not occur in the order mentioned in the accompanying drawings, may occur simultaneously, etc.

[0117] The terms "computer program medium", "computer usable medium", "computer readable medium" and "computer program product" are generally used to refer to media such as main memory, secondary memory, removable storage drives, hard disks installed in hard drives, and signals. These computer program products are devices for providing software to computer systems. Computer readable media allow computer systems to read data, instructions, messages or message packets, and other computer readable information from computer readable media. Computer readable media may include, for example, non-volatile memory such as floppy disks, ROMs, flash memory, disk drive memory, CD-ROMs, and other permanent storage. For example, it is useful for transmitting information (such as data and computer instructions) between computer systems. Computer program instructions may be stored in a computer readable medium, which may instruct a computer, other programmable data processing device, or other device to act in a specific manner so that the instructions stored in the computer readable medium produce an article of manufacture including instructions for implementing the functions / actions specified in the flowchart and / or block diagram box or box.

[0118] As will be appreciated by those skilled in the art, various aspects of the embodiments may be implemented as systems, methods or computer program products. Therefore, various aspects of the embodiments may take the form of complete hardware embodiments, complete software embodiments (including firmware, resident software, microcode, etc.) or embodiments combining software aspects and hardware aspects, which may generally be referred to herein as "circuits", "modules" or "systems". In addition, various aspects of the embodiments may take the form of computer program products implemented in one or more computer-readable media, and one or more computer-readable media have computer-readable program codes implemented thereon.

[0119] Any combination of one or more computer-readable media may be utilized. A computer-readable medium may be a computer-readable storage medium. A computer-readable storage medium may be, for example, but not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or apparatus, or any suitable combination of the foregoing. More specific examples (a non-exhaustive list) of computer-readable storage media would include the following: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing. In the context of this document, a computer-readable storage medium may be any tangible medium that may contain or store a program used by or in conjunction with an instruction execution system, device, or apparatus.

[0120] The computer program code for performing the operations of various aspects of the embodiments can be written in any combination of one or more programming languages ​​(including object-oriented programming languages ​​such as Java, Smalltalk, C++, etc., and traditional processing programming languages ​​such as "C" programming language or similar programming languages). The program code can be executed entirely on the user's computer, partially on the user's computer, as an independent software package, partially on the user's computer and partially on a remote computer, or completely on a remote computer or server. In the latter case, the remote computer can be connected to the user's computer through any type of network (including a local area network (LAN) or a wide area network (WAN)), or can be connected to an external computer (e.g., through the Internet using an Internet service provider).

[0121] Aspects of the embodiments are described above with reference to flowchart illustrations and / or block diagrams of methods, devices (systems) and computer program products. It should be understood that each frame of the flowchart illustration and / or block diagram and the combination of frames in the flowchart illustration and / or block diagram can be implemented by computer program instructions. These computer program instructions can be provided to a special-purpose computer or other programmable data processing device to produce a machine, so that instructions executed by a processor of the computer or other programmable data processing device create a device for implementing the functions / actions specified in the flowchart and / or block diagram frame or frame.

[0122] These computer program instructions may also be stored in a computer-readable medium, which may instruct a computer, other programmable data processing device, or other apparatus to act in a specific manner so that the instructions stored in the computer-readable medium produce an article of manufacture including instructions for implementing the functions / actions specified in the flowchart and / or block diagram frame or frames.

[0123] The computer program instructions may also be loaded onto a computer, other programmable data processing device, or other apparatus to cause a series of operational steps to be performed on the computer, other programmable device, or other apparatus to produce a computer-implemented process, such that the instructions executed on the computer or other programmable device provide a process for implementing the functions / actions specified in the flowchart and / or block diagram box or boxes.

[0124] The flow chart and block diagram in the accompanying drawings illustrate the architecture, function and operation of the possible implementation of the system, method and computer program product according to various embodiments. In this regard, each frame in the flow chart or block diagram may represent a module, segment or part of an instruction, which includes one or more executable instructions for realizing the specified logical function. In some alternative embodiments, the function mentioned in the frame may not occur in the order mentioned in the figure. For example, the two frames shown in succession can actually be performed substantially at the same time, or the frame can sometimes be performed in reverse order, depending on the function involved. It should also be noted that each frame of the block diagram and / or the flow chart diagram and the combination of the frames in the block diagram and / or the flow chart diagram can be implemented by a system based on special-purpose hardware, wherein the system based on special-purpose hardware performs a specified function or action or performs a combination of special-purpose hardware and computer instructions.

[0125] Unless explicitly stated otherwise, reference to an element in the singular form in a claim is not intended to mean "one and only one," but rather "one or more." All structural and functional equivalents to the elements of the above-described exemplary embodiments that are currently known or later come to be known to one of ordinary skill in the art are intended to be covered by the present claims. Unless an element is expressly recited using the phrase "means for" or "step for," no claim element herein is to be interpreted pursuant to the provisions of 35 U.S.C. § 112 (f).

[0126] The terms used herein are only for the purpose of describing specific embodiments and are not intended to limit the disclosed technology. As used herein, unless the context clearly indicates otherwise, the singular forms "one", "an" and "the" are intended to also include plural forms. It will be further understood that the terms "include" and / or "comprise" when used in this specification specify the presence of the features, integers, steps, operations, elements and / or parts, but do not exclude the presence or addition of one or more other features, integers, steps, operations, elements, parts and / or their groups.

[0127] The corresponding structures, materials, actions, and equivalents of all means or steps and functional elements in the above claims are intended to include any structure, material, or action for performing the function in combination with other claimed elements specifically claimed. The description of the embodiments has been presented for the purpose of illustration and description, but is not intended to be exhaustive or limited to the embodiments in the disclosed form. Many modifications and variations will be apparent to those of ordinary skill in the art without departing from the scope and spirit of the disclosed technology.

[0128] Although embodiments have been described with reference to certain versions thereof; however, other versions are possible. Therefore, the spirit and scope of the appended claims should not be limited to the description of the preferred versions contained herein.

Claims

1. A method comprising: detecting at least one object displayed within at least one input frame (305) of an input video; cropping at least one cropped image (335, 345) including the at least one object from the at least one input frame (305); generating at least one training image (355) by overlaying simulated text on the at least one cropped image (335, 345); as well as The at least one training image (355) is provided to a pruned convolutional neural network (CNN), wherein the pruned CNN learns from the at least one training image (355) to reconstruct objects and text regions during image super-resolution.

2. The method according to claim 1, wherein: The step of detecting at least one object displayed in at least one input frame of the input video comprises: Sobel edge detection is applied to a predetermined number of input frames (305) of the input video to generate a probabilistic static map, wherein the probabilistic static map is noise-free and contains only one or more static objects detected within the predetermined number of input frames (305).

3. The method according to claim 1 or 2, wherein: The step of cropping at least one cropped image including the at least one object from the at least one input frame comprises: determining a center of each of the one or more detected static objects based on the probabilistic static map; generating a list including each pixel location of each center determined; Randomly sampling pixel positions from the list; and A cropped image (335) is generated that includes at least one of the one or more detected static objects, wherein the randomly sampled pixel position is a center of the cropped image (335).

4. The method according to any one of claims 1 to 3, wherein: The step of detecting at least one object displayed in at least one input frame of the input video comprises: detecting and locating one or more objects within the at least one input frame (305) of the input video using a deep learning model for You Only Look Once (YOLO) based object detection; and An output frame is provided that includes the one or more objects and one or more bounding boxes corresponding to the one or more objects.

5. The method according to any one of claims 1 to 4, wherein: The step of cropping at least one cropped image including the at least one object from the at least one input frame comprises: generating a list including each of the one or more bounding boxes; randomly selecting a bounding box from the list; and A cropped image (345) is generated that includes at least one of the one or more objects, wherein the randomly selected bounding box corresponds to the object included in the cropped image (345).

6. The method according to any one of claims 1 to 5, wherein: Each object occupies less than five percent of the entire area of ​​the at least one input frame (305).

7. The method according to any one of claims 1 to 6, wherein: Each object is one of an icon, a map, a logo, a number, or text.

8. A processor-readable medium comprising a program, wherein the program, when executed by a processor, causes the processor to perform the method according to any one of claims 1 to 7.

9. A system comprising: at least one processor (910); as well as A processor-readable memory device (930) storing instructions, wherein the instructions, when executed by the at least one processor (910), cause the at least one processor (910) to perform operations comprising: detecting at least one object displayed within at least one input frame (305) of an input video; cropping at least one cropped image (335, 345) including the at least one object from the at least one input frame (305); generating at least one training image (355) by overlaying simulated text on the at least one cropped image (335, 345); and The at least one training image (355) is provided to a pruned convolutional neural network (CNN), wherein the pruned CNN learns from the at least one training image (355) to reconstruct objects and text regions during image super-resolution.

10. The system according to claim 9, wherein: The step of detecting at least one object displayed in at least one input frame of the input video comprises: Sobel edge detection is applied to a predetermined number of input frames (305) of the input video to generate a probabilistic static map, wherein the probabilistic static map is noise-free and contains only one or more static objects detected within the predetermined number of input frames (305).

11. The system according to claim 9 or 10, wherein: The step of cropping at least one cropped image including the at least one object from the at least one input frame comprises: determining a center of each of the one or more detected static objects based on the probabilistic static map; generating a list including each pixel location of each center determined; Randomly sampling pixel positions from the list; and A cropped image (335) is generated that includes at least one of the one or more detected static objects, wherein the randomly sampled pixel position is a center of the cropped image (335).

12. A system according to any one of claims 9 to 11, wherein: The step of detecting at least one object displayed in at least one input frame of the input video comprises: detecting and locating one or more objects within the at least one input frame (305) of the input video using a deep learning model for You Only Look Once (YOLO) based object detection; and An output frame is provided that includes the one or more objects and one or more bounding boxes corresponding to the one or more objects.

13. A system according to any one of claims 9 to 12, wherein: The step of cropping at least one cropped image including the at least one object from the at least one input frame comprises: generating a list including each of the one or more bounding boxes; randomly selecting a bounding box from the list; and A cropped image (345) is generated that includes at least one of the one or more objects, wherein the randomly selected bounding box corresponds to the object included in the cropped image (345).

14. A system according to any one of claims 9 to 13, wherein: Each object occupies less than five percent of the entire area of ​​the at least one input frame (305).

15. A system according to any one of claims 9 to 14, wherein: Each object is one of an icon, a map, a logo, a number, or text.