Method of expanding output area of video and device for performing same
Patent Information
- Application Number
- PCT/KR2025/000148
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-12-16
- Filing Date
- 2025-01-03
- Publication Date
- 2025-10-02
AI Technical Summary
The issue of letterboxing and distortion occurs when videos with different aspect ratios are output, leading to reduced user immersion and content loss.
An electronic device determines a generation target area based on object characteristics across multiple frames, using artificial intelligence models to align and expand the video output area, aligning objects spatially and temporally to fill letterbox areas and correct distortion.
Enhances user experience by eliminating letterboxing and distortion, providing a seamless video output that maintains image quality and immersion.
Smart Images

Figure KR2025000148_02102025_PF_FP_ABST
Abstract
Description
Method for expanding the output area of a video and device for doing so
[0001] The present disclosure relates to a method for expanding an output area of a video and a device for performing the same, and more particularly, to a method for expanding an output area of a video based on a plurality of frames included in the video and a device for performing the same.
[0002] When the aspect ratio of the video to be output is different from that of the display on which the video is output, or when a video shot with a standard lens is output in a panoramic view, there is a problem in which the video content is lost or distorted, or letter boxes are formed in the output video.
[0003] A method according to one embodiment may include a step of determining an area in which an image is to be expanded for a current frame among a plurality of frames included in the video as a generation target area.
[0004] A method according to one embodiment may include obtaining object properties of at least one object included in the current frame.
[0005] A method according to one embodiment may include a step of obtaining an object characteristic of at least one object included in frames other than the current frame among the plurality of frames, based on an object characteristic of at least one object included in the current frame.
[0006] A method according to one embodiment may include a step of generating at least one object in the generation target area based on an object characteristic of at least one object included in the current frame and an object characteristic of at least one object included in the other frames.
[0007] An electronic device according to one embodiment may include a memory storing a program or at least one instruction and at least one processor executing at least one instruction stored in the memory.
[0008] According to one embodiment, the electronic device can determine an area in which an image is to be expanded for a current frame among a plurality of frames included in a video as a generation target area by having the at least one processor execute at least one command stored in the memory.
[0009] According to one embodiment, the electronic device can obtain object characteristics of at least one object included in the current frame by having the at least one processor execute at least one command stored in the memory.
[0010] According to one embodiment, the electronic device can obtain object characteristics of at least one object included in frames other than the current frame among the plurality of frames based on object characteristics of at least one object included in the current frame by having the at least one processor execute at least one command stored in the memory.
[0011] According to one embodiment, the electronic device can generate the at least one object in the generation target area based on object characteristics of at least one object included in the current frame and object characteristics of at least one object included in the other frames by having the at least one processor execute at least one command stored in the memory.
[0012] Figure 1 is a drawing showing an example in which a letter box is formed in an output video.
[0013] FIG. 2 is a flowchart illustrating an operating method of an electronic device according to one embodiment.
[0014] FIG. 3 is a drawing for explaining an operating method of an electronic device according to one embodiment.
[0015] FIG. 4 is a flowchart illustrating an operating method of an electronic device according to one embodiment.
[0016] FIG. 5 is a diagram illustrating a method for an electronic device according to one embodiment to obtain a correlation between an object characteristic of at least one object included in a current frame and an object characteristic of at least one object included in other frames.
[0017] FIG. 6 is a flowchart illustrating an operation method of an electronic device according to one embodiment.
[0018] FIG. 7 is a drawing for explaining an operating method of an electronic device according to one embodiment.
[0019] Figure 8 is a flowchart for explaining an operating method of an electronic device according to one embodiment.
[0020] FIG. 9 is a diagram illustrating a method for an electronic device to create at least one object in a creation target area according to one embodiment.
[0021] FIG. 10 is a drawing for explaining an operating method of an electronic device according to one embodiment.
[0022] FIG. 11 is a block diagram illustrating components of an electronic device according to one embodiment.
[0023] The terms used in this specification will be briefly explained, and the present invention will be described in detail.
[0024] The terms used in this invention have been selected from widely used, current terms, taking into account the functions of the invention. However, these terms may vary depending on the intentions of those skilled in the art, precedents, the emergence of new technologies, etc. Furthermore, in certain cases, terms may be arbitrarily selected by the applicant, in which case their meanings will be described in detail in the relevant description of the invention. Therefore, the terms used in this invention should not be defined simply as names, but rather based on their inherent meanings and the overall content of the invention.
[0025] When a part of the specification is said to "include" a component, this does not exclude other components, but rather implies the inclusion of other components, unless otherwise specifically stated. Furthermore, terms such as "part," "module," etc., used throughout the specification refer to a unit that processes at least one function or operation, which may be implemented in hardware, software, or a combination of hardware and software.
[0026] Additionally, the description 'at least one of A, B, and C' means that it can be any one of 'A', 'B', 'C', 'A and B', 'A and C', 'B and C', and 'A, B, and C'.
[0027] It should be understood that the combinations of blocks and sequence diagrams in each flowchart can be performed by one or more computer programs containing computer-executable instructions. The one or more computer programs may be stored entirely in a single memory or may be divided and stored in multiple different memories.
[0028] All functions or operations described in this document may be performed by a single processor or a combination of processors. A single processor or a combination of processors is a circuitry that performs processing, and may include circuitry such as an Application Processor (AP), a Communication Processor (CP), a Graphical Processing Unit (GPU), a Neural Processing Unit (NPU), a Microprocessor Unit (MPU), a System on Chip (SoC), or an Integrated Chip (IC).
[0029] A processor may include various processing circuits and / or multiple processors. For example, the term “processor” as used herein, including in the claims, may include various processing circuits, including at least one processor. At least one processor, one or more processors, may be individually and / or collectively configured to perform the various functions described herein in a distributed fashion. As used herein, “processor,” “at least one processor,” and “one or more processors” may be configured to perform multiple functions. However, these terms encompass, without limitation, situations where one processor performs some of the functions and other processor(s) perform other parts of the functions, and situations where a single processor may perform all of the functions. Furthermore, the at least one processor may include a combination of processors that perform various of the disclosed functions in a distributed manner. At least one processor may execute program instructions to achieve or perform various functions.
[0030] Below, with reference to the attached drawings, embodiments of the present invention are described in detail so that those skilled in the art can easily implement the present invention. However, the present invention can be implemented in various different forms and is not limited to the embodiments described herein. In the drawings, parts irrelevant to the description have been omitted to clearly explain the present invention, and similar parts have been designated with similar reference numerals throughout the specification.
[0031] Figure 1 is a drawing showing an example in which a letter box is formed in an output video.
[0032] There is a problem of letter box formation when the aspect ratio of the video to be output and the display on which the video is output are different.
[0033] Letterbox refers to the area displayed on the display along with the video to maintain the image quality of the video to be displayed when the aspect ratio of the video to be displayed and the display on which the video is displayed are different.
[0034] Letterboxing can refer to black bars that appear along with a video on a display when the aspect ratio of the video you are trying to output is different from that of the display on which the video is being output.
[0035] Letter boxes may be formed on at least one of the left, right, top, or bottom of the output video.
[0036] For example, when a video (110) with a width-to-height ratio of 4:3 is to be output to a display (120) with a width-to-height ratio of 16:10, letter boxes (122) may be formed on the left and right sides of the video to maintain the image quality of the video (110).
[0037] For example, if you want to output a video with a 16:9 aspect ratio to a display with a 16:10 aspect ratio, letterboxes may be formed on the left and right sides of the video.
[0038] For example, if you want to output a video with a width-to-height ratio of 1:1 to a display with a width-to-height ratio of 16:9, letterboxes may be formed on the left, right, top, and bottom of the video.
[0039] Users who view video content through a display with letterboxing face problems such as a significant reduction in user usability, such as a decrease in immersion in the content due to letterboxing.
[0040] Additionally, when attempting to display a video shot with a standard lens in panoramic view, distortion, such as loss of content contained in the video, may occur due to differences in the aspect ratios of the video shot with the standard lens and the display providing the panoramic view. Here, a standard lens may refer to a lens with a focal length of 50 mm.
[0041] Therefore, in order to provide higher usability and satisfaction to users watching videos, there is a need to expand the output area of the video to cover areas where letter boxes are formed or distortion occurs on the display on which the video is output.
[0042] Hereinafter, the present disclosure will describe a method for expanding an output area of a video based on a plurality of frames included in the video and an electronic device for performing the same.
[0043] FIG. 2 is a flowchart illustrating an operating method of an electronic device according to one embodiment.
[0044] In step 210, the electronic device can determine an area in which an image is to be expanded for a current frame among multiple frames included in the video as a generation target area.
[0045] The current frame refers to a frame in the video output to the display that has an area where image expansion is possible.
[0046] For example, the current frame may mean the frame containing letterboxing if the output video contains letterboxing.
[0047] The current frame can change over time based on motion information in the video.
[0048] For example, if a video contains frame t1, frame t2, and frame t3, the current frame at time t1 may be frame t1, the current frame at time t2 may be frame t2, and the current frame at time t3 may be frame t3. However, this is not limited thereto.
[0049] Motion information may refer to information indicating pixel-level movement between consecutive frames. For example, motion information may refer to information on the direction and distance each pixel moved between the previous frame and the current frame. Motion information may include, but is not limited to, information about direction information and / or vector fields indicating pixel-level movement between consecutive frames.
[0050] The target generation area may refer to the area in which the image can be expanded in the video output to the display.
[0051] For example, if the video output to the current frame contains a letter box, the target area to be generated may mean the letter box, but is not limited thereto.
[0052] An electronic device according to one embodiment can determine a target region for generation based on comparing an aspect ratio of a video and an aspect ratio of a display on which the video is output.
[0053] For example, when an electronic device wants to output a video with a width-to-height ratio of 4:3 to a display with a width-to-height ratio of 16:10, it can determine the letter boxes formed on the left and right sides of the output video as the target areas to be generated.
[0054] For example, when an electronic device wants to output a video with a width-to-height ratio of 1:1 to a display with a width-to-height ratio of 16:9, it may determine the letter boxes formed on the left, right, top, and bottom of the output video as the target areas for generation. However, this is not limited thereto.
[0055] An electronic device according to one embodiment may determine a target area to be generated in units of at least one of a pixel or a patch. A patch may be composed of a plurality of pixels.
[0056] In step 220, the electronic device can obtain object characteristics of at least one object included in the current frame.
[0057] An electronic device according to one embodiment can recognize an object in units of at least one of pixels or patches.
[0058] For example, if the current frame contains content related to mountains, rivers, and the sky, the current frame may include the mountains, rivers, and the sky as objects. For example, if the current frame contains content related to people, desks, and food, the current frame may include people, desks, and food as objects.
[0059] For example, if the current frame contains content related to a person, the current frame may include as an object at least one of a patch consisting of pixels constituting a person or a set of pixels constituting a person.
[0060] For example, if the current frame contains content related to a river, the current frame may contain as an object at least one of a pixel constituting the river or a patch consisting of a set of pixels constituting the river.
[0061] Object characteristics may include at least one of location information or visual information of at least one object included in the current frame.
[0062] Location information can include information indicating the positional relationships between objects. For example, location information can include information about which objects are located above, below, left, and right of a specific object, and the order in which the objects are arranged.
[0063] The visual information may include at least one of color information, texture information, shape information, pattern information, or size information of at least one object.
[0064] Color information may include information about the (R, G, B) values of at least one object.
[0065] Texture information may include at least one of information about the roughness, glossiness, or hardness of the surface of at least one object.
[0066] Shape information may include edge information of at least one object.
[0067] The pattern information may include geometric pattern information of at least one object.
[0068] The size information may include relative size information between at least one object.
[0069] An electronic device according to one embodiment can identify at least one object contained in a current frame.
[0070] An electronic device according to one embodiment can identify a type of at least one object included in a current frame.
[0071] An electronic device according to one embodiment can obtain location information indicating a location relationship between identified objects.
[0072] An electronic device according to one embodiment can obtain object characteristics of at least one object included in a current frame based on positional information between at least one identified object.
[0073] In step 230, the electronic device can obtain object characteristics of at least one object included in frames other than the current frame among a plurality of frames based on object characteristics of at least one object included in the current frame.
[0074] An electronic device according to one embodiment can store information about a plurality of frames included in a video in a memory.
[0075] Multiple frames may include frame t0, frame t1, frame t2, ..., frame tn.
[0076] Other frames refer to the remaining frames (t0, t2, t3,..., tn) excluding the current frame (t1) among multiple frames.
[0077] A specific method for an electronic device according to one embodiment to obtain object characteristics of at least one object included in frames other than the current frame among a plurality of frames based on object characteristics of at least one object included in the current frame will be described with reference to FIG. 3.
[0078] FIG. 3 is a drawing for explaining an operating method of an electronic device according to one embodiment.
[0079] FIG. 3 is a drawing illustrating an example of an electronic device according to one embodiment obtaining object characteristics of at least one object included in other frames when a letter box is formed due to a mismatch in the aspect ratio of the video (310) and the aspect ratio of the display on which the video is output.
[0080] An electronic device according to one embodiment can determine a target region (331, 333) for generating an image to be expanded for a current frame (330) among a plurality of frames (320, 330, 340) included in a video.
[0081] For convenience of explanation, only three frames (320, 330, 340) are illustrated in FIG. 3, but the number of multiple frames included in the video is not limited thereto.
[0082] An electronic device according to one embodiment can obtain object characteristics (332, 334) of at least one object included in a current frame (330).
[0083] For example, an electronic device according to one embodiment can obtain lecture object properties (332) included in the current frame (330).
[0084] Here, the river object characteristics (332) may include location information related to being located on the left side of the mountain and being located below the sky.
[0085] Here, the lecture object characteristics (332) may include at least one of lecture color information, texture information, shape information, pattern information, or size information.
[0086] For example, an electronic device according to one embodiment can obtain object characteristics (334) of a mountain included in a current frame (330).
[0087] Here, the object characteristics (334) of the mountain may include location information related to being located on the right side of the river and being located below the sky.
[0088] Here, the object characteristics (334) of the mountain may include at least one of color information, texture information, shape information, pattern information, or size information of the mountain.
[0089] An electronic device according to one embodiment can obtain object characteristics (322, 324, 342, 344) of at least one object included in frames (320, 340) other than the current frame (330) among a plurality of frames (320, 330, 340), based on object characteristics (332, 334) of at least one object included in the current frame (330).
[0090] Object characteristics may include at least one of positional information or visual information of the object.
[0091] An electronic device according to one embodiment can obtain object characteristics (322, 324, 342, 344) of at least one object included in frames (320, 340) other than the current frame (330) among a plurality of frames (320, 330, 340), based on location information of at least one object included in the current frame (330).
[0092] For example, if an electronic device acquires lecture location information included in the current frame (330), it can acquire lecture object properties (332, 342) included in other frames (320, 340) based on the lecture location information.
[0093] For example, if the electronic device obtains location information related to a river being located to the left of a mountain and below the sky, it can obtain object properties (332, 342) for the river being located to the left of a mountain and below the sky included in other frames (320, 340).
[0094] For example, if the electronic device obtains location information related to a mountain being located to the right of a river and below the sky, it can obtain object properties (332, 342) for the mountain being located to the right of a river and below the sky included in other frames (320, 340).
[0095] Referring again to Figure 2,
[0096] In step 240, the electronic device can generate at least one object in the generation target area based on object characteristics of at least one object included in the current frame and object characteristics of at least one object included in other frames.
[0097] FIG. 4 is a flowchart illustrating an operating method of an electronic device according to one embodiment.
[0098] Steps 230 and 240 of FIG. 4 correspond to steps 230 and 240 of FIG. 2, respectively.
[0099] In step 241, the electronic device can obtain a correlation between an object characteristic of at least one object included in the current frame and an object characteristic of at least one object included in other frames.
[0100] A specific method for an electronic device according to one embodiment to obtain a correlation between an object characteristic of at least one object included in a current frame and an object characteristic of at least one object included in other frames will be described with reference to FIG. 5.
[0101] FIG. 5 is a diagram illustrating a method for an electronic device according to one embodiment to obtain a correlation between an object characteristic of at least one object included in a current frame and an object characteristic of at least one object included in other frames.
[0102] An electronic device according to one embodiment may include a first artificial intelligence model (500).
[0103] An electronic device according to one embodiment can input a plurality of frames included in a video as input to a first artificial intelligence model (500).
[0104] An electronic device according to one embodiment may set object properties of at least one object included in a current frame to a query matrix (query, 510).
[0105] An electronic device according to one embodiment may set object properties of at least one object included in a current frame as a query matrix (510) in the form of [1xNxC].
[0106] N represents the number of patches in the video.
[0107] C represents the number of channels of an object's characteristics. Here, a channel can mean a type of object characteristic.
[0108] An electronic device according to one embodiment can set object characteristics of at least one object included in a current frame and object characteristics of at least one object included in other frames as a key matrix (key, 520) and a value matrix (value, 530).
[0109] An electronic device according to one embodiment may set object characteristics of at least one object included in a current frame and object characteristics of at least one object included in other frames as a key matrix (520) in the form of [TxCxN].
[0110] An electronic device according to one embodiment can set object characteristics of at least one object included in a current frame and object characteristics of at least one object included in other frames as a value matrix (530) in the form of [TxNxC].
[0111] T represents the number of frames contained in the video.
[0112] An electronic device according to one embodiment can obtain an attention score matrix (540) based on a query matrix (510) and a key matrix (520).
[0113] An electronic device according to one embodiment can obtain an attention score matrix (540) based on a matrix multiplication operation.
[0114] An electronic device according to one embodiment can obtain an attention score matrix (540) in the form of [TxNxN] based on a matrix multiplication operation.
[0115] An electronic device according to one embodiment can obtain an attention weight matrix (550) by performing scaling and normalization on an attention score matrix (540).
[0116] An electronic device according to one embodiment can obtain an output (560) by multiplying an attention weight matrix (550) by a value matrix (530).
[0117] An electronic device according to one embodiment can obtain an output (560) by multiplying an attention weight matrix (550) by a value matrix (530) based on a matrix multiplication operation.
[0118] An electronic device according to one embodiment can obtain an output (560) in the form of [TxNxC] by multiplying an attention weight matrix (550) by a value matrix (530) based on a matrix multiplication operation.
[0119] The attention score matrix (540), attention weight matrix (550) and / or output (560) may indicate a correlation between at least one object included in the current frame and at least one object included in other frames.
[0120] Accordingly, an electronic device according to one embodiment sets an object characteristic of at least one object included in a current frame as a query matrix, and sets the object characteristic of at least one object included in the current frame and the object characteristic key matrix and value matrix of at least one object included in other frames, thereby obtaining a correlation between the object characteristic of at least one object included in the current frame and the object characteristic of at least one object included in other frames.
[0121] Referring again to Figure 4,
[0122] In step 243, the electronic device can generate at least one object in the generation target area based on the correlation.
[0123] An electronic device according to one embodiment may generate at least one object in a generation target area based on object characteristics of at least one object included in a current frame and object characteristics of at least one object included in other frames having a high degree of similarity based on correlation.
[0124] FIG. 6 is a flowchart illustrating an operation method of an electronic device according to one embodiment.
[0125] Steps 230 and 240 of FIG. 6 correspond to steps 230 and 240 of FIG. 2, respectively.
[0126] In step 242, the electronic device can spatially align object characteristics of at least one object included in the current frame and object characteristics of at least one object included in other frames in the generation target area.
[0127] Here, spatial alignment may mean aligning object properties of at least one object contained in the current frame and object properties of at least one object contained in other frames to the target region of the generation.
[0128] A method for an electronic device according to one embodiment to spatially align object characteristics of at least one object included in a current frame and object characteristics of at least one object included in other frames in a generation target area will be described with reference to FIG. 7.
[0129] FIG. 7 is a drawing for explaining an operating method of an electronic device according to one embodiment.
[0130] An electronic device according to one embodiment can determine a generation target area (731, 733) for expanding an image for a current frame.
[0131] An electronic device according to one embodiment can obtain object characteristics (732, 734) of at least one object included in a current frame and object characteristics (722, 724, 742, 744) of at least one object included in other frames.
[0132] An electronic device according to one embodiment can spatially align object characteristics (732, 734) of at least one object included in a current frame and object characteristics (722, 724, 742, 744) of at least one object included in other frames in a generation target area (731, 733).
[0133] In one embodiment, an electronic device may align the object properties of at least one object included in a current frame, which correspond to a 4x4 patch centered at (a2, b2) in the current frame, and the object properties of at least one object included in other frames, which correspond to a 4x4 patch centered at (a1, b1) and / or (a3, b3) in the other frames, and the generation target area is set to a 4x4 patch centered at (a0, b0), the object properties of at least one object included in the current frame corresponding to the 4x4 patch at the location (a2, b2) and the object properties of at least one object included in other frames corresponding to the 4x4 patch at the location (a1, b1) and / or (a3, b3) in the 4x4 patch centered at the generation target area (a0, b0). However, the sizes and positions of the patches of the object properties included in the current frame and / or the object properties included in other frames are not limited thereto.
[0134] For example, if the characteristics of a lecture object included in the current frame are characteristics (732) corresponding to a 4x4 patch centered at (a2, b2) in the current frame, and the characteristics of a lecture object included in other frames are characteristics (722, 742) corresponding to a 4x4 patch centered at (a1, b1) and / or (a3, b3) in other frames, and the generation target area is determined to be a patch (731) centered at (a0, b0) and having a size of 4x4, the electronic device can align the characteristics of the lecture object included in the current frame corresponding to the 4x4 patch at the location (a2, b2) and the characteristics of the lecture object included in other frames corresponding to the 4x4 patch at the location (a1, b1) and / or (a3, b3) to the patch (731) centered at the location (a0, b0) and having a size of 4x4.
[0135] For example, if the object characteristic of a mountain included in the current frame is a characteristic (734) corresponding to a 4x4 sized patch centered at (c2, d2) in the current frame, and the object characteristic of a mountain included in other frames is a characteristic (724, 744) corresponding to a 4x4 sized patch centered at (c1, d1) and / or (c3, d3) in other frames, and the target area for generation is determined (733) to be a patch centered at (c0, d0) and having a size of 4x4, the object characteristic (734) of the mountain included in the current frame corresponding to the 4x4 sized patch at the location (c2, d2) and the object characteristic (724, 744) of the mountain included in other frames corresponding to the 4x4 sized patch at the location (c1, d1) and / or (c3, d3) are determined to be a 4x4 sized patch centered at the location (c0, d0). It can be aligned to patch (733).
[0136] Referring again to Figure 6,
[0137] In step 244, the electronic device can generate at least one object in the generation target area based on spatial alignment of object characteristics of at least one object included in the current frame and object characteristics of at least one object included in other frames.
[0138] Figure 8 is a flowchart for explaining an operating method of an electronic device according to one embodiment.
[0139] Steps 230 and 240 of FIG. 8 correspond to steps 230 and 240 of FIG. 2, respectively.
[0140] In step 245, the electronic device can time-seriesly align object characteristics of at least one object included in the current frame and object characteristics of at least one object included in other frames based on motion information of the video.
[0141] Motion information may refer to information indicating pixel-level movement between consecutive frames. For example, motion information may refer to information on the direction and distance each pixel moved between the previous and current frames. Motion information may include, but is not limited to, information about direction information and / or vector fields indicating pixel-level movement between consecutive frames.
[0142] Sequential alignment may mean aligning at least one object included in the current frame and at least one object included in other frames in output order based on change information of the position of at least one object included in the frames between consecutive frames.
[0143] An electronic device according to one embodiment can time-seriesly align object characteristics of at least one object included in a current frame and object characteristics of at least one object included in other frames based on an elementwise multiplication operation.
[0144] A specific method for an electronic device according to one embodiment to time-seriesly align object characteristics of at least one object included in a current frame and object characteristics of at least one object included in other frames based on motion information of a video will be described with reference to FIG. 9.
[0145] FIG. 9 is a diagram illustrating a method for an electronic device to create at least one object in a creation target area according to one embodiment.
[0146] An electronic device according to one embodiment may include a second artificial intelligence model (900).
[0147] An electronic device according to one embodiment may set object properties of at least one object included in a current frame to a query matrix (510).
[0148] The query matrix (510) of FIG. 9 corresponds to the query matrix (510) of FIG. 5.
[0149] An electronic device according to one embodiment can set the output (560) obtained in FIG. 5 as a key matrix (920) and a value matrix (930) of a second artificial intelligence model (900).
[0150] An electronic device according to one embodiment can concatenate (910) the components included in the output (560) obtained in FIG. 5 to set the output (560) obtained in FIG. 5 as a key matrix (920) and a value matrix (930) of a second artificial intelligence model (900).
[0151] An electronic device according to one embodiment can obtain an attention score matrix (940) based on a query matrix (510) and a key matrix (920).
[0152] An electronic device according to one embodiment can obtain an attention score matrix (940) based on elementwise multiplication.
[0153] An electronic device according to one embodiment can obtain an attention weight matrix (950) by performing scaling and normalization on an attention score matrix (940).
[0154] An electronic device according to one embodiment can obtain an output (960) by multiplying an attention weight matrix (950) by a value matrix (930) based on an element-wise multiplication operation.
[0155] The attention score matrix (940), attention weight matrix (950), and / or output (960) may represent a correlation on the time axis between at least one object included in the current frame and at least one object included in other frames.
[0156] The correlation on the time axis may mean the output order relationship between at least one object included in the current frame and at least one object included in other frames based on motion information of the video.
[0157] An embodiment related to a first artificial intelligence model (Fig. 5, 500) and a second artificial intelligence model (Fig. 9, 900) included in an electronic device according to one embodiment is described separately for convenience of explanation, and an embodiment related to a first artificial intelligence model (Fig. 5, 500) and a second artificial intelligence model (Fig. 9, 900) can operate based on one artificial intelligence model.
[0158] Referring again to Figure 8,
[0159] In step 247, the electronic device can generate at least one object in the generation target area based on object characteristics of at least one object included in the current frame and object characteristics of at least one object included in other frames being aligned in time series.
[0160] A specific method for generating at least one object in a generation target area based on the time-series alignment of object characteristics of at least one object included in a current frame and object characteristics of at least one object included in other frames by an electronic device according to one embodiment will be described with reference to FIG. 10.
[0161] FIG. 10 is a drawing for explaining an operating method of an electronic device according to one embodiment.
[0162] An electronic device according to one embodiment can determine a generation target area (1031, 1033) for expanding an image for a current frame.
[0163] An electronic device according to one embodiment can obtain object characteristics (1032, 1034) of at least one object included in a current frame and object characteristics (1022, 1024, 1042, 1044) of at least one object included in other frames.
[0164] An electronic device according to one embodiment can time-seriesly align object characteristics (1032, 1034) of at least one object included in a current frame and object characteristics (1022, 1024, 1042, 1044) of at least one object included in other frames in a generation target area (1031, 1033) based on motion information of a video.
[0165] An electronic device according to one embodiment can obtain a correlation on a time axis between at least one object included in a current frame and at least one object included in other frames based on the embodiment described above in FIG. 9.
[0166] The correlation on the time axis may mean the output order relationship between at least one object included in the current frame and at least one object included in other frames based on motion information of the video.
[0167] An electronic device according to one embodiment can identify, based on a correlation on a time axis between at least one object included in a current frame and at least one object included in other frames, which object properties of at least one object included in the current frame and / or other frames are object properties of an object included in a frame immediately preceding the current frame (1022, 1024), an object property included in the current frame (1032, 1034), or an object property included in a frame immediately following the current frame (1042, 1044).
[0168] An electronic device according to one embodiment can generate at least one object in a generation target area in a time series manner based on an identification result.
[0169] FIG. 11 is a block diagram illustrating components of an electronic device according to one embodiment.
[0170] An electronic device (1100) according to one embodiment may include a memory (1120) storing a program or at least one instruction and at least one processor (1110) executing at least one instruction stored in the memory (1120).
[0171] The memory (1120) may store various data, programs, or applications for driving and controlling the electronic device (1110) according to one embodiment. The memory (1120) may include, for example, a non-volatile memory including at least one of a flash memory type, a hard disk type, a multimedia card micro type, a card type memory (e.g., an SD or XD memory, etc.), a ROM (Read-Only Memory), and an EEPROM (Electrically Erasable Programmable Read-Only Memory), a PROM (Programmable Read-Only Memory), and a volatile memory such as a RAM (Random Access Memory) or an SRAM (Static Random Access Memory).
[0172] The memory (1120) may store instructions, data structures, and program codes that can be read by the processor (1110). In the following embodiments, the processor (1110) may be implemented by executing instructions or codes of a program stored in the memory (1120).
[0173] The processor (1110) is a configuration that controls a series of processes to allow the electronic device (1110) to operate according to the embodiments described below, and may be composed of one or more processors.
[0174] The processor (1110) may be composed of hardware components that perform arithmetic, logic, and input / output operations and signal processing. One or more processors included in the processor (1110) may be circuitry such as a System on Chip (SoC), an Integrated Circuit (IC), etc. The processor (210) may be composed of at least one of, for example, a Central Processing Unit (CPU), a microprocessor, a Graphic Processing Unit (GPU), Application Specific Integrated Circuits (ASICs), Digital Signal Processors (DSPs), Digital Signal Processing Devices (DSPDs), Programmable Logic Devices (PLDs), and Field Programmable Gate Arrays (FPGAs), but is not limited thereto.
[0175] The processor (1110) can write data to the memory (1120), read data stored in the memory (1120), and process data according to predefined operation rules, particularly by executing a program or at least one command stored in the memory (1120).
[0176] According to one embodiment, at least one processor (1110) executes at least one command stored in a memory (1120), so that the electronic device (1100) can determine an area in which an image is to be expanded for a current frame among a plurality of frames included in a video as a generation target area.
[0177] The current frame refers to a frame in the video output to the display that has an area where image expansion is possible.
[0178] For example, the current frame may mean the frame containing letterboxing if the output video contains letterboxing.
[0179] The current frame can change over time based on motion information in the video.
[0180] For example, if a video contains frame t1, frame t2, and frame t3, the current frame at time t1 may be frame t1, the current frame at time t2 may be frame t2, and the current frame at time t3 may be frame t3. However, this is not limited thereto.
[0181] Motion information may refer to information indicating pixel-level movement between consecutive frames. For example, motion information may refer to information on the direction and distance each pixel moved between the previous frame and the current frame. Motion information may include, but is not limited to, information about direction information and / or vector fields indicating pixel-level movement between consecutive frames.
[0182] The target generation area may refer to the area in which the image can be expanded in the video output to the display.
[0183] For example, if the video output to the current frame contains a letter box, the target area to be generated may mean the letter box, but is not limited thereto.
[0184] According to one embodiment, at least one processor (1110) executes at least one command stored in a memory (1120), so that the electronic device (1100) can determine a target region for generation based on comparing an aspect ratio of a video and an aspect ratio of a display on which the video is output.
[0185] For example, when the electronic device (1100) wants to output a video having a width-to-height ratio of 4:3 to a display having a width-to-height ratio of 16:10 by having at least one processor (1110) execute at least one command stored in the memory (1120), the electronic device can determine the letter boxes formed on the left and right sides of the video to be output as the generation target area.
[0186] For example, when the electronic device (1100) wants to output a video having a width-to-height ratio of 1:1 on a display having a width-to-height ratio of 16:9 by having at least one processor (1110) execute at least one command stored in the memory (1120), the electronic device may determine letter boxes formed on the left, right, top, and bottom of the output video as the generation target area. However, the present invention is not limited thereto.
[0187] According to one embodiment, at least one processor (1110) executes at least one command stored in a memory (1120), thereby enabling the electronic device (1100) to determine a target area to be generated in units of at least one pixel or patch. A patch may be composed of multiple pixels.
[0188] According to one embodiment, at least one processor (1110) executes at least one command stored in a memory (1120), thereby enabling the electronic device (1100) to obtain object characteristics of at least one object included in a current frame.
[0189] According to one embodiment, at least one processor (1110) executes at least one command stored in a memory (1120), thereby enabling the electronic device (1100) to recognize an object in units of at least one pixel or patch.
[0190] For example, if the current frame contains content related to mountains, rivers, and the sky, the current frame may include the mountains, rivers, and the sky as objects. For example, if the current frame contains content related to people, desks, and food, the current frame may include people, desks, and food as objects.
[0191] For example, if the current frame contains content related to a person, the current frame may include as an object at least one of a patch consisting of pixels constituting a person or a set of pixels constituting a person.
[0192] For example, if the current frame contains content related to a river, the current frame may contain as an object at least one of a pixel constituting the river or a patch consisting of a set of pixels constituting the river.
[0193] Object characteristics may include at least one of location information or visual information of at least one object included in the current frame.
[0194] Location information can include information indicating the positional relationships between objects. For example, location information can include information about which objects are located above, below, left, and right of a specific object, and the order in which the objects are arranged.
[0195] The visual information may include at least one of color information, texture information, shape information, pattern information, or size information of at least one object.
[0196] Color information may include information about the (R, G, B) values of at least one object.
[0197] The texture information may include at least one of information about the roughness, glossiness, or hardness of the surface of at least one object.
[0198] Shape information may include edge information of at least one object.
[0199] The pattern information may include geometric pattern information of at least one object.
[0200] The size information may include relative size information between at least one object.
[0201] According to one embodiment, at least one processor (1110) executes at least one command stored in a memory (1120), thereby enabling the electronic device (1100) to identify at least one object included in the current frame.
[0202] According to one embodiment, at least one processor (1110) executes at least one command stored in a memory (1120), thereby enabling the electronic device (1100) to identify the type of at least one object included in the current frame.
[0203] According to one embodiment, at least one processor (1110) executes at least one command stored in a memory (1120), thereby enabling the electronic device (1100) to obtain location information indicating a location relationship between identified objects.
[0204] According to one embodiment, at least one processor (1110) executes at least one command stored in a memory (1120), thereby allowing the electronic device (1100) to store information about a plurality of frames included in a video in the memory.
[0205] According to one embodiment, at least one processor (1110) executes at least one instruction stored in a memory (1120), thereby causing the electronic device (1100) to generate a plurality of frames, including frame t0, frame t1, frame t2, ..., frame tn.
[0206] Other frames refer to the remaining frames (t0, t2, t3,..., tn) excluding the current frame (t1) among multiple frames.
[0207] According to one embodiment, at least one processor (1110) executes at least one command stored in a memory (1120), so that the electronic device (1100) can determine a target region for generating an image to be expanded for a current frame among a plurality of frames included in a video.
[0208] According to one embodiment, at least one processor (1110) executes at least one command stored in a memory (1120), thereby enabling the electronic device (1100) to obtain object characteristics of at least one object included in a current frame.
[0209] For example, the electronic device (1100) can obtain the characteristics of a lecture object included in the current frame by having at least one processor (1110) execute at least one command stored in the memory (1120).
[0210] Here, the river object properties may include location information related to being located on the left side of the mountain and being located below the sky.
[0211] Here, the lecture object characteristics may include at least one of lecture color information, texture information, shape information, pattern information, or size information.
[0212] For example, by having at least one processor (1110) execute at least one command stored in a memory (1120), the electronic device (1100) can obtain object characteristics of a mountain included in the current frame.
[0213] Here, the object properties of the mountain may include location information related to being located on the right side of the river and being located below the sky.
[0214] Here, the object characteristics of the mountain may include at least one of color information, texture information, shape information, pattern information, or size information of the mountain.
[0215] According to one embodiment, at least one processor (1110) executes at least one command stored in a memory (1120), so that the electronic device (1100) can obtain object characteristics of at least one object included in frames other than the current frame among a plurality of frames based on object characteristics of at least one object included in the current frame.
[0216] Object characteristics may include at least one of positional information or visual information of the object.
[0217] According to one embodiment, at least one processor (1110) executes at least one command stored in a memory (1120), so that the electronic device (1100) can obtain object characteristics of at least one object included in frames other than the current frame among a plurality of frames based on location information of at least one object included in the current frame.
[0218] For example, when at least one processor (1110) executes at least one command stored in the memory (1120), the electronic device (1100) can obtain lecture location information included in the current frame, and can obtain lecture object characteristics included in other frames based on the lecture location information.
[0219] For example, when at least one processor (1110) executes at least one command stored in the memory (1120), and the electronic device (1100) obtains location information related to a river being located to the left of a mountain and below the sky, the electronic device can obtain object characteristics for a river located to the left of a mountain and below the sky included in other frames.
[0220] For example, when at least one processor (1110) executes at least one command stored in the memory (1120), and the electronic device (1100) obtains location information related to a mountain being located to the right of a river and below the sky, the electronic device can obtain object characteristics for a mountain located to the right of a river and below the sky included in other frames.
[0221] According to one embodiment, at least one processor (1110) executes at least one command stored in the memory (1120), so that the electronic device (1100) can generate at least one object in the generation target area based on object characteristics of at least one object included in the current frame and object characteristics of at least one object included in other frames.
[0222] According to one embodiment, at least one processor (1110) executes at least one command stored in a memory (1120), thereby enabling the electronic device (1100) to obtain a correlation between an object characteristic of at least one object included in a current frame and an object characteristic of at least one object included in other frames.
[0223] According to one embodiment, at least one processor (1110) executes at least one command stored in a memory (1120), thereby allowing the electronic device (1100) to input a plurality of frames included in a video as input to the first artificial intelligence model.
[0224] According to one embodiment, at least one processor (1110) executes at least one command stored in a memory (1120), thereby enabling the electronic device (1100) to set object characteristics of at least one object included in a current frame as a query matrix.
[0225] According to one embodiment, at least one processor (1110) executes at least one command stored in a memory (1120), thereby enabling the electronic device (1100) to set object characteristics of at least one object included in the current frame as a query matrix in the form of [1xNxC].
[0226] N represents the number of patches in the video.
[0227] C represents the number of channels of an object's characteristics. Here, a channel can mean a type of object characteristic.
[0228] According to one embodiment, at least one processor (1110) executes at least one command stored in a memory (1120), so that the electronic device (1100) can set object characteristics of at least one object included in a current frame and object characteristics of at least one object included in other frames as a key matrix and a value matrix.
[0229] According to one embodiment, at least one processor (1110) executes at least one command stored in the memory (1120), so that the electronic device (1100) can set object characteristics of at least one object included in the current frame and object characteristics of at least one object included in other frames as a key matrix in the form of [TxCxN].
[0230] According to one embodiment, at least one processor (1110) executes at least one command stored in the memory (1120), so that the electronic device (1100) can set object characteristics of at least one object included in the current frame and object characteristics of at least one object included in other frames as a value matrix in the form of [TxNxC].
[0231] T represents the number of frames contained in the video.
[0232] According to one embodiment, at least one processor (1110) executes at least one instruction stored in a memory (1120), thereby enabling the electronic device (1100) to obtain an attention score matrix based on a query matrix and a key matrix.
[0233] According to one embodiment, at least one processor (1110) executes at least one instruction stored in a memory (1120), thereby enabling the electronic device (1100) to obtain an attention score matrix based on a matrix multiplication operation.
[0234] According to one embodiment, at least one processor (1110) executes at least one instruction stored in a memory (1120), thereby enabling the electronic device (1100) to obtain an attention score matrix in the form of [TxNxN] based on a matrix multiplication operation.
[0235] According to one embodiment, at least one processor (1110) executes at least one instruction stored in a memory (1120), thereby enabling the electronic device (1100) to obtain an attention weight matrix by performing scaling and normalization on an attention score matrix (540).
[0236] According to one embodiment, at least one processor (1110) executes at least one instruction stored in a memory (1120), thereby causing the electronic device (1100) to obtain an output by multiplying an attention weight matrix by a value matrix.
[0237] According to one embodiment, at least one processor (1110) executes at least one instruction stored in a memory (1120), so that the electronic device (1100) can obtain an output by multiplying an attention weight matrix by a value matrix based on a matrix multiplication operation.
[0238] According to one embodiment, at least one processor (1110) executes at least one instruction stored in a memory (1120), so that the electronic device (1100) can obtain an output in the form of [TxNxC] by multiplying an attention weight matrix by a value matrix based on a matrix multiplication operation.
[0239] The attention score matrix, attention weight matrix and / or output may represent a correlation between at least one object contained in the current frame and at least one object contained in other frames.
[0240] Accordingly, by executing at least one command stored in the memory (1120) by at least one processor (1110) according to one embodiment, the electronic device (1100) sets the object characteristics of at least one object included in the current frame as a query matrix, and sets the object characteristics of at least one object included in the current frame and the object characteristics of at least one object included in other frames as a key matrix and a value matrix, thereby obtaining a correlation between the object characteristics of at least one object included in the current frame and the object characteristics of at least one object included in other frames.
[0241] According to one embodiment, at least one processor (1110) executes at least one command stored in a memory (1120), thereby enabling the electronic device (1100) to generate at least one object in a generation target area based on a correlation.
[0242] According to one embodiment, at least one processor (1110) executes at least one command stored in the memory (1120), so that the electronic device (1100) can generate at least one object in a generation target area based on the object characteristics of at least one object included in the current frame and the object characteristics of at least one object included in other frames that have a high degree of similarity based on a correlation.
[0243] According to one embodiment, at least one processor (1110) executes at least one command stored in the memory (1120), so that the electronic device (1100) can spatially align object characteristics of at least one object included in the current frame and object characteristics of at least one object included in other frames in a generation target area.
[0244] Here, spatial alignment may mean aligning object properties of at least one object contained in the current frame and object properties of at least one object contained in other frames to the target region of the generation.
[0245] According to one embodiment, at least one processor (1110) executes at least one command stored in a memory (1120), thereby enabling the electronic device (1100) to determine a target region for generating an image to be expanded for the current frame.
[0246] According to one embodiment, at least one processor (1110) executes at least one command stored in a memory (1120), thereby enabling the electronic device (1100) to obtain object characteristics of at least one object included in a current frame and object characteristics of at least one object included in other frames.
[0247] According to one embodiment, at least one processor (1110) executes at least one command stored in the memory (1120), so that the electronic device (1100) can spatially align object characteristics of at least one object included in the current frame and object characteristics of at least one object included in other frames in a generation target area.
[0248] In one embodiment, when at least one processor (1110) executes at least one command stored in the memory (1120), the electronic device (1100) may align the object characteristics of at least one object included in a current frame, which correspond to a 4x4 patch centered at (a2, b2) in the current frame, and the object characteristics of at least one object included in other frames, which correspond to a 4x4 patch centered at (a1, b1) and / or (a3, b3) in the other frame, and the generation target area is set to a 4x4 patch centered at (a0, b0), the object characteristics of at least one object included in the current frame corresponding to the 4x4 patch at the location (a2, b2) and the object characteristics of at least one object included in other frames corresponding to the 4x4 patch at the location (a1, b1) and / or (a3, b3) in the 4x4 patch centered at the generation target area (a0, b0). However, the size and location of the patch of object properties included in the current frame and / or object properties included in other frames are not limited thereto.
[0249] For example, if the characteristics of a lecture object included in the current frame are characteristics corresponding to a 4x4 patch centered at (a2, b2) in the current frame, and the characteristics of a lecture object included in other frames are characteristics corresponding to a 4x4 patch centered at (a1, b1) and / or (a3, b3) in the other frames, and the generation target area is determined to be a 4x4 patch centered at (a0, b0), the electronic device can align the characteristics of the lecture object included in the current frame corresponding to the 4x4 patch at the location (a2, b2) and the characteristics of the lecture object included in other frames corresponding to the 4x4 patch at the locations (a1, b1) and / or (a3, b3) to the 4x4 patch centered at (a0, b0) as the generation target area.
[0250] For example, if the object characteristics of a mountain included in the current frame are characteristics corresponding to a 4x4 patch centered at (c2, d2) in the current frame, and the object characteristics of mountains included in other frames are characteristics corresponding to a 4x4 patch centered at (c1, d1) and / or (c3, d3) in the other frames, and the generation target area is determined to be a 4x4 patch centered at (c0, d0), the electronic device can align the object characteristics of the mountain included in the current frame corresponding to the 4x4 patch at the location (c2, d2) and the object characteristics of the mountains included in other frames corresponding to the 4x4 patch at the location (c1, d1) and / or (c3, d3) to the 4x4 patch centered at (c0, d0) as the generation target area.
[0251] According to one embodiment, at least one processor (1110) executes at least one command stored in the memory (1120), so that the electronic device (1100) can generate at least one object in a generation target area based on spatial alignment of object characteristics of at least one object included in a current frame and object characteristics of at least one object included in other frames.
[0252] According to one embodiment, at least one processor (1110) executes at least one command stored in a memory (1120), so that the electronic device (1100) can time-seriesly align object characteristics of at least one object included in a current frame and object characteristics of at least one object included in other frames based on motion information of a video.
[0253] Motion information may refer to information indicating pixel-level movement between consecutive frames. For example, motion information may refer to information on the direction and distance each pixel moved between the previous frame and the current frame. Motion information may include, but is not limited to, information about direction information and / or vector fields indicating pixel-level movement between consecutive frames.
[0254] Sorting in a time series may mean arranging at least one object included in the current frame and at least one object included in other frames in output order based on change information of the position of at least one object included in the frames between consecutive frames.
[0255] According to one embodiment, at least one processor (1110) executes at least one command stored in a memory (1120), so that the electronic device (1100) can time-seriesly align object characteristics of at least one object included in a current frame and object characteristics of at least one object included in other frames based on an elementwise multiplication operation.
[0256] According to one embodiment, at least one processor (1110) executes at least one command stored in a memory (1120), thereby enabling the electronic device (1100) to set object characteristics of at least one object included in a current frame as a query matrix of a second artificial intelligence model.
[0257] According to one embodiment, at least one processor (1110) executes at least one instruction stored in the memory (1120), thereby allowing the electronic device (1100) to set the output obtained in FIG. 5 as a key matrix and a value matrix of the second artificial intelligence model.
[0258] According to one embodiment, at least one processor (1110) executes at least one command stored in a memory (1120), so that the electronic device (1100) can concatenate the components included in the output obtained in FIG. 5, thereby setting the output obtained in FIG. 5 as a key matrix and a value matrix of the second artificial intelligence model.
[0259] According to one embodiment, at least one processor (1110) executes at least one instruction stored in a memory (1120), thereby enabling the electronic device (1100) to obtain an attention score matrix based on a query matrix and a key matrix.
[0260] According to one embodiment, at least one processor (1110) executes at least one instruction stored in a memory (1120), thereby enabling the electronic device (1100) to obtain an attention score matrix based on an elementwise multiplication operation.
[0261] According to one embodiment, at least one processor (1110) may execute at least one instruction stored in a memory (1120) so that the electronic device (1100) may obtain an attention weight matrix by performing scaling and normalization on an attention score matrix.
[0262] According to one embodiment, at least one processor (1110) executes at least one instruction stored in a memory (1120), so that the electronic device (1100) can obtain an output by multiplying an attention weight matrix by a value matrix based on an element-wise multiplication operation.
[0263] The attention score matrix, attention weight matrix and / or output may represent a temporal correlation between at least one object contained in the current frame and at least one object contained in other frames.
[0264] The correlation on the time axis may mean the output order relationship between at least one object included in the current frame and at least one object included in other frames based on motion information of the video.
[0265] According to one embodiment, at least one processor (1110) executes at least one command stored in the memory (1120), so that the electronic device (1100) can generate at least one object in the generation target area based on the time-series alignment of object characteristics of at least one object included in the current frame and object characteristics of at least one object included in other frames.
[0266] According to one embodiment, at least one processor (1110) executes at least one command stored in a memory (1120), thereby enabling the electronic device (1100) to determine a target region for generating an image to be expanded for the current frame.
[0267] According to one embodiment, at least one processor (1110) executes at least one command stored in a memory (1120), thereby enabling the electronic device (1100) to obtain object characteristics of at least one object included in a current frame and object characteristics of at least one object included in other frames.
[0268] According to one embodiment, at least one processor (1110) executes at least one command stored in the memory (1120), so that the electronic device (1100) can time-seriesly align object characteristics of at least one object included in the current frame and object characteristics of at least one object included in other frames in the generation target area based on motion information of the video.
[0269] According to one embodiment, at least one processor (1110) executes at least one command stored in a memory (1120), thereby enabling the electronic device (1100) to obtain a correlation on the time axis between at least one object included in a current frame and at least one object included in other frames.
[0270] The correlation on the time axis may mean the output order relationship between at least one object included in the current frame and at least one object included in other frames based on motion information of the video.
[0271] According to one embodiment, at least one processor (1110) executes at least one command stored in the memory (1120), so that the electronic device (1100) can identify, based on a correlation on a time axis between at least one object included in a current frame and at least one object included in other frames, which object properties of the at least one object included in the current frame and / or other frames are object properties of an object included in a frame immediately preceding the current frame, an object property of an object included in the current frame, or an object property of an object included in a frame immediately following the current frame.
[0272] According to one embodiment, at least one processor (1110) executes at least one command stored in a memory (1120), so that the electronic device (1100) can generate at least one object in a generation target area in a time series manner based on the identification result.
[0273] An operating method of an electronic device according to an embodiment of the present disclosure may be implemented in the form of program commands that can be executed through various computer means and recorded on a computer-readable medium. The computer-readable medium may include program commands, data files, data structures, etc., either singly or in combination. The program commands recorded on the medium may be those specially designed and configured for the present invention or may be those known and available to those skilled in the art of computer software. Examples of computer-readable recording media include magnetic media such as hard disks, floppy disks, and magnetic tapes, optical media such as CD-ROMs and DVDs, magneto-optical media such as floptical disks, and hardware devices specially configured to store and execute program commands, such as ROMs, RAMs, and flash memories. Examples of program commands include not only machine language codes generated by a compiler, but also high-level language codes that can be executed by a computer using an interpreter, etc.
[0274] Additionally, the operating method of the electronic device according to the disclosed embodiments may be provided as a computer program product. The computer program product may be traded as a product between a seller and a buyer.
[0275] A computer program product may include a software program and a computer-readable storage medium on which the software program is stored. For example, a computer program product may include a product in the form of a software program (e.g., a downloadable app) distributed electronically by an electronic device manufacturer or through an electronic marketplace (e.g., Google Play Store, App Store). For electronic distribution, at least a portion of the software program may be stored on a storage medium or temporarily created. In this case, the storage medium may be a storage medium of a manufacturer's server, an electronic marketplace server, or a relay server that temporarily stores the software program.
[0276] In a system comprising a server and a client device, the computer program product may include a storage medium of the server or a storage medium of the client device. Alternatively, if a third device (e.g., a smartphone) exists that is communicatively connected to the server or the client device, the computer program product may include a storage medium of the third device. Alternatively, the computer program product may include a software program itself that is transmitted from the server to the client device or the third device, or from the third device to the client device.
[0277] In this case, one of the server, the client device, and the third device may execute the computer program product to perform the method according to the disclosed embodiments. Alternatively, two or more of the server, the client device, and the third device may execute the computer program product to perform the method according to the disclosed embodiments in a distributed manner.
[0278] For example, a server (e.g., a cloud server or an artificial intelligence server, etc.) may execute a computer program product stored on the server, thereby controlling a client device in communication with the server to perform a method according to the disclosed embodiments.
[0279] Although the embodiments have been described in detail above, the scope of the present invention is not limited thereto, and various modifications and improvements made by those skilled in the art using the basic concept of the present invention defined in the following claims also fall within the scope of the present invention.
[0280] A method according to one embodiment may include a step of determining an area in which an image is to be expanded for a current frame among a plurality of frames included in the video as a generation target area.
[0281] A method according to one embodiment may include obtaining object properties of at least one object included in the current frame.
[0282] A method according to one embodiment may include a step of obtaining an object characteristic of at least one object included in frames other than the current frame among the plurality of frames, based on an object characteristic of at least one object included in the current frame.
[0283] A method according to one embodiment may include a step of generating at least one object in the generation target area based on an object characteristic of at least one object included in the current frame and an object characteristic of at least one object included in the other frames.
[0284] According to one embodiment, the object characteristics may include at least one of positional information or visual information of the at least one object.
[0285] A method according to one embodiment may include obtaining a correlation between an object characteristic of at least one object included in the current frame and an object characteristic of at least one object included in the other frames.
[0286] A method according to one embodiment may include a step of generating at least one object in the generation target area based on the correlation.
[0287] A method according to one embodiment may include a step of spatially aligning object properties of at least one object included in the current frame and object properties of at least one object included in the other frames to the generation target region.
[0288] A method according to one embodiment may include a step of generating at least one object in the generation target area based on spatial alignment of object characteristics of at least one object included in the current frame and object characteristics of at least one object included in the other frames.
[0289] A method according to one embodiment may include the step of identifying at least one object included in the current frame.
[0290] A method according to one embodiment may include obtaining location information between at least one of the identified objects.
[0291] A method according to one embodiment may include a step of obtaining object characteristics of at least one object included in the current frame based on positional information between the at least one object identified.
[0292] A method according to one embodiment may include a step of obtaining object characteristics of at least one object included in the other frames based on positional information of at least one object included in the current frame.
[0293] A method according to one embodiment may include a step of temporally aligning object characteristics of at least one object included in the current frame and object characteristics of at least one object included in the other frames based on motion information of the video.
[0294] A method according to one embodiment may include a step of generating at least one object in the generation target area based on object characteristics of at least one object included in the current frame and object characteristics of at least one object included in the other frames being aligned in time series.
[0295] A method according to one embodiment may include a step of aligning object properties of at least one object included in the current frame and object properties of at least one object included in the other frames in a time series manner based on an elementwise multiplication operation.
[0296] The method according to one embodiment may include a step of determining the target region for generation based on comparing an aspect ratio of the video and an aspect ratio of a display on which the video is output.
[0297] Visual information according to one embodiment may include at least one of color information, texture information, shape information, pattern information, or size information of at least one object.
[0298] An electronic device according to one embodiment may include a memory storing a program or at least one instruction and at least one processor executing at least one instruction stored in the memory.
[0299] According to one embodiment, the electronic device can determine an area in which an image is to be expanded for a current frame among a plurality of frames included in a video as a generation target area by having the at least one processor execute at least one command stored in the memory.
[0300] According to one embodiment, the electronic device can obtain object characteristics of at least one object included in the current frame by having the at least one processor execute at least one command stored in the memory.
[0301] According to one embodiment, the electronic device can obtain object characteristics of at least one object included in frames other than the current frame among the plurality of frames based on object characteristics of at least one object included in the current frame by having the at least one processor execute at least one command stored in the memory.
[0302] According to one embodiment, the electronic device can generate the at least one object in the generation target area based on object characteristics of at least one object included in the current frame and object characteristics of at least one object included in the other frames by having the at least one processor execute at least one command stored in the memory.
[0303] According to one embodiment, the object characteristics may include at least one of positional information or visual information of the at least one object.
[0304] According to one embodiment, the electronic device can obtain a correlation between an object characteristic of at least one object included in the current frame and an object characteristic of at least one object included in the other frames by having the at least one processor execute at least one command stored in the memory.
[0305] According to one embodiment, the electronic device can generate the at least one object in the generation target area based on the correlation by having the at least one processor execute at least one command stored in the memory.
[0306] According to one embodiment, the electronic device can spatially align object characteristics of at least one object included in the current frame and object characteristics of at least one object included in the other frames in the generation target area by having the at least one processor execute at least one command stored in the memory.
[0307] According to one embodiment, the electronic device can generate the at least one object in the generation target area based on spatial alignment of object characteristics of the at least one object included in the current frame and object characteristics of the at least one object included in the other frames by executing at least one command stored in the memory.
[0308] According to one embodiment, the electronic device can identify at least one object included in the current frame by having the at least one processor execute at least one command stored in the memory.
[0309] According to one embodiment, the electronic device can obtain location information between the at least one identified object by having the at least one processor execute at least one command stored in the memory.
[0310] According to one embodiment, the electronic device can obtain object characteristics of at least one object included in the current frame based on positional information between the at least one identified object by having the at least one processor execute at least one command stored in the memory.
[0311] According to one embodiment, the electronic device can obtain object characteristics of at least one object included in the other frames based on location information of at least one object included in the current frame by having the at least one processor execute at least one command stored in the memory.
[0312] According to one embodiment, the electronic device can temporally align object characteristics of at least one object included in the current frame and object characteristics of at least one object included in the other frames based on motion information of the video by having the at least one processor execute at least one command stored in the memory.
[0313] According to one embodiment, the electronic device can generate the at least one object in the generation target area based on the object characteristics of the at least one object included in the current frame and the object characteristics of the at least one object included in the other frames being aligned in time series by the at least one processor executing at least one command stored in the memory.
[0314] According to one embodiment, the electronic device can time-seriesly align object characteristics of at least one object included in the current frame and object characteristics of at least one object included in the other frames based on an elementwise multiplication operation by having the at least one processor execute at least one command stored in the memory.
[0315] According to one embodiment, the electronic device can determine the target region for generation based on comparing an aspect ratio of the video with an aspect ratio of a display on which the video is output, by having the at least one processor execute at least one command stored in the memory.
[0316] Visual information according to one embodiment may include at least one of color information, texture information, shape information, pattern information, or size information of at least one object.
Claims
1. In the method of expanding the output area of a video, A step of determining an area in which an image is to be expanded for the current frame among multiple frames included in the above video as a generation target area; A step of obtaining object properties of at least one object included in the current frame; A step of obtaining object characteristics of at least one object included in frames other than the current frame among the plurality of frames based on object characteristics of at least one object included in the current frame; and A step of generating at least one object in the generation target area based on object characteristics of at least one object included in the current frame and object characteristics of at least one object included in the other frames; A method wherein the object characteristics include at least one of positional information or visual information of the at least one object.
2. In the first paragraph, the step of creating at least one object, A step of obtaining a correlation between an object characteristic of at least one object included in the current frame and an object characteristic of at least one object included in the other frames; and A method comprising: a step of generating at least one object in the generation target area based on the correlation; 3. In the first or second paragraph, the step of creating at least one object comprises: A step of spatially aligning object characteristics of at least one object included in the current frame and object characteristics of at least one object included in the other frames in the generation target area; and A method comprising: generating at least one object in the generation target area based on spatial alignment of object characteristics of at least one object included in the current frame and object characteristics of at least one object included in the other frames.
4. In any one of the first to third paragraphs, the step of obtaining object properties of at least one object included in the current frame comprises: A step of identifying at least one object included in the current frame; A step of obtaining location information between at least one object identified above; and A method comprising: obtaining object characteristics of at least one object included in the current frame based on positional information between at least one object identified above.
5. In any one of paragraphs 1 to 4, the step of obtaining object characteristics of at least one object included in the other frames comprises: A method comprising: a step of obtaining object characteristics of at least one object included in the other frames based on position information of at least one object included in the current frame.
6. In any one of paragraphs 1 to 5, the step of creating at least one object comprises: A step of temporally aligning object characteristics of at least one object included in the current frame and object characteristics of at least one object included in the other frames based on motion information of the video; and A method comprising: a step of generating at least one object in the generation target area based on the object characteristics of at least one object included in the current frame and the object characteristics of at least one object included in the other frames being aligned in time series; 7. A computer-readable recording medium having recorded thereon a program for performing the method of any one of clauses 1 to 6 on a computer.
8. A computer program stored on a recording medium to perform the method of any one of claims 1 to 6, and performed by a computing device.
9. In the electronic device (1100), a memory (1120) storing a program or at least one instruction; and At least one processor (1110) that executes at least one instruction stored in the memory (1120), The electronic device (1100) executes at least one instruction stored in the memory (1120) by the at least one processor (1110). Among the multiple frames included in the video, the area to be expanded for the current frame is determined as the generation target area, Obtaining object properties of at least one object included in the current frame, Based on the object characteristics of at least one object included in the current frame, the object characteristics of at least one object included in frames other than the current frame among the plurality of frames are obtained, Generate at least one object in the generation target area based on object characteristics of at least one object included in the current frame and object characteristics of at least one object included in the other frames, An electronic device wherein the object characteristics include at least one of positional information or visual information of the at least one object.
10. In the 9th paragraph, the electronic device (1100) generates the at least one object by having the at least one processor (1110) execute at least one command stored in the memory (1120). Obtaining a correlation between an object characteristic of at least one object included in the current frame and an object characteristic of at least one object included in the other frames, An electronic device that creates at least one object in the creation target area based on the above correlation.
11. In the 9th or 10th paragraph, the electronic device (1100) generates the at least one object by executing at least one instruction stored in the memory (1120) by the at least one processor (1110). Spatially align object properties of at least one object included in the current frame and object properties of at least one object included in the other frames in the generation target area, An electronic device that creates at least one object in the generation target area based on spatial alignment of object characteristics of at least one object included in the current frame and object characteristics of at least one object included in the other frames.
12. In any one of the 9th to 11th paragraphs, the electronic device (1100) obtains object characteristics of at least one object included in the current frame by having the at least one processor (1110) execute at least one command stored in the memory (1120). Identify at least one object contained in the current frame, Obtaining location information between at least one object identified above, An electronic device that obtains object characteristics of at least one object included in the current frame based on positional information between at least one object identified above.
13. In any one of the 9th to 12th paragraphs, the electronic device (1100) obtains object characteristics of at least one object included in the other frames by having the at least one processor (1110) execute at least one command stored in the memory (1120). An electronic device that obtains object characteristics of at least one object included in the other frames based on location information of at least one object included in the current frame.
14. In any one of the 9th to 13th paragraphs, the electronic device (1100) generates the at least one object by having the at least one processor (1110) execute at least one command stored in the memory (1120). Based on the motion information of the video, the object characteristics of at least one object included in the current frame and the object characteristics of at least one object included in the other frames are temporally aligned, An electronic device that creates at least one object in the generation target area based on the object characteristics of at least one object included in the current frame and the object characteristics of at least one object included in the other frames being aligned in time series.
15. In the 14th paragraph, the electronic device (1100) sequentially arranges object characteristics of at least one object included in the current frame and object characteristics of at least one object included in the other frames by executing at least one command stored in the memory (1120). An electronic device that sequentially aligns object characteristics of at least one object included in the current frame and object characteristics of at least one object included in the other frames based on an elementwise multiplication operation.