Electronic device for providing video search result and operation method thereof
The electronic device generates thumbnail videos based on spatio-temporal information to address the challenge of intuitively identifying content in video search results, improving user experience by specifying the time and location of the search target within the video.
Patent Information
- Application Number
- PCT/KR2025/008793
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-07-30
- Filing Date
- 2025-06-24
- Publication Date
- 2026-02-05
AI Technical Summary
Conventional video search results do not allow users to intuitively confirm whether the searched content exists within the results, as thumbnails are typically displayed from a fixed frame or the entire video, failing to indicate the presence of the desired content along the timeline.
An electronic device generates thumbnail videos by extracting image frames based on spatio-temporal information of the search target object, specifying the time period and location within the video where the object is detected, allowing users to intuitively verify the presence of the content.
Enables users to efficiently identify the presence of search target content within video search results, enhancing user convenience by providing clear and intuitive thumbnail videos.
Smart Images

Figure KR2025008793_05022026_PF_FP_ABST
Abstract
Description
Electronic device for providing video search results and method of operation thereof
[0001] The present disclosure relates to an electronic device for searching for videos and providing search results, and a method for operating the same. Specifically, the present disclosure discloses an electronic device for searching for videos matching a search query entered by a user, and generating and displaying thumbnail videos representing the search results based on the time interval and location where the search target object is detected, and a method for operating the same.
[0002] Unlike the past, when information was primarily provided through text or still images, video has recently become the central tool for information provision and content consumption. In particular, with the increasing processing speed and storage capacity of mobile devices, the number of people storing video content, either directly filmed or downloaded from the network, has increased. This has led to a growing demand for video search capabilities among users.
[0003] Unlike still images like photographs, videos require users to navigate a video's timeline to view objects of interest (e.g., people, animals, objects, devices, buildings, etc.) or scenes related to specific actions or events. For example, a specific event, such as a toast in a video, can occur at various locations along the video's timeline, such as at the beginning, middle, or end. Therefore, even when users search for videos related to objects, actions, behaviors, situations, or events of interest, they face the challenge of not knowing at what point in the video the content they are searching for exists.
[0004] Conventional video search results are presented in the form of thumbnails (e.g., thumbnail images or thumbnail videos). Regardless of the user's interest in the searched content, thumbnails are displayed by displaying the first image frame of the video, playing image frames corresponding to a preset time (e.g., 5 seconds) within the video as a thumbnail video, or playing the entire video as a thumbnail video. The thumbnail images displayed in video search results alone do not allow users to intuitively confirm whether the searched content exists within the search results.
[0005] One aspect of the present disclosure discloses a method for providing video search results by an electronic device. An operating method of an electronic device according to an embodiment of the present disclosure may include a step of receiving a search query regarding a video to be searched for from a user. An operating method of an electronic device according to an embodiment of the present disclosure may include a step of searching for at least one video matching a keyword extracted from the search query among a plurality of videos pre-stored in a memory. An operating method of an electronic device according to an embodiment of the present disclosure may include a step of recognizing a time period in which a search target object corresponding to the keyword is detected from at least one searched video, and identifying a region in which the search target object is located within the recognized time period, thereby obtaining spatio-temporal information regarding the search target object. An operating method of an electronic device according to an embodiment of the present disclosure may include a step of generating a thumbnail video by extracting a plurality of image frames from at least one searched video based on the time period and region information. An operating method of an electronic device according to an embodiment of the present disclosure may include a step of displaying the generated thumbnail video.
[0006] Another aspect of the present disclosure discloses an electronic device that provides video search results. According to one embodiment of the present disclosure, the electronic device may include a user input unit that receives a user input for entering a search query; at least one processor including processing circuitry; a memory that stores one or more instructions; and a display. The one or more instructions are individually or collectively executed by the at least one processor, so that the electronic device can: search for at least one video matching a search query received through the user input unit among a plurality of videos previously stored in the memory. The one or more instructions are individually or collectively executed by the at least one processor (120), so that the electronic device can recognize a time period in which a search target object corresponding to a keyword extracted from the search query is detected from at least one searched video, and identify an area in which the search target object is located within the recognized time period, thereby obtaining spatio-temporal information about the search target object. The electronic device can generate a thumbnail video by extracting a plurality of image frames from at least one searched video based on the section and area information by individually or collectively executing one or more of the above commands by at least one processor (120). The electronic device can display the thumbnail video on the display by individually or collectively executing one or more of the above commands by at least one processor (120).
[0007] Another aspect of the present disclosure provides a computer program product including a computer-readable storage medium. The storage medium may include instructions readable by an electronic device (100) to perform the following operations: searching for at least one video matching a keyword extracted from a search query input by a user among a plurality of previously stored videos; recognizing a time period in which a search target object corresponding to the keyword is detected from the at least one searched video, and identifying a region in which the search target object is located within the recognized time period, thereby obtaining spatio-temporal information about the search target object; generating a thumbnail video by extracting a plurality of image frames from the at least one searched video based on the time and region information; and displaying the generated thumbnail video.
[0008] The present disclosure can be readily understood by the combination of the following detailed description and the accompanying drawings, wherein reference numerals refer to structural elements.
[0009] FIG. 1 is a diagram illustrating an operation of an electronic device according to one embodiment of the present disclosure to search for a video matching a search query and provide a video search result.
[0010] FIG. 2 is a flowchart illustrating a method for an electronic device according to one embodiment of the present disclosure to search for a video matching a search query and provide a video search result.
[0011] FIG. 3 is a block diagram illustrating components of an electronic device according to one embodiment of the present disclosure.
[0012] FIG. 4 is a flowchart illustrating a method for an electronic device to search for a video matching a search query according to an embodiment of the present disclosure.
[0013] FIG. 5 is a diagram illustrating an operation of an electronic device according to one embodiment of the present disclosure to search for a video matching a search query.
[0014] FIG. 6 is a flowchart illustrating a method for an electronic device according to an embodiment of the present disclosure to obtain spatio-temporal information regarding a search target object.
[0015] FIG. 7 is a diagram illustrating an operation of an electronic device according to one embodiment of the present disclosure to obtain section and area information regarding a search target object.
[0016] FIG. 8 is a diagram illustrating an operation of an electronic device according to an embodiment of the present disclosure to obtain section and area information when multiple search target objects are detected.
[0017] FIG. 9 is a flowchart illustrating a method for an electronic device to generate a thumbnail video according to one embodiment of the present disclosure.
[0018] FIG. 10 is a diagram illustrating an operation of an electronic device generating a thumbnail video according to one embodiment of the present disclosure.
[0019] FIG. 11 is a flowchart illustrating a method for an electronic device to generate a thumbnail video according to an embodiment of the present disclosure.
[0020] FIG. 12 is a diagram illustrating an operation of an electronic device generating a thumbnail video according to one embodiment of the present disclosure.
[0021] FIG. 13 is a flowchart illustrating a method for an electronic device according to an embodiment of the present disclosure to generate a thumbnail video using a plurality of time intervals during which a search target object is detected.
[0022] Figure 14 is a diagram schematically illustrating multiple time intervals in which a search target object is detected on the entire timeline of a video and a thumbnail video generated using multiple time intervals.
[0023] Figure 15 is a diagram schematically illustrating a thumbnail video generated using multiple time intervals in which a search target object is detected on the entire timeline of the video and at least two time intervals among the multiple time intervals.
[0024] Figure 16 is a diagram schematically illustrating a thumbnail video generated by adjusting the length of at least two time intervals among multiple time intervals in which a search target object is detected on the entire timeline of the video.
[0025] FIG. 17 is a diagram illustrating an operation of an electronic device according to one embodiment of the present disclosure to store or delete a thumbnail video.
[0026] FIG. 18 is a diagram illustrating an operation of an electronic device displaying a thumbnail video according to one embodiment of the present disclosure.
[0027] FIG. 19 is a diagram illustrating an operation of an electronic device according to one embodiment of the present disclosure to play an original video when a thumbnail video is selected.
[0028] The terms used in the embodiments of this specification have been selected from widely used, current terms, taking into account the functions of the present disclosure. However, these terms may vary depending on the intentions of those skilled in the art, precedents, the emergence of new technologies, etc. Furthermore, in certain cases, terms may be arbitrarily selected by the applicant, and in such cases, their meanings will be described in detail in the description of the relevant embodiments. Therefore, the terms used in this specification should not be defined simply as names of terms, but rather based on their meanings and the overall content of the present disclosure.
[0029] Singular expressions may include plural expressions unless the context clearly indicates otherwise. Terms used herein, including technical or scientific terms, have the same meaning as commonly understood by a person of ordinary skill in the art described herein.
[0030] Throughout this disclosure, when a part is said to "include" a component, this does not exclude other components, but rather implies the inclusion of other components, unless otherwise specifically stated. Furthermore, terms such as "part," "module," etc., used herein refer to a unit that processes at least one function or operation, which may be implemented in hardware or software, or a combination of hardware and software.
[0031] As used herein, the expression "configured to" can be used interchangeably with, for example, "suitable for," "having the capacity to," "designed to," "adapted to," "made to," or "capable of." The term "configured to" does not necessarily mean something is "specifically designed to" in terms of hardware. Instead, in some contexts, the expression "a system configured to" can mean that the system is "capable of" in conjunction with other devices or components. For example, the phrase "a processor configured to perform A, B, and C" can mean a dedicated processor for performing the operations (e.g., an embedded processor), or a general-purpose processor (e.g., a CPU or an application processor) that can perform the operations by executing one or more software programs stored in memory.
[0032] Additionally, when a component is referred to as being "connected" or "connected" to another component in the present disclosure, it should be understood that the component may be directly connected or connected to the other component, but may also be connected or connected via another component in between, unless otherwise specifically stated.
[0033] All functions or operations described in the present disclosure may be individually processed by a single processor and / or collectively processed by multiple processors. A single processor or a combination of multiple processors may include circuitry that performs processing, such as an Application Processor (AP), a Communication Processor (CP), a Graphical Processing Unit (GPU), a Neural Processing Unit (NPU), a Microprocessor Unit (MPU), a System on Chip (SoC), an Integrated Chip (IC), etc.
[0034] It should be understood that the blocks and combinations of flowcharts illustrated in this disclosure can be implemented by one or more computer programs containing computer-executable instructions. The one or more computer programs may be stored entirely in a single memory, or may be divided and stored across multiple different memories.
[0035] In this disclosure, "video" refers to image data containing an object moving over time. A video may include multiple image frames arranged and displayed in time sequence. In one embodiment of the present disclosure, a video may further include audio data, such as voice or music.
[0036] In this disclosure, a "search query" refers to a signal that a user requests to search for a desired object from storage or a database where video data is stored. In one embodiment of the present disclosure, the search query may include keywords indicating the content to be searched.
[0037] In this disclosure, a "keyword" refers to a word or phrase containing information about at least one of the search target objects, actions, behaviors, situations, or events that a user wishes to search for. In one embodiment of the present disclosure, a keyword may be obtained by analyzing a search query.
[0038] In this disclosure, a "search target object" refers to an object that a user wishes to search for. A search target object may include, for example, a person, an animal, an object, food, a device, or a building. However, the present disclosure is not limited thereto.
[0039] In the present disclosure, a "content identifier" refers to unique identification information for distinguishing a video from other videos. In one embodiment of the present disclosure, the content identifier may include classification or identification information of at least one of the contents included in the video, such as an object, an action, an action, a situation, or an event. In one embodiment of the present disclosure, the content identifier may also include information regarding a time period during which the object, action, action, situation, or event is detected. However, the present disclosure is not limited thereto, and the content identifier may further include identification information regarding a specific voice or music.
[0040] A content identifier may, for example, include metadata about a video. In this case, the content identifier may include tags related to object detection results (e.g., bounding boxes) or image categories (e.g., portrait, landscape, etc.).
[0041] The content identifier may be extracted at or before the time the video is acquired and stored with the video. However, this is not limited to the above, and the content identifier may also be extracted in real time through video analysis results by at least one processor of the electronic device at the time the search query is entered.
[0042] In the present disclosure, functions related to 'artificial intelligence' are operated through a processor and memory. The processor may be composed of one or more processors. In this case, one or more processors may be a general-purpose processor such as a CPU, AP, or DSP (Digital Signal Processor), a graphics-only processor such as a GPU or VPU (Vision Processing Unit), or an artificial intelligence-only processor such as an NPU. One or more processors control the processing of input data according to predefined operation rules or artificial intelligence models stored in memory. Alternatively, if one or more processors are artificial intelligence-only processors, the artificial intelligence-only processor may be designed with a hardware structure specialized for processing a specific artificial intelligence model.
[0043] The predefined operation rules or artificial intelligence models are characterized by being created through learning. Here, being created through learning means that the basic artificial intelligence model is trained using a learning algorithm using a plurality of learning data, thereby creating a predefined operation rules or artificial intelligence model set to perform a desired characteristic (or purpose). This learning may be performed on the device itself on which the artificial intelligence according to the present disclosure is performed, or may be performed through a separate server and / or system. Examples of the learning algorithm include, but are not limited to, supervised learning, unsupervised learning, semi-supervised learning, or reinforcement learning.
[0044] In the present disclosure, an 'artificial intelligence model' may be composed of a plurality of neural network layers. Each of the plurality of neural network layers has a plurality of weight values, and performs neural network operations through operations between the operation results of the previous layer and the plurality of weights. The plurality of weights of the plurality of neural network layers may be optimized based on the learning results of the artificial intelligence model. For example, the plurality of weights may be updated so that the loss value or cost value obtained from the artificial intelligence model during the learning process is reduced or minimized. The artificial neural network model may include a deep neural network (DNN), and examples thereof include, but are not limited to, a convolutional neural network, a recurrent neural network, a restricted Boltzmann machine, a deep belief network, a bidirectional recurrent deep neural network, or deep Q-networks.
[0045] Below, embodiments of the present disclosure are described in detail with reference to the attached drawings so that those skilled in the art can easily implement the present disclosure. However, the present disclosure may be implemented in various different forms and is not limited to the embodiments described herein.
[0046] Hereinafter, embodiments of the present disclosure will be described in detail with reference to the drawings.
[0047] FIG. 1 is a diagram illustrating an operation of an electronic device (100) according to one embodiment of the present disclosure to search for a video matching a search query and provide a video search result.
[0048] Referring to FIG. 1, the electronic device (100) may be a mobile device such as a smart phone, a tablet PC, a laptop computer, a digital camera, an e-book terminal, a digital broadcasting terminal, a PDA (Personal Digital Assistant), a PMP (Portable Multimedia Player), a navigation device, or an MP3 player. In one embodiment of the present disclosure, the electronic device (100) may also be a home appliance such as a TV, an air conditioner, a robot vacuum cleaner, or a clothes manager. However, the present disclosure is not limited thereto, and in one embodiment of the present disclosure, the electronic device (100) may also be implemented as a wearable device such as a smart watch, an eyeglass-type augmented reality device (e.g., AR (Augmented Reality) glasses), a head-mounted device (HMD), or a body-attached device (e.g., a skin pad).
[0049] Referring to FIG. 1, an electronic device (100) receives a search query input from a user (1) regarding a video to be searched (operation ①).
[0050] The electronic device (100) stores a plurality of previously stored videos (V1, V2, ..., V n ) searches for videos matching the search query entered by the user (1) (Action ②).
[0051] The electronic device (100) retrieves the searched video (V k ) obtains spatio-temporal information about the search target object (10) among the plurality of image frames constituting the image frame (Operation ③).
[0052] The electronic device (100) generates a thumbnail video (40) based on the acquired section and area information and displays the generated thumbnail video (40) (operation ④).
[0053] Hereinafter, referring to FIGS. 1 and 2 together, the electronic device (100) of the present disclosure displays a video (V) matching a search query entered by a user. k ) and search for the searched video (V k ) to obtain section and area information about a search target object, and to create and display a thumbnail video (40) representing the search result based on the obtained section and area information. The function and / or operation will be described in detail.
[0054] FIG. 2 is a flowchart illustrating a method in which an electronic device (100) according to one embodiment of the present disclosure searches for a video matching a search query and provides a video search result.
[0055] In step S210, the electronic device (100) receives a search query regarding a video to be searched from the user. In the present disclosure, a "search query" refers to a signal that the user requests to search for a target of search from storage or a database where video data is stored. In one embodiment of the present disclosure, the search query may include information regarding at least one of an object, action, behavior, situation, or event that is the target of the search.
[0056] The electronic device (100) can obtain keywords by analyzing a search query. In one embodiment of the present disclosure, a 'keyword' may be composed of a word or a phrase.
[0057] The search target object can be a word, such as "car," "dog," "ball," or "computer," or a phrase, such as "a jumping dog," "a person kicking a ball," or "a car in motion." The action or behavior of the search target object can include, for example, "jumping," "running," "waving," or "kicking a ball." The situation or event can include, for example, "a birthday party," "a golf swing," or "a toast." However, the search target object, action, action, situation, or event is not limited to the examples listed above.
[0058] The electronic device (100) may receive a user input for entering a search query through a user's keyboard input, mouse input, or touch input. However, the present invention is not limited thereto, and in one embodiment of the present disclosure, the electronic device (100) may also receive a search query input through a user's voice input. Referring to the embodiment illustrated in operation ① of FIG. 1, the electronic device (100) may receive a user's voice input for uttering a search target content, such as an object (e.g., a puppy), an action (e.g., a jump), a situation, or an event that the user (1) wishes to search for, through a microphone. The electronic device (100) may perform automatic speech recognition (ASR) to convert the acoustic signal of the received voice input into text, and may analyze the text through a natural language understanding (NLU) model to obtain keywords (e.g., 'puppy', 'jump') and intents (e.g., 'search') from the search query.
[0059] In one embodiment of the present disclosure, the electronic device (100) may receive a search query through a portion of an image or video.
[0060] Referring back to FIG. 2, in step S220, the electronic device (100) searches for at least one video that matches a keyword extracted from a search query among a plurality of previously stored videos. The plurality of videos may be acquired in advance before the user's search query is input and stored in a data storage (138, see FIG. 3) within the memory (130, see FIG. 3) of the electronic device (100). However, the present disclosure is not limited thereto, and in one embodiment of the present disclosure, the plurality of videos may be stored in a web storage or a cloud server.
[0061] In one embodiment of the present disclosure, a content identifier including information about at least one of an object, action, behavior, situation, or event included in the video may be pre-extracted from each of a plurality of videos and stored together with the plurality of videos in storage. In the present disclosure, a 'content identifier' refers to unique identification information for distinguishing a video from other videos, and may include, for example, metadata of the video. The content identifier may include tag information regarding an object detection result (e.g., a bounding box) or an image category (e.g., a portrait shot, a landscape shot, etc.).
[0062] The electronic device (100) can analyze a search query and convert the extracted keywords to correspond to the type of content identifier for each of the plurality of videos. The electronic device (100) can search for at least one video having a content identifier that matches the converted keyword by comparing the converted keywords with the content identifiers for each of the plurality of videos.
[0063] However, the present invention is not limited thereto, and content identifiers may not be extracted in advance from multiple videos. In one embodiment of the present disclosure, the electronic device (100) may perform analysis on each of the multiple videos using a deep neural network model, such as an object detection model, to detect a search target object, action, behavior, situation, or event corresponding to a search query, and register (e.g., 'tagging') a category of the object based on the detection result, thereby obtaining a content identifier.
[0064] In one embodiment of the present disclosure, the electronic device (100) may convert input search query data into a feature vector through preprocessing, extract a feature vector including information about content existing in each of a plurality of videos, and compare the feature vector converted from the search query data with the feature vector extracted about content in the videos to calculate a similarity. For example, the feature vector of each of the plurality of videos may be extracted per frame or from the entire video. Based on the calculated similarity, the electronic device (100) may search for a video having the highest similarity to the search query or may search for at least one video having a similarity exceeding a preset threshold.
[0065] Referring to operation ② of FIG. 1, the electronic device (100) stores a plurality of videos (V1, V2, ..., V) in a memory (130, see FIG. 3) and displays them through a display (140). n ) videos matching keywords extracted from search queries, for example, 'puppy' and 'jump' (V k ) can be searched. Figure 1 shows one video (V k ) is shown as being searched, but it is not limited to this and two or more multiple videos may be searched.
[0066] Referring back to FIG. 2, in step S230, the electronic device (100) recognizes a time period in which a search target object corresponding to a keyword is detected from at least one searched video and the location of the search target object within the time period, thereby obtaining spatio-temporal information. In one embodiment of the present disclosure, the electronic device (100) may perform object detection to identify a plurality of image frames in which a search target object is detected among the entire image frames constituting at least one searched video. In one embodiment of the present disclosure, the electronic device (100) may use an artificial intelligence model to detect a search target object from the entire image frames of at least one video and display a bounding box in the area in which the object is detected. In one embodiment of the present disclosure, when the electronic device (100) detects the entire image, such as a situation or event, the electronic device (100) may determine the entire image frame as an object detection area and display a bounding box in the entire image frame.
[0067] The electronic device (100) can recognize a time period including the start and end times of a plurality of image frames in which an object is detected. The electronic device (100) can obtain location information about an object region in which a search target object exists from each of the plurality of image frames within the time period. The electronic device (100) can obtain period and region information based on the location information about the time period and object region of the plurality of image frames.
[0068] In one embodiment of the present disclosure, the electronic device (100) can track a search target object whose position and size change through frame-to-frame tracking for a plurality of consecutive image frames within a time interval, thereby obtaining a spatio-temporal tube for the search target object.
[0069] Referring to operation ③ illustrated in FIG. 1, the electronic device (100) searches for the searched video (V k ) detects a search target object (10) (e.g., 'puppy') and an action (e.g., 'jump') among the entire image frames included in the image frame, and displays a bounding box (20), and displays a plurality of image frames (f1, f2, ..., f) in which the search target object (10) is detected. k ) and identify multiple image frames (f1, f2, ..., f k ) is the starting point (t1) and the ending point (t k) is the ending point. k ) can obtain time interval information including a plurality of image frames (f1, f2, ..., f) within the time interval. In addition, the electronic device (100) can obtain time interval information including a plurality of image frames (f1, f2, ..., f) within the time interval. k ) can obtain location information of an object area where a search target object (10) detected in the object area exists. The object area may be, for example, a display area of a bounding box (20), but is not limited thereto. The electronic device (100) may obtain a plurality of image frames (f1, f2, ..., f k ) can obtain interval and area information based on the time interval and object area location information.
[0070] In one embodiment of the present disclosure, the search target object (10) (in the embodiment shown in FIG. 1, a 'jumping puppy') may change in position and size in successive image frames, and the electronic device (100) may change the position and size of a plurality of image frames (f1, f2, ..., fk ) can be tracked to obtain a section and area tube (30) regarding the search target object (10).
[0071] In one embodiment of the present disclosure, when performing frame tracking, a plurality of image frames (f1, f2, ..., f) are used to improve processing speed. k ) Instead of tracking the entire image, multiple image frames (f1, f2, ..., f) are sampled. k ) can be used to selectively extract only some image frames to perform frame tracking.
[0072] Referring back to FIG. 2, in step S240, the electronic device (100) generates a thumbnail video by extracting a plurality of image frames from at least one video based on the acquired section and area information. Referring also to operation ④ of FIG. 1, the electronic device (100) extracts a plurality of image frames (f1, f2, ..., f) based on the section and area tube (30). k ) and extract multiple image frames (f1, f2, ..., f k ) can be used to generate a thumbnail video (40). In one embodiment of the present disclosure, the thumbnail video (40) comprises a plurality of image frames (f1, f2, ..., f k ) can be implemented as a thumbnail video that is sequentially listed and sequentially displayed over time. The electronic device (100) can display the thumbnail video (40) in a preview manner through the display (140).
[0073] In one embodiment of the present disclosure, the electronic device (100) may obtain an object area, which is an area occupied by a search target object, from each of a plurality of image frames extracted based on section and area information, and transform the object area based on size information of the width and height of a thumbnail image for preview. The electronic device (100) may perform image transformation, such as enlarging, reducing, or cropping the object area, based on preset sizes of the width and height of the thumbnail image for preview, for example. The electronic device (100) may generate a thumbnail video using images of the transformed object area.
[0074] In one embodiment of the present disclosure, when a plurality of time intervals in which a search target object is detected are acquired based on interval and area information, the electronic device (100) can select at least two time intervals among the plurality of time intervals to generate a thumbnail video.
[0075] In step S250 of FIG. 2, the electronic device (100) displays a thumbnail video (40, see FIG. 1). If no user input for selecting a searched video is received, the thumbnail video (40) may be automatically displayed via a graphical user interface (UI) that indicates the video search results. If a user input for selecting the thumbnail video (40) is received, the electronic device (100) may play the original video corresponding to the thumbnail video (40).
[0076] The electronic device (100) can store thumbnail videos together with the original videos in storage within the memory (130, see FIG. 3). In one embodiment of the present disclosure, the stored thumbnail videos are stored in the storage for a preset period of time (e.g., 24 hours or a week) and may be automatically deleted after the preset period of time has elapsed. However, the present disclosure is not limited thereto, and depending on the storage method, the thumbnail videos may be permanently stored in the storage unless deleted by the user.
[0077] Unlike still images like photographs, videos require users to navigate a video's timeline to view objects of interest (e.g., people, animals, objects, devices, buildings, etc.) or scenes related to specific actions or events. For example, a specific event, such as a toast in a video, can occur at various locations along the video's timeline, such as at the beginning, middle, or end. Therefore, even when users search for videos related to objects, actions, behaviors, situations, or events of interest, they face the challenge of not knowing at what point in the video the content they are searching for exists.
[0078] In conventional video search results, thumbnail images are displayed by displaying the first image frame of the video, playing image frames corresponding to a preset time (e.g., 5 seconds) within the video in the form of a thumbnail video, or playing the entire video as a thumbnail video, regardless of the search target content that the user is interested in. Therefore, there is a problem in that the user cannot intuitively confirm whether the search target content exists within the search results based solely on the thumbnail image displayed in the video search results.
[0079] The purpose of the present disclosure is to provide an electronic device (100) that provides video search results so that a user can intuitively check the search target content of a video that he or she wishes to search for, and an operating method thereof.
[0080] The electronic device (100) according to the embodiment illustrated in FIGS. 1 and 2 provides search results of a video that a user wishes to search for through a thumbnail video created by specifying a time period and an object area in which a search target object is displayed, thereby enabling the user to intuitively check the video search results and providing a technical effect that improves user convenience.
[0081] FIG. 3 is a block diagram illustrating the configuration of an electronic device (100) according to one embodiment of the present disclosure.
[0082] Referring to FIG. 3, the electronic device (100) may include a user input interface (110), a processor (120), a memory (130), and a display (140). The user input interface (110), the processor (120), the memory (130), and the display (140) may be electrically and / or physically connected to each other, respectively. In FIG. 3, only essential components for explaining the function and / or operation of the electronic device (100) are illustrated, and the components included in the electronic device (100) are not limited as illustrated in FIG. 3. In one embodiment of the present disclosure, the electronic device (100) may further include a camera that includes a lens module, an image sensor, and an image processing module, and is configured to capture an object to obtain a still image or a moving image. In one embodiment of the present disclosure, when the electronic device (100) is implemented as a portable device or a mobile device, the electronic device (100) may further include a battery that supplies driving power to the user input interface (110), the processor (120), and the display (140).
[0083] The user input interface (110) is configured to receive user input for entering a search query from a user. The user input interface (110) may be configured as a control panel including hardware components such as a keyboard, a key pad, a mouse, a trackball, a jog dial, a jog switch, or a touch pad, but is not limited thereto. In one embodiment of the present disclosure, the user input interface (160) may also be configured as a touch screen that receives touch input and displays a graphical user interface (GUI). In this case, the display (140) may be integrated with the user input interface (110).
[0084] The user input interface (110) may receive user input for entering a search query consisting of the name of a search target object or words, phrases, and / or sentences describing the object. For example, the user input interface (110) may receive a search query entered via a keyboard or touchscreen.
[0085] In one embodiment of the present disclosure, the electronic device (100) further includes a microphone and may receive a voice input for uttering a search query through the microphone. In this case, the microphone may receive a voice input uttered by a user and obtain a voice signal from the received voice input. The processor (120) may convert a sound component of the voice input received through the microphone into an acoustic signal and remove noise (e.g., non-speech components) from the acoustic signal to obtain the voice signal.
[0086] However, it is not limited thereto, and the user input interface (110) may also receive a portion of a photo or video as a search query.
[0087] The processor (120) can execute one or more instructions of a program stored in the memory (130). The processor (120) may be composed of hardware components that perform arithmetic, logic, and input / output operations, as well as image processing. Although the processor (120) is illustrated as a single element in FIG. 3 , it is not limited thereto. In one embodiment of the present disclosure, the processor (120) may be composed of one or more elements.
[0088] The processor (120) may include various processing circuits and / or multiple processors. For example, the term "processor" as used in this disclosure, including the claims, may include various processing circuits, including at least one processor. One or more processors in at least one processor may be configured to perform various functions described in this disclosure, individually and / or collectively, in a distributed fashion. As used herein, "processor," "at least one processor," and "one or more processors" may be configured to perform various functions. However, these terms encompass, without limitation, situations where one processor performs some of the functions and other processor(s) perform other parts of the functions, and situations where a single processor may perform all of the functions. Furthermore, at least one processor may include a combination of processors that perform various functions of the disclosed functions in a distributed manner. At least one processor may execute program instructions to achieve or perform various functions.
[0089] The processor (120) may be implemented as a general-purpose processor such as a CPU (Central Processing Unit), an AP (Application Processor), a DSP (Digital Signal Processor), a graphics-only processor such as a GPU (Graphics Processing Unit), a VPU (Vision Processing Unit), or an AI-only processor such as an NPU (Neural Processing Unit), for example. The processor (120) may be controlled to process input data according to predefined operating rules or an AI model. Alternatively, if the processor (120) is an AI-only processor, the AI-only processor may be designed with a hardware structure specialized for processing a specific AI model.
[0090] The memory (130) may be configured as at least one type of storage medium among, for example, a flash memory type, a hard disk type, a multimedia card micro type, a card type memory (e.g., SD or XD memory, etc.), a RAM (Random Access Memory), a SRAM (Static Random Access Memory), a ROM (Read-Only Memory), an EEPROM (Electrically Erasable Programmable Read-Only Memory), a PROM (Programmable Read-Only Memory), or an optical disk.
[0091] The memory (130) may store instructions related to functions and / or operations for the electronic device (100) to search for a video matching a search query, obtain spatio-temporal information about a search target object from the searched video, extract a plurality of image frames based on the spatio-temporal information from the searched video, generate a thumbnail video using the extracted plurality of image frames, and display the generated thumbnail video. In one embodiment of the present disclosure, the memory (130) may store at least one of instructions, an algorithm, a data structure, a program code, and an application program that can be read by the processor (120). The instructions, algorithms, data structures, and program codes stored in the memory (130) may be implemented in a programming or scripting language such as, for example, C, C++, Java, or an assembler.
[0092] The memory (130) may store instructions, algorithms, data structures, or program codes related to the video search module (132), the section and area information acquisition module (134), and the thumbnail video generation module (136). The 'module' included in the memory (130) refers to a unit that processes a function or operation performed by the processor (120), and this may be implemented as software such as instructions, algorithms, data structures, or program codes. In one embodiment of the present disclosure, the memory (130) may also include data storage (138).
[0093] The processor (120) may be implemented by executing instructions or program codes stored in the memory (130). Hereinafter, the functions and / or operations performed by the processor (120) by executing instructions or program codes of each of the plurality of modules stored in the memory (130) will be described in detail.
[0094] The video search module (132) is configured with commands or program codes for executing functions and / or operations for searching for videos matching a search query input by a user. In one embodiment of the present disclosure, the video search module (132) may be configured to extract keywords by analyzing a search query and search for at least one video matching the keyword among a plurality of videos. The processor (120) may search for at least one video matching the search query among a plurality of videos previously stored in the data storage (138) of the memory (130) by executing the commands or program codes of the video search module (132).
[0095] A 'search query' refers to a signal that a user requests to search for a target from a storage or database where video data is stored. In one embodiment of the present disclosure, the search query may include information regarding at least one of an object, action, behavior, situation, or event that is the target of the search. In one embodiment of the present disclosure, the processor (120) may obtain a search query from a user input received through the user input interface (110) and extract a keyword by analyzing the obtained search query. In one embodiment of the present disclosure, a 'keyword' may be composed of a word or a phrase.
[0096] In one embodiment of the present disclosure, the video search module (132) may include a language model, such as an automatic speech recognition (ASR) model configured to convert a voice signal into text, and a natural language understanding (NLU) model configured to analyze text to recognize intents and slots (e.g., 'keywords'). When the user input interface (110) receives a voice signal uttering a search query through a microphone, the processor (120) may perform ASR using the ASR model of the video search module (132) to convert the acoustic signal of the received voice input into text, and may obtain keywords from the search query by analyzing the text using the natural language understanding model (NLU model).
[0097] In one embodiment of the present disclosure, the video search module (132) may receive a search query as input through a part of an image or video, and search for a video matching the input image or video.
[0098] The processor (120) may search for at least one video that matches a keyword extracted from a search query among a plurality of previously stored videos. The plurality of videos may be acquired in advance before a user's search query is input and stored in a data storage (138) within the memory (130). However, the present disclosure is not limited thereto, and in one embodiment of the present disclosure, the plurality of videos may be stored in a web storage or a cloud server. In one embodiment of the present disclosure, a content identifier including information about at least one of an object, action, behavior, situation, or event included in the video may be extracted in advance from each of the plurality of videos and stored together with the plurality of videos in the storage. In the present disclosure, a 'content identifier' refers to unique identification information for distinguishing a video from other videos, and may include, for example, metadata of the video. The content identifier may include tag information regarding an object detection result (e.g., a bounding box) or an image category (e.g., a portrait shot, a landscape shot, etc.).
[0099] The processor (120) can convert keywords extracted from a search query to correspond to the type of content identifier for each of a plurality of videos. The processor (120) can search for at least one video having a content identifier that matches the converted keyword by comparing the converted keyword with the content identifier for each of the plurality of videos.
[0100] However, the present invention is not limited thereto, and content identifiers may not be extracted in advance from multiple videos. In one embodiment of the present disclosure, the processor (120) may perform analysis on each of the multiple videos using a deep neural network model, such as an object detection model, to detect a search target object, action, behavior, situation, or event corresponding to a search query, and register a category of the object (e.g., 'tagging') based on the detection result, thereby obtaining a content identifier.
[0101] A specific embodiment in which the processor (120) searches for at least one video by comparing a keyword and a content identifier will be described in detail in FIGS. 4 and 5.
[0102] In one embodiment of the present disclosure, the processor (120) may convert input search query data into a feature vector through preprocessing, and extract a feature vector including information about content existing in each of a plurality of videos. For example, the feature vector of each of the plurality of videos may be extracted per frame or from the entire video. The processor (120) may compare the feature vector converted from the search query with the feature vector extracted about content in the videos to calculate a similarity. The similarity between vectors may be calculated using, for example, cosine similarity, Jaccard similarity, Euclidean similarity, or Manhattan similarity measurement methods, or a known similarity measurement algorithm. Based on the calculated similarity, the processor (120) may search for a video having the maximum similarity with the search query, or may search for at least one video having a similarity exceeding a preset threshold.
[0103] The spatio-temporal information acquisition module (134) is configured with commands or program codes for executing a function and / or operation of acquiring a time period in which a search target object is detected from a searched video and location information in which the search target object exists. In the present disclosure, the 'spatio-temporal information' may include time period information of a plurality of image frames in which a search target object is detected among all image frames constituting the video and location information of an object region in which the search target object exists in each of the plurality of image frames within the time period. Here, the 'time period information' may include time information of a start time and an end time of a plurality of image frames in which the search target object is detected. The 'object region location information' may include two-dimensional location coordinate information of a region in which a search target object exists in each of the plurality of image frames. For example, the location information of the object area may include the first location coordinate information of the upper left (e.g., [x1, y1]) and the second location coordinate information of the lower right (e.g., [x2, y2]) among the two-dimensional location coordinate information of the bounding box of the search target object. The processor (120) may obtain the section and area information of the search target object from at least one searched video by executing the instructions or program code of the section and area information obtaining module (134).
[0104] In one embodiment of the present disclosure, the processor (120) may perform object detection to identify a plurality of image frames in which a search target object is detected among all image frames constituting at least one searched video. In one embodiment of the present disclosure, the processor (120) may detect a search target object from all image frames of at least one video using an artificial intelligence model, and display a bounding box in an area in which the object is detected. The 'artificial intelligence model' may be an artificial neural network model that is trained in advance to specify an object detection area in which an object is detected from an input image when a learning image is input, by a learning algorithm, and to output classification information according to the recognition result of the object included in the object detection area. The artificial neural network model may be implemented as, for example, a convolutional neural network model (CNN). However, the present invention is not limited thereto, and the artificial intelligence model may be implemented by at least one of, for example, a recurrent neural network (RNN), a restricted Boltzmann machine, a deep belief network, a bidirectional recurrent deep neural network, and a deep Q-network. In one embodiment of the present disclosure, when detecting an entire image, such as a situation or event, the processor (120) may determine the entire image frame as an object detection area and display a bounding box on the entire image frame.
[0105] The processor (120) can recognize a time period including the start and end times of a plurality of image frames in which an object is detected. The processor (120) can obtain location information about an object region in which a search target object exists from each of the plurality of image frames within the time period. The processor (120) can obtain period and region information based on the location information about the time period and object region of the plurality of image frames.
[0106] In one embodiment of the present disclosure, the processor (120) can track a search target object whose position and size change through frame-to-frame tracking for a plurality of image frames sequentially arranged within a time interval, thereby obtaining a spatio-temporal tube for the search target object. In this case, when performing frame tracking, instead of processing tracking for all of the plurality of image frames, the processor (120) can selectively perform frame tracking only for some of the frames by sampling some of the plurality of image frames in order to improve processing speed.
[0107] A specific embodiment in which the processor (120) obtains section and area information regarding the search target object from the searched video will be described in detail in FIGS. 6 to 8.
[0108] The thumbnail video generation module (136) is configured with commands or program codes for executing a function and / or operation of generating a thumbnail video using image frames of a specific section and a specific area among the searched video. In one embodiment of the present disclosure, the thumbnail video generation module (136) may selectively extract image frames corresponding to a specific time section and a specific object area from among all image frames constituting the video based on section and area information, and may generate a thumbnail video using the extracted image frames. The processor (120) may generate a thumbnail video by extracting image frames from at least one searched video by executing the commands or program codes of the thumbnail video generation module (136).
[0109] In one embodiment of the present disclosure, the processor (120) may obtain an object region of a search target object from each of a plurality of image frames extracted based on section and region information among the entire image frames of the searched video, and transform the object region based on size information of the width and height of a thumbnail image for preview. The processor (120) may perform image transformation, such as enlarging, reducing, or cropping the object region based on preset sizes of the width and height of the thumbnail image for preview, for example. The processor (120) may generate a thumbnail video using images of the transformed object region.
[0110] In one embodiment of the present disclosure, the processor (120) may scale an object region extracted from each of a plurality of image frames based on a preset width and height of a thumbnail image for preview, and generate a thumbnail video using images of the scaled object region. A specific embodiment in which the processor (120) scales an object region to generate a thumbnail video will be described in detail with reference to FIGS. 9 and 10 .
[0111] In one embodiment of the present disclosure, the processor (120) may set a region of interest by scaling the maximum values of the width and height of the object region using information about the preset width and height of the thumbnail image for preview, and align the center point of the object region with the center point of the region of interest. The processor (120) may crop each of a plurality of image frames to the size of the aligned region of interest to generate a thumbnail video. If a part of the region of interest is located outside the image frame according to the result of aligning the object region and the region of interest, the processor (120) may move the region of interest into the image frame based on offset information of the region of interest located outside the image frame. A specific embodiment in which the processor (120) sets a region of interest, aligns the object region and the region of interest, and crops a plurality of image frames according to the result of aligning to generate a thumbnail video will be described in detail with reference to FIGS. 11 and 12 .
[0112] In one embodiment of the present disclosure, when a plurality of time intervals in which a search target object is detected are acquired based on interval and area information, the processor (120) may select at least two time intervals among the plurality of time intervals to generate a thumbnail video. The processor (120) may arrange a plurality of image frames corresponding to each of the at least two selected time intervals, for example, in the order of at least one of time, the size of the object area, the size of the action and movement of the search target object, a preset user preference, or search history information. The processor (120) may generate a thumbnail video using the plurality of image frames listed in the arrangement order. A specific embodiment of generating a thumbnail video when there are a plurality of time intervals in which a search target object is detected by the processor (120) will be described in detail with reference to FIGS. 13 to 16.
[0113] Data storage (138) is a storage device within memory (130) that stores image data including multiple videos. Within data storage (138), multiple videos and content identifiers for each of the multiple videos may be stored together.
[0114] Data storage (138) may be composed of non-volatile memory. Non-volatile memory refers to a storage medium that stores and maintains information even when power is not supplied, and can use the stored information again when power is supplied. Non-volatile memory may include, for example, at least one of flash memory, a hard disk, a solid state drive (SSD), a multimedia card micro type, a card type memory (e.g., SD or XD memory), a read only memory (ROM), a magnetic memory, a magnetic disk, and an optical disk.
[0115] Although data storage (138) is depicted as a component included within memory (130) in FIG. 3, the present disclosure is not limited to the configuration shown in the drawing. In one embodiment of the present disclosure, data storage (138) may be configured as a database within the electronic device (100), which is a separate component from memory (130).
[0116] However, the present disclosure is not limited thereto, and in one embodiment of the present disclosure, the data storage (138) may be implemented as a web storage or cloud server that is accessible via a network and performs a storage function. In this case, the electronic device (100) further includes a communication interface configured to perform wired and wireless data communication, and may perform data transmission and reception by communicating with the web storage or cloud server through the communication interface. The processor (120) may access a plurality of videos from the web storage or cloud server, and search for at least one video that matches a search query among the plurality of videos.
[0117] The processor (120) may store the thumbnail video generated based on the section and area information in the data storage (138) together with the original video. In one embodiment of the present disclosure, the stored thumbnail video may be stored in the storage for a preset period of time (e.g., 24 hours, 3 days, or 1 week) and may be automatically deleted after the preset period of time has elapsed. However, the present invention is not limited thereto, and depending on the storage method, the thumbnail video may be permanently stored in the storage unless deleted by the user.
[0118] In one embodiment of the present disclosure, when receiving a user input for saving a thumbnail video, the processor (120) may permanently store the thumbnail video without deleting it from the data storage (138). A specific embodiment in which the processor (120) determines whether to save the thumbnail video based on the user input is described in detail in FIG. 17.
[0119] The display (140) is configured to display a thumbnail video under the control of the processor (120). The display (140) may be implemented as at least one of, for example, a liquid crystal display, a thin film transistor-liquid crystal display, an organic light-emitting diode, a flexible display, a 3D display, and an electrophoretic display.
[0120] If no user input is received to select a specific video from among the searched videos, the display (140) can automatically play a thumbnail video through a graphical user interface (UI) that provides video search results. In one embodiment of the present disclosure, the display (140) can play the entire thumbnail video or only a portion of the timeline. For example, if the thumbnail video is a 5-second video, the display (140) can play the thumbnail video for 5 seconds. However, the present invention is not limited thereto, and the display (140) can sequentially display image frames corresponding to a portion of the thumbnail video, for example, 2 seconds, in a time-series manner.
[0121] When a user input for selecting a thumbnail video is received, the display (140) can play the original video corresponding to the thumbnail video.
[0122] FIG. 4 is a flowchart illustrating a method for an electronic device (100) to search for a video matching a search query according to one embodiment of the present disclosure.
[0123] Steps S410 to S430 illustrated in FIG. 4 are operations that specify the operation of step S220 of FIG. 2. Step S410 illustrated in FIG. 4 may be performed after the operation of step S210 of FIG. 2 is performed. After the operation of step S430 illustrated in FIG. 4 is performed, step S230 of FIG. 2 may be performed.
[0124] FIG. 5 is a diagram illustrating an operation of an electronic device (100) according to one embodiment of the present disclosure to search for a video matching a search query.
[0125] Hereinafter, the function and / or operation of the electronic device (100) will be described in detail with reference to FIGS. 4 and 5.
[0126] In step S410 of FIG. 4, the electronic device (100) analyzes the search query to extract a keyword representing at least one of an object, action, behavior, situation, or event that is the target of the search from the search query. In the present disclosure, a 'keyword' represents a word or phrase that includes information about at least one of a search target object, action, behavior, situation, or event that a user wants to search for. In the present disclosure, a 'search target object' means an object that a user wants to search for, and may mean, for example, a person, an animal, an object, food, a device, or a building. The action or behavior of the search target object may include, for example, an action such as jumping, running, waving, or kicking a ball. The situation or event may include, for example, a birthday party, a golf swing, a toast, etc. However, the search target object, action, behavior, situation, or event is not limited to the examples listed above.
[0127] The electronic device (100) can extract keywords from a search query entered by a user through a keyboard, mouse, or touchscreen. In one embodiment of the present disclosure, the electronic device (100) can receive a voice signal uttering a search query through a microphone and extract keywords from the received voice signal. Referring to the embodiment illustrated in FIG. 5, the electronic device (100) includes a microphone (112) and can receive a voice input (500) such as "Search for a jumping puppy" from a user through the microphone (112). The microphone (112) can convert the sound component of the voice input (500) into an acoustic signal (510). The processor (120, see FIG. 3) of the electronic device (100) can convert an acoustic signal (510) into text by performing ASR (automatic speech recognition) using an ASR model, and can obtain keywords (520) from a search query by analyzing the text using a natural language understanding model (NLU model). In the embodiment illustrated in FIG. 5, the electronic device (100) can obtain 'puppy' and 'jump' as keywords (520).
[0128] However, it is not limited to the embodiment illustrated in FIG. 5, and the electronic device (100) may receive a search query through a portion of an image or video, perform image processing on the input image or video, or extract keywords by analyzing the image or video using an artificial intelligence model.
[0129] Referring back to FIG. 4, in step S420, the electronic device (100) converts the extracted keywords to correspond to the type of content identifier for each of the plurality of videos. The plurality of videos may be stored in the memory (130) of the electronic device (100). In one embodiment of the present disclosure, a content identifier including information about at least one of an object, action, behavior, situation, or event included in the video may be pre-extracted from each of the plurality of videos and stored together with the plurality of videos in the storage. In the present disclosure, a 'content identifier' refers to unique identification information for distinguishing a video from other videos, and may include, for example, metadata of the video. However, the present disclosure is not limited thereto, and the content identifier may include, for example, data of the image (e.g., a bounding box), voice, or audio signal type.
[0130] Content identifiers corresponding to multiple videos may not be extracted in advance. In one embodiment of the present disclosure, the electronic device (100) may perform analysis on each of the multiple videos using a deep neural network model, such as an object detection model, to detect objects, actions, behaviors, situations, or events from the multiple videos, and register (e.g., 'tagging') a category of the object based on the detection result, thereby obtaining a content identifier.
[0131] Referring to the embodiment illustrated in FIG. 5, a plurality of videos (530-1, 530-2, 530-3, ..., 530-n) may be stored in the data storage (138). Each of the plurality of videos (530-1, 530-2, 530-3, ..., 530-n) may be stored together with a content identifier (540-1, 540-2, 540-3, ..., 540-n) that includes identification information of the video. The content identifier (540-1, 540-2, 540-3, ..., 540-n) may include metadata such as an object detection result (e.g., a two-dimensional position coordinate value of a bounding box) or an image category (e.g., a portrait shot, a landscape shot, etc.). In the embodiment illustrated in FIG. 5, the first content identifier (540-1) for the first video (530-1) may include information about an object (e.g., a person) detected in the video and an action or action of the object (e.g., cheers). The second content identifier (540-2) for the second video (530-2) may include information about a dog detected in the video and the dog's action or action, such as jumping. The third content identifier (540-3) for the third video (530-3) may include information about a person detected in the video and the person's action or action, such as playing soccer. The n-th content identifier (540-n) for the n-th video (530-n) may include information about a person detected in the video and the person's action or action, such as eating.
[0132] The processor (120, see FIG. 3) of the electronic device (100) can convert the extracted keyword (520) to correspond to the type of the content identifier (540-1, 540-2, 540-3, ..., 540-n). In the embodiment illustrated in FIG. 5, the content identifier (540-1, 540-2, 540-3, ..., 540-n) is stored in the form of metadata corresponding to a plurality of videos (530-1, 530-2, 530-3, ..., 530-n), and thus the processor (120) can convert the keyword (520) to the type of the metadata. However, it is not limited thereto, and if the content identifier (540-1, 540-2, 540-3, ..., 540-n) is stored as a data type of an image or an audio signal, the processor (120) can convert the keyword (520) into an image or an audio signal.
[0133] Referring to step S430 of FIG. 4, the electronic device (100) searches for at least one video having a content identifier that matches the converted keyword by comparing the converted keyword with the content identifier. Referring also to the embodiment illustrated in FIG. 5, the processor (120, see FIG. 3) of the electronic device (100) can search for a video having a content identifier that matches the keyword (520) among the content identifiers (540-1, 540-2, 540-3, ..., 540-n) corresponding to a plurality of videos (530-1, 530-2, 530-3, ..., 530-n) stored in the data storage (138). For example, if the keyword (520) includes 'puppy' as a search target object and 'jump' as an action of the object, the processor (120) can identify a second content identifier (540-2) including information about puppy and jumping among the content identifiers (540-1, 540-2, 540-3, ..., 540-n), and search for a second video (530-2) having the second content identifier (540-2).
[0134] FIG. 6 is a flowchart illustrating a method for an electronic device (100) according to one embodiment of the present disclosure to obtain spatio-temporal information regarding a search target object.
[0135] Steps S610 to S640 illustrated in FIG. 6 are operations that specify the operation of step S230 of FIG. 2. Step S610 illustrated in FIG. 6 may be performed after the operation of step S220 of FIG. 2 is performed. After the operation of step S640 illustrated in FIG. 6 is performed, step S240 of FIG. 2 may be performed.
[0136] FIG. 7 is a diagram illustrating an operation of an electronic device (100) according to one embodiment of the present disclosure to obtain section and area information regarding a search target object (700).
[0137] Hereinafter, the function and / or operation of the electronic device (100) will be described in detail with reference to FIGS. 6 and 7.
[0138] In step S610 of FIG. 6, the electronic device (100) performs object detection to identify a plurality of image frames in which a search target object is detected among the entire image frames of at least one searched video. In one embodiment of the present disclosure, the electronic device (100) can detect a search target object corresponding to a keyword among the entire image frames constituting at least one searched video. In the present disclosure, a 'search target object' is an object that a user wishes to search by inputting a search query, and may include, for example, a person, an animal, an object, food, a device, or a building. In one embodiment of the present disclosure, the search target object may be obtained from a keyword. For example, if the keywords are 'puppy' and 'jump', the search target object may be 'puppy', which is an object among the keywords.
[0139] The electronic device (100) can detect a search target object from each of the entire image frames of at least one video using vision recognition technology. In the present disclosure, 'vision recognition' refers to image signal processing that inputs an image to an artificial intelligence model and detects an object from the input image, classifies the object into a specific category, or segments the object through inference using the artificial intelligence model. The 'artificial intelligence model' may be an artificial neural network model that is trained in advance to specify an object detection area in which an object is detected from an input image when a learning image is input, and output classification information according to the recognition result of the object included in the object detection area by a learning algorithm. The artificial neural network model may be implemented as, for example, a convolutional neural network model (CNN). However, the present invention is not limited thereto, and the artificial intelligence model may be implemented as at least one of, for example, a recurrent neural network (RNN), a restricted Boltzmann machine, a deep belief network, a bidirectional recurrent deep neural network, and a deep Q-network. The electronic device (100) may display an object detection area in which a search target object is detected among the entire image frames as a bounding box.
[0140] However, the present invention is not limited thereto, and the electronic device (100) may recognize a search target object from all image frames using a machine learning model among vision recognition technologies. The electronic device (100) may recognize an object using, for example, HOG feature extraction using a support vector machine (SVM) machine learning model, a BoW (Bag of Words) model using features such as SURF and MSER, or the Viola-Jones algorithm.
[0141] In one embodiment of the present disclosure, the electronic device (100) may detect the search target object only for the selected frames by sampling some of the image frames, rather than detecting the search target object from all of the image frames included in the video, in order to improve the processing speed.
[0142] Referring to the embodiment illustrated in FIG. 7, when the keywords extracted from the search query are 'puppy' and 'jump', the processor (120, see FIG. 3) of the electronic device (100) uses an artificial intelligence model to extract all image frames (f1 to f) included in the video. 10 ) can be detected as a search target object (700). The processor (120) processes the image frames (f1 to f 10 ) is applied as input to the artificial intelligence model, and inferencing is performed using the artificial intelligence model, thereby generating image frames (f1 to f 10 ) can detect a search target object (700) (e.g., 'a jumping puppy'). As a result of object recognition using an artificial intelligence model, the entire image frames (f1 to f 10) the search target object (700) can be detected only in the third image frame (f3) to the seventh image frame (f7). The processor (120) can display a bounding box (710) in the object area where the search target object (700) is detected in the third image frame (f3) to the seventh image frame (f7).
[0143] In FIG. 7, the entire image frames are illustrated as 10 frames, and the image frames in which the search target object (700) is detected are illustrated as 5 frames including the third image frame (f3) to the seventh image frame (f7), but the number of frames and the frames in which the search target object (700) is detected are merely examples for convenience of explanation, and the present disclosure is not limited to as illustrated in FIG. 7.
[0144] Referring back to FIG. 6, in step S620, the electronic device (100) recognizes a time section including a start time and an end time on the timeline of the identified plurality of image frames. Referring also to FIG. 7, the processor (120) of the electronic device (100) can recognize a third time point (t3) in which a third image frame (f3) in which a search target object (700) is detected in the timeline of a video is output as a start time point, and a seventh time point (t7) in which a seventh image frame (f7) is output as an end time point. The processor (120) can obtain time information regarding the third time point (t3), which is a start time point, and the seventh time point (t7), which is an end time point.
[0145] In step S630 of FIG. 6, the electronic device (100) obtains location information about an object region in which a search target object exists from each of a plurality of image frames within a time interval. The electronic device (100) can obtain a two-dimensional location coordinate value of a bounding box in which a search target object is detected from each of the plurality of image frames. Referring also to the embodiment illustrated in FIG. 7, the processor (120) of the electronic device (100) can obtain a two-dimensional location coordinate value of a bounding box (710) in which a search target object (700) is detected from each of a plurality of image frames (f3 to f7) output in a time interval (t3 to t7) in which a search target object (700) is detected. The location information of the object area may include, for example, a first location coordinate value (e.g., [x1, y1]) at the upper left and a second location coordinate value (e.g., [x2, y2]) at the lower right among the two-dimensional location coordinate values of the bounding box (710).
[0146] In step S640 of FIG. 6, the electronic device (100) obtains spatio-temporal information based on the time intervals and positional information on the object area of a plurality of image frames. In the present disclosure, the 'spatio-temporal information' may include the time interval information of a plurality of image frames recognized in step S620 among the entire image frames constituting the video and the positional information of the object area where the search target object exists obtained in step S630. Referring also to the embodiment illustrated in FIG. 7, the processor (120) of the electronic device (100) obtains spatio-temporal information based on the time intervals and positional information on the object area of the plurality of image frames constituting the video. 10) can obtain information about the time interval (t3 to t7) on the time line of the third image frame (f3) to the seventh image frame (f7) in which the search target object (700) is detected, and section and region information including the location information (e.g., two-dimensional location coordinate values) of the object region in which the search target object (700) is detected in the image frames (f3 to f7) in the time interval (t3 to t7).
[0147] The position and size of the search target object may change in image frames that are listed in time series. In one embodiment of the present disclosure, the electronic device (100) may track the search target object whose position and size change in a time section on a timeline of a plurality of image frames in which the search target object is detected through frame-to-frame tracking for a plurality of image frames, thereby obtaining a spatio-temporal tube for the search target object. In the embodiment illustrated in FIG. 7, the processor (120) may perform frame tracking for each of the third image frame (f3) to the seventh image frame (f7) to track the object area of the search target object (700), thereby obtaining a spatio-temporal tube (720). In order to improve processing speed when performing frame tracking, the processor (120) may sample some of the image frames (f3 to f7 in the embodiment of FIG. 7) and perform frame tracking only on the selected frames, instead of processing tracking for all of the image frames.
[0148] FIG. 8 is a diagram illustrating an operation of an electronic device (100) according to one embodiment of the present disclosure to obtain section and area information when multiple search target objects (801, 802) are detected.
[0149] Fig. 8 is identical to Fig. 7 except that there are multiple search target objects (801, 802) compared to the embodiment illustrated in Fig. 7, and therefore, any description overlapping with Fig. 7 will be omitted.
[0150] Referring to FIG. 8, if the keywords extracted from the search query are 'puppy' and 'jump', the electronic device (100) searches for a video matching the keywords and displays the entire image frames (f1 to f) included in the searched video. 10 ) can detect a plurality of search target objects (801, 802), i.e., jumping puppies. In the embodiment illustrated in FIG. 8, the jumping puppies may be a plurality including a first search target object (801) (e.g., a first puppy) and a second search target object (802) (e.g., a second puppy). However, the number of search target objects (801, 802) is exemplary and is not limited as illustrated in FIG. 8.
[0151] The electronic device (100) uses vision recognition technology to capture the entire image frames (f1 to f) of a video. 10 ) can detect multiple search target objects (801, 802) from each other. For example, the first search target object (801) is detected in the third image frame (f3) to the seventh image frame (f7), and the second search target object (802) is detected in the eighth image frame (f8) to the tenth image frame (f 10 ) can be detected.
[0152] The processor (120, see FIG. 3) of the electronic device (100) can obtain time interval information on the time line of image frames in which each of the plurality of search target objects (801, 802) is detected. For example, the first search target object (801) is detected in the first time interval between the third time point (t3) at which the third image frame (f3) is output and the seventh time point (t7) at which the seventh image frame (f7) is output, and the second search target object (802) is detected in the first time interval between the eighth time point (t8) at which the eighth image frame (f8) is output and the tenth image frame (f 10 ) is output at the 10th time point (t 10 ) can be detected in the second time interval between. The processor (120) detects the first time interval (t3 to t7) and the second time interval (t8 to t) for the first search target object (801) among the plurality of search target objects (801, 802). 10 ) can be obtained.
[0153] The processor (120) can obtain position information about an object region in which a plurality of search target objects (801, 802) exist within a time interval in which each of the plurality of search target objects (801, 802) is detected. For example, the processor (120) obtains two-dimensional position coordinate value information of a first object region (811) in which a first search target object (801) exists in image frames (f3 to f7) within a first time interval (t3 to t7), and obtains two-dimensional position coordinate value information of a first object region (811) in which a first search target object (801) exists within a second time interval (t8 to t 10 ) image frames (f8 to f 10 ) can obtain two-dimensional position coordinate value information of the second object area (812) in which the second search target object (802) exists.
[0154] The processor (120) can obtain section and area information for each of a plurality of search target objects (801, 802). For example, the processor (120) can obtain a first section and area tube (821) based on information about a first time interval (t3 to t7) and a first object area (811) for a first search target object (801). The processor (120) can obtain a second time interval (t8 to t) for a second search target object (802). 10 ) and the second object area (812) can be used to obtain the second section and area tube (822).
[0155] FIG. 9 is a flowchart illustrating a method for an electronic device (100) to generate a thumbnail video according to one embodiment of the present disclosure.
[0156] Steps S910 to S930 illustrated in FIG. 9 are operations that specify the operation of step S240 of FIG. 2. Step S910 illustrated in FIG. 9 may be performed after the operation of step S230 of FIG. 2 is performed. After the operation of step S930 illustrated in FIG. 9 is performed, step S250 of FIG. 2 may be performed.
[0157] FIG. 10 is a diagram illustrating an operation of an electronic device (100) according to one embodiment of the present disclosure to generate a thumbnail video.
[0158] Hereinafter, the function and / or operation of the electronic device (100) will be described in detail with reference to FIGS. 9 and 10.
[0159] In step S910 of FIG. 9, the electronic device (100) obtains an object region of a search target object from each of a plurality of image frames based on spatio-temporal information. The 'object region' refers to a region in which a search target object is detected and exists in the entire region of an image frame. In one embodiment of the present disclosure, the electronic device (100) can obtain two-dimensional position coordinate value information of an object region in which a search target object exists. Referring also to the embodiment illustrated in FIG. 10, the processor (120, see FIG. 3) of the electronic device (100) can obtain position information of object regions (1010-1 to 1010-5) in which a search target object (1000) exists from a plurality of image frames (f1 to f5). The position and size of the object regions (1010-1 to 1010-5) can change in the plurality of image frames (f1 to f5). For example, the position and size of the first object area (1010-1), which is a bounding box in which the search target object (1000) is detected in the first image frame (f1), may be different from the position and size of the second object area (1010-2) in the second image frame (f2).
[0160] In step S920 of FIG. 9, the electronic device (100) transforms the object area based on the size information of the thumbnail image for preview. In one embodiment of the present disclosure, the electronic device (100) may obtain information about the width and height based on the resolution of the thumbnail image for preview, and may transform the size of the object area using the information about the width and height. Transforming the object area may mean, for example, image processing that expands, crops, resizes, or scales the size of the object area. Referring also to the embodiment illustrated in FIG. 10, the thumbnail image (1020) for preview may be transformed into a width (W) according to the resolution. thumb ) and height (Hthumb ) can be preset. The processor (120) of the electronic device (100) sets the width (W) of the thumbnail image for preview thumb ) and height (H thumb ) can be scaled based on the size of the object area (1010-1 to 1010-5). Scaling factor (s xi , s yi ) can be calculated by the following equation 1.
[0161]
[0162] For example, if the first object area (1010-1) has a size of a first width (W1) and a first height (H1), the first scaling factor (s x1 , s y1 )Is can be calculated. In the same way, the scaling factor of the second object area (1010-2) to the fifth object area (1010-5) can be calculated by Equation 1.
[0163] The processor (120) calculates the scaling factor (s xi , s yi ) can be used to scale the size of the object area (1010-1 to 1010-5), thereby obtaining images of the scaled object area (1030-1 to 1030-5).
[0164] Referring again to FIG. 9, in step S930, the electronic device (100) generates a thumbnail video using images of the converted object area. Referring also to the embodiment illustrated in FIG. 10, the processor (120) of the electronic device (100) can generate a thumbnail video by chronologically listing images of the scaled object area (1030-1 to 1030-5).
[0165] In the embodiment shown in FIGS. 9 and 10, the electronic device (100) previews the object area (1010-1 to 1010-5) in which the search target object (1000) exists, with a width (W) of a thumbnail image (1020) for preview. thumb ) and height (H thumb ) is scaled based on the size of the object area (1030-1 to 1030-5), and a thumbnail video is generated using images of the scaled object area (1030-1 to 1030-5), so that a thumbnail video showing only a portion that the user wants to search among the entire image frames of the video can be provided as a search result. Accordingly, the electronic device (100) can enable the user to intuitively check the video search results and improve user convenience.
[0166] However, in the embodiments shown in FIGS. 9 and 10, the width (W) of the thumbnail image (1020) for preview is determined without considering the aspect ratio of the object area (1010-1 to 1010-5). thumb ) and height (H thumb ) is applied by considering only the size of the object area (1030-1 to 1030-5), so the ratio of the search target object (1000) in the thumbnail video may be distorted compared to the original.
[0167] FIG. 11 is a flowchart illustrating a method for an electronic device (100) to generate a thumbnail video according to one embodiment of the present disclosure.
[0168] Steps S1110 to S1130 illustrated in FIG. 11 are operations that embody the operation of step S920 of FIG. 9. Step S1110 of FIG. 11 may be performed after the operation of step S910 of FIG. 9 is performed. Steps S1140 to S1170 illustrated in FIG. 11 are operations that embody the operation of step S930 of FIG. 9.
[0169] FIG. 12 is a diagram illustrating an operation of an electronic device (100) according to one embodiment of the present disclosure to generate a thumbnail video.
[0170] Hereinafter, the function and / or operation of the electronic device (100) will be described in detail with reference to FIGS. 11 and 12.
[0171] In step S1110 of FIG. 11, the electronic device (100) obtains the maximum values of the width and height of the object region. In one embodiment of the present disclosure, the electronic device (100) may obtain two-dimensional position coordinate value information of the object region in which the search target object exists, and obtain the maximum values of the width and height of the object region based on the two-dimensional position coordinate value information. Referring also to the embodiment illustrated in FIG. 12, the processor (120, see FIG. 3) of the electronic device (100) may obtain the maximum values of the width and height of each of the object regions (1210-1 to 1210-5) in which the search target object (1200) exists from a plurality of image frames (f1 to f5). The size of the object area (1210-1 to 1210-5) changes in multiple image frames (f1 to f5), and accordingly, the maximum values of the width and height of the object area (1210-1 to 1210-5) may also be different for each of the multiple image frames (f1 to f5).
[0172] In step S1120 of FIG. 11, the electronic device (100) sets the region of interest by scaling the maximum values of the width and height of the object region using information about the width and height set for the thumbnail image for preview. Referring also to the embodiment illustrated in FIG. 12, the thumbnail image (1220) for preview may have a width (W) depending on the resolution. thumb ) and height (H thumb ) can be preset. The processor (120) of the electronic device (100) sets the width (W) of the thumbnail image for preview thumb ) and height (H thumb) by scaling the maximum values of the width and height of the object area (1210-1 to 1210-5). The scaling factor for determining the area of interest can be calculated by the following equation 2.
[0173]
[0174] For example, the maximum width of the first object area (1210-1) is W max1 , and the maximum height is H max1 In this case, the first scaling factor (scale 1) is can be calculated as follows. In the same manner, the scaling factor of the second object area (1210-2) to the fifth object area (1210-5) can be calculated by Equation 2.
[0175] Region of interest (ROI) per image frame f ) can be calculated by the following equation 3.
[0176]
[0177] Referring to Equation 3, the first region of interest (ROI1) is defined by the first scaling factor (scale 1) and the width of the thumbnail for preview (W thumb ) and height (H thumb ) can be determined based on the width (W) of the first region of interest (ROI1). For example, the width (W) of the first region of interest (ROI1) ROI1 ) is the size of the first scaling factor (scale 1) and the width of the thumbnail for preview (W thumb ) is calculated as the product of the height of the first region of interest (ROI1) (H ROI1 ) is the size of the first scaling factor (scale 1) and the height of the thumbnail for preview (H thumb ) can be calculated as a product of the widths of the second region of interest (ROI2) to the fifth region of interest (ROI5). Similarly, the width (W ROI n ) and height (H ROI n ) can be determined by Equation 3.
[0178] Referring again to FIG. 11, in step S1130, the electronic device (100) aligns the center point of the object area and the center point of the area of interest. Referring also to the embodiment illustrated in FIG. 12, the processor (120) of the electronic device (100) aligns the center point (C) of the object area (1210). obj ) and region of interest (ROI) f ) of the center point (C ROI ) can be aligned. As a result of aligning the center point, the object area (1210) and the region of interest (ROI) f ) can also be aligned.
[0179] In step S1140 of FIG. 11, the electronic device (100) determines whether the entire area of the region of interest is included within the image frame based on the alignment result.
[0180] If it is determined that the entire area of the region of interest is included within the image frame (step S1150), the electronic device (100) crops each of the plurality of image frames to the size of the aligned region of interest, thereby generating a thumbnail video.
[0181] If it is determined that a part of the region of interest is located outside the image frame (step S1160), the electronic device (100) moves the region of interest into the image frame based on offset information of the region of interest located outside the image frame. The offset information can be obtained based on two-dimensional position coordinate value information of the region of interest located outside the image frame among the entire region of interest. Referring to the embodiment illustrated in FIG. 12, the region of interest (ROI) f ) may exist outside the image frame (f). In this case, the processor (120) of the electronic device (100) may select a region of interest (ROI f ) based on the two-dimensional position coordinate values of the area existing outside the image frame (f) and the region of interest (ROI) f) can obtain the offset size (△x, △y) between the two. The processor (120) uses the offset size (△x, △y) to determine the region of interest (ROI) f ) can be moved within the image frame (f). As a result of the movement, the region of interest (ROI) f ')'s center point (C ROI ') is the 2D position coordinate value of the center point (C) of the object area (1210) obj ) can be moved by an offset size (△x, △y) in the X-axis and Y-axis compared to the two-dimensional position coordinate values.
[0182] In step S1170 of FIG. 11, the electronic device (100) crops each of the plurality of image frames to the size of the moved region of interest, thereby generating a thumbnail video. Referring also to the embodiment illustrated in FIG. 12, the processor (120) of the electronic device (100) crops each of the plurality of image frames (f1 to f5) to the size of the moved region of interest (ROI) f By cropping to the size of '), multiple object areas (1230-1 to 1230-5) can be obtained. The processor (120) can generate a thumbnail video by listing the cropped multiple object areas (1230-1 to 1230-5) in time series.
[0183] The electronic device (100) according to the embodiment illustrated in FIGS. 11 and 12 has a region of interest (ROI) scaled in proportion to the maximum values of the width and height of the object area (1210-1 to 1210-5). f ) and obtain the region of interest (ROI) f ) by cropping multiple image frames (f1 to f5) based on the size of the image frame, thereby providing a thumbnail video having an aspect ratio that reflects the original ratio of the search target object (1200).
[0184] FIG. 13 is a flowchart illustrating a method for an electronic device (100) according to one embodiment of the present disclosure to generate a thumbnail video using a plurality of time intervals in which a search target object is detected.
[0185] Steps S1310 to S1330 illustrated in FIG. 13 are operations that embody the operation of step S240 of FIG. 2. Step S1310 illustrated in FIG. 13 may be performed after the operation of step S230 of FIG. 2 is performed. After the operation of step S1330 illustrated in FIG. 13 is performed, step S250 of FIG. 2 may be performed.
[0186] In step S1310, if there are multiple time intervals in which the search target object is detected based on spatio-temporal information, the electronic device (100) selects at least two time intervals from among the multiple time intervals. In one embodiment of the present disclosure, the search target object may be detected in image frames from multiple different time intervals among the entire image frames of the video. In this case, the multiple time intervals may be discontinuous on the entire timeline of the video.
[0187] FIG. 14 is a diagram schematically illustrating a plurality of time intervals (T1, T2, T3) in which a search target object is detected on the entire timeline of a video, and a thumbnail video (M1, M2, M3) generated using the plurality of time intervals (T1, T2, T3). Referring to step S1310 of FIG. 13 together with FIG. 14, a search target object can be detected in image frames in a plurality of non-consecutive time intervals (T1, T2, T3) on the entire timeline of the video. In FIG. 14, the plurality of time intervals (T1, T2, T3) are illustrated as a total of three time intervals, but this is exemplary, and the present disclosure is not limited to that illustrated in FIG. 14.
[0188] The electronic device (100) can select at least two of the multiple time intervals or all of the multiple time intervals. Referring also to FIG. 14, the processor (120, see FIG. 3) of the electronic device (100) can select all of the multiple time intervals (T1, T2, T3). However, this is not limited thereto.
[0189] FIG. 15 is a diagram schematically illustrating a plurality of time intervals (T1, T2, T3) in which a search target object is detected on the entire timeline of a video, and a thumbnail video (M1, M2, M3) generated using at least two time intervals (T1, T3) among the plurality of time intervals (T1, T2, T3). Referring to step S1310 of FIG. 13 together with FIG. 15, the processor (120) of the electronic device (100) may select a first time interval (T1) and a third time interval (T3) among the plurality of time intervals (T1, T2, T3). However, the present invention is not limited to the case illustrated in FIG. 15, and the processor (120) may also select the first time interval (T1) and the second time interval (T2), or the second time interval (T2) and the third time interval (T3).
[0190] In step S1320 of FIG. 13, the electronic device (100) arranges a plurality of image frames corresponding to each of at least two selected time intervals in an order determined by at least one of time, size of an object area, action and movement size of a search target object, user preference, or search history information.
[0191] In one embodiment of the present disclosure, the electronic device (100) can list a plurality of image frames corresponding to at least two time intervals on a timeline based on the time order of at least two selected time intervals.
[0192] In one embodiment of the present disclosure, the electronic device (100) determines the order of the at least two time sections based on the order of the sizes of object areas in which a search target object exists in each of the plurality of image frames corresponding to at least two selected time sections, and arranges the plurality of image frames corresponding to the at least two time sections on a time line based on the determined order. For example, the electronic device (100) determines the order of the at least two time sections in the order of the sizes of the object areas, and arranges the plurality of image frames corresponding to the at least two time sections on a time line according to the determined order.
[0193] Referring to the embodiment illustrated in FIG. 15, the size of the object area of the search target object detected in the image frames corresponding to the third time interval (T3) may be larger than the size of the object area in the image frames corresponding to the first time interval (T1). In this case, the processor (120, see FIG. 3) of the electronic device (100) may arrange the third time interval (T3) in front of the first time interval (T1) on the timeline, and may arrange the image frames corresponding to the third time interval (T3) to be output before the image frames corresponding to the first time interval (T1). The electronic device (100) according to one embodiment of the present disclosure may provide an effect of displaying a high-quality thumbnail video that allows a user to perceive a higher definition by arranging an image frame having a large size of the search target object in the front part of the timeline to generate a thumbnail video.
[0194] In one embodiment of the present disclosure, the electronic device (100) may determine the order of the at least two time intervals on the timeline based on the degree to which the size or type of the action / movement of the search target object detected in each of the plurality of image frames corresponding to the at least two selected time intervals matches the keyword extracted from the search query. For example, if the keywords extracted from the search query input by the user are 'puppy' and 'jump', the electronic device (100) may identify an image frame having the highest similarity to the keyword among the action or movement (e.g., 'jump') of the search target object (e.g., 'puppy') detected from each of the plurality of image frames corresponding to the at least two time intervals, and may place the time interval in which the identified image frame is output earlier on the timeline than other time intervals. Referring also to the embodiment illustrated in FIG. 15, the action or movement of the search target object may be most similar to the keyword extracted from the search query in the third time interval (T3). In this case, the processor (120) of the electronic device (100) may place the third time interval (T3) before the first time interval (T1) on the timeline. For example, the magnitude of the action or movement of the search target object in the plurality of image frames corresponding to the third time interval (T3) may be greater than the magnitude of the action or movement of the search target object detected in the plurality of image frames corresponding to the first time interval (T1). In this case, the processor (120) may place the third time interval (T3) before the first time interval (T1) on the timeline.
[0195] In one embodiment of the present disclosure, the electronic device (100) can determine the order of at least two time intervals among a plurality of time intervals on a timeline based on user preference. The user preference can be preset by user input. However, the present invention is not limited thereto, and the user preference can also be determined based on search history information. In this case, the user preference can be determined, for example, based on at least one of the most frequently searched keyword, search target object, action, or behavior by analyzing the user's search history information. For example, if the user's search history information shows that 'puppy' is searched most frequently, the processor (120) of the electronic device (100) can place the time interval including the image frame in which the puppy is searched at the front of the at least two time intervals on the timeline. Referring to the embodiment illustrated in FIG. 15, when analyzing the user's search history information and finding that 'puppy' is searched with the highest frequency, the processor (120) can place the third time section (T3) containing the image frame in which the puppy is searched among at least two time sections (T1, T3) before the first time section (T1).
[0196] In one embodiment of the present disclosure, the electronic device (100) can adjust the time of each of at least two selected time intervals based on the size of the object area, the size of the action and movement of the search target object, or user preference. FIG. 16 is a diagram schematically illustrating a plurality of time intervals (T1, T2, T3) in which a search target object is detected on the entire timeline of a video, and a thumbnail video (M1, M2, M3) generated by adjusting the length of at least two time intervals (T1, T3) among the plurality of time intervals (T1, T2, T3). Referring also to FIG. 16, the processor (120) of the electronic device (100) can, for example, increase the length of a third time interval (T3) including an image frame having a maximum object area size among the at least two selected time intervals (T1, T3), and decrease the length of a first time interval (T1) including an image frame having a small object area size. In this case, the processor (120) can place the third time interval (T3) in front of the first time interval (T1) on the timeline.
[0197] For example, the processor (120) may increase the length of a third time interval (T3) including image frames in which the size of the action and movement of the search target object detected from the image frames is large, and may decrease the length of a first time interval (T1) including image frames in which the size of the action and movement is relatively small. For example, the processor (120) may increase the length of a third time interval (T3) including image frames in which the search target object has a high user preference (e.g., a search target object matching a keyword with a high search history), and may decrease the length of the first time interval (T1).
[0198] Referring back to FIG. 13, in step S1330, the electronic device (100) generates a thumbnail video using a plurality of image frames listed in an array order. Referring also to the timeline illustrated in FIG. 14, the processor (120) of the electronic device (100) generates a first video (M1) using a plurality of image frames included in a first time section (T1), generates a second video (M2) using a plurality of image frames included in a second time section (T2), generates a third video (M3) using a plurality of image frames included in a third time section (T3), and can generate a thumbnail video by merging the first video (M1) to the third video (M3). Referring to the timeline illustrated in FIG. 15, the processor (120) may generate a third video (M3) using a plurality of image frames included in a third time section (T3) arranged in the front on the timeline, generate a first video (M1) using a plurality of image frames included in a first time section (T1), and merge the third video (M3) and the first video (M1) in order on the timeline to generate a thumbnail video. Referring to the timeline illustrated in FIG. 16, the processor (120) may generate a third video (M3) using a plurality of image frames included in a third time section (T3) arranged in the front on the timeline and having an adjusted length, generate a first video (M1) using a plurality of image frames included in a first time section (T1) having an adjusted length, and generate a thumbnail video by merging the third video (M3) and the first video (M1) in order on the timeline.
[0199] The electronic device (100) according to the embodiment illustrated in FIGS. 13 to 16 selects at least two time intervals from among a plurality of time intervals, arranges the selected at least two time intervals in an order determined based on at least one of the size of an object area, the size of an action / movement of a search target object, user preference, or search history information as well as a time order, and generates a thumbnail video by listing a plurality of image frames in the arranged order, thereby providing a high-quality thumbnail video to a user, and thereby improving user convenience.
[0200] FIG. 17 is a diagram illustrating an operation of an electronic device (100) according to one embodiment of the present disclosure to store or delete thumbnail videos (1710-1, 1710-2, ..., 1710-n).
[0201] Referring to FIG. 17, a plurality of original videos (1700-1, 1700-2, ..., 1700-n) may be stored in the data storage (138). The electronic device (100) may store a plurality of thumbnail videos (1710-1, 1710-2, ..., 1710-n) generated corresponding to each of the plurality of original videos (1700-1, 1700-2, ..., 1700-n) in the data storage (138). In one embodiment of the present disclosure, the electronic device (100) may temporarily store the plurality of thumbnail videos (1710-1, 1710-2, ..., 1710-n) in the data storage (138) only for a preset time. The preset time may be, for example, 24 hours, 3 days, or a week, but is not limited to the above examples.
[0202] The electronic device (100) may display a graphical user interface (UI) that indicates search results for videos on a display (140). The graphical UI may include a search query (1720) and a plurality of thumbnail videos (1710-1, 1710-2, ..., 1710-n) for each of a plurality of videos searched for matching the search query (1720). A download icon (1730) may be displayed on each of the plurality of thumbnail videos (1710-1, 1710-2, ..., 1710-n) for receiving a user input for saving the thumbnail video.
[0203] When a user input regarding a download icon (1730) is received, a thumbnail video corresponding to the download icon (1730) for which the user input was received may be permanently stored in the data storage (138). In the embodiment illustrated in FIG. 17, the download icon (1730) displayed on the second thumbnail video (1710-2) among the plurality of thumbnail videos (1710-1, 1710-2, ..., 1710-n) is selected by the user input, and in this case, the electronic device (100) may store the second thumbnail video (1710-2) in the data storage (138). If the user does not delete the second thumbnail video (1710-2), the second thumbnail video (1710-2) may be permanently stored in the data storage (138).
[0204] The electronic device (100) can delete the remaining thumbnail videos (1710-1, 1710-3, ..., 1710-n) among the plurality of thumbnail videos (1710-1, 1710-2, ..., 1710-n) for which the download icon (1730) is not selected. In one embodiment of the present disclosure, the electronic device (100) can automatically delete the thumbnail videos (1710-1, 1710-3, ..., 1710-n) at a second time point when a preset time (e.g., 24 hours, 3 days, or 1 week, etc.) has elapsed from the first time point at which the thumbnail videos (1710-1, 1710-3, ..., 1710-n) are stored in the data storage (138).
[0205] However, this is not limited thereto. In one embodiment of the present disclosure, depending on the storage method, thumbnail videos (1710-1, 1710-3, ..., 1710-n) may not be automatically deleted, but may be deleted only by user input.
[0206] FIG. 18 is a diagram illustrating an operation of an electronic device (100) displaying a thumbnail video according to one embodiment of the present disclosure.
[0207] Referring to FIG. 18, the electronic device (100) may display a graphic user interface (UI) that indicates a search result of a video on the display (140). The graphic UI may include a search query (1820) and a plurality of thumbnail videos (1810-1, 1810-2, ..., 1810-n) for each of a plurality of videos searched for matching the search query (1820). When the electronic device (100) displays a plurality of thumbnail videos (1810-1, 1810-2, ..., 1810-n), the electronic device (100) may display a first image frame of each of the plurality of thumbnail videos (1810-1, 1810-2, ..., 1810-n). By displaying the first image frame of multiple thumbnail videos (1810-1, 1810-2, ..., 1810-n), the user can easily and intuitively determine whether a search target object exists in the videos displayed as video search results.
[0208] However, it is not limited thereto. In one embodiment of the present disclosure, the electronic device (100) can automatically play a plurality of thumbnail videos (1810-1, 1810-2, ..., 1810-n) in a preview manner as a video search result. In this case, the electronic device (100) can play the entire play time of the plurality of thumbnail videos (1810-1, 1810-2, ..., 1810-n). However, it is not limited thereto, and the electronic device (100) can play only a part of the entire play time of each of the plurality of thumbnail videos (1810-1, 1810-2, ..., 1810-n). For example, the electronic device (100) can play the first 5 seconds of the total play time of each of the plurality of thumbnail videos (1810-1, 1810-2, ..., 1810-n).
[0209] When a user input is received to select a specific thumbnail video among multiple thumbnail videos (1810-1, 1810-2, ..., 1810-n) displayed as a video search result, the electronic device (100) can play the original video corresponding to the thumbnail video. A specific embodiment of playing the original video will be described in detail in FIG. 19.
[0210] FIG. 19 is a diagram illustrating an operation of an electronic device (100) according to one embodiment of the present disclosure to play an original video (1900) when a thumbnail video is selected.
[0211] Referring to FIG. 19, when a user input for selecting one of a plurality of thumbnail videos displayed through a graphic UI indicating a video search result is received, the electronic device (100) may enlarge and play an original video (1900) corresponding to the selected thumbnail video on the entire screen area of the display (140). The electronic device (100) may sequentially display all image frames in chronological order starting from the first image frame of the original video (1900). However, the present disclosure is not limited thereto, and in one embodiment of the present disclosure, the electronic device (100) may selectively play only a plurality of image frames corresponding to one or a plurality of time sections used to generate a thumbnail video among the entire image frames of the original video (1900). When playing back multiple image frames corresponding to multiple time sections, the electronic device (100) may play back multiple image frames corresponding to a first time section among the multiple time sections, and then play back the remaining image frames of the original video (1900) until the time line ends. In one embodiment of the present disclosure, the electronic device (100) may sequentially play back multiple image frames corresponding to multiple time sections in the order of the time sections.
[0212] In one embodiment of the present disclosure, the electronic device (100) may display a graphical UI visualizing information about each section of the video while playing the original video (1900). Referring to the embodiment illustrated in FIG. 19, the graphical UI may include at least one of a section content information UI (1910), a timeline UI (1920), and a play bar UI (1930).
[0213] The section content information UI (1910) is an interface that visualizes and displays, in the form of characters or icons, content included in at least one time section used to generate a thumbnail video among the entire timeline of the original video (1900). The section content information UI (1910) can visualize and display information about a search target object, action, behavior, situation, or event included in a specific time section corresponding to a keyword extracted from a search query of the original video (1900). For example, the section content information UI (1910) can include a first section content information UI (1910-1) that visualizes an eating action of a search target object (e.g., a person), a second section content information UI (1910-2) that visualizes a specific action, and a third section content information UI (1910-3) that visualizes a crowd that is a search target object.
[0214] When a user input for selecting a section content information UI (1910) is received, the electronic device (100) can skip to a time section corresponding to the selected section content information UI among the entire timeline of the original video (1900) and play back image frames of the corresponding time section. For example, when a user input for selecting a second section content information UI (1910-2) is received, the electronic device (100) can skip to a time section matching the 'action' keyword among the entire timeline of the original video (1900) and play back only the image frames of the corresponding time section.
[0215] The timeline UI (1920) is a graphical interface that visualizes the start and end times of the time sections corresponding to the section content information UI (1910). The timeline UI (1920) can visualize the area occupied by each of at least one time section on the entire timeline of the original video (1900). When a user input for selecting the section content information UI (1910) is received, the electronic device (100) can display the time section corresponding to the selected section content information UI (1910) through the timeline UI (1920). When there are multiple time sections corresponding to the selected section content information UI (1910), the electronic device (100) can display the timeline UIs (1920-1, 1920-2) on the multiple time sections on the timeline. For example, when a user input is received to select a second section content information UI (1910-2) for the 'action' keyword, the electronic device (100) may display a first timeline UI (1920-1) on a first time section matching the 'action' keyword among the entire timeline of the original video (1900), and may display a second timeline UI (1920-2) on a second time section.
[0216] The Play Bar UI (1930) is a graphical interface that visually displays the point in time of the image frame being played within the entire timeline of the original video (1900).
[0217] In the embodiment illustrated in FIG. 19, the electronic device (100) provides graphic UIs (1910, 1920, 1930) that visually display a time interval in which a search target object that the user wants to search for is detected, content information about the time interval, and the current play time when playing an original video (1900), thereby enabling the user to easily and intuitively grasp various pieces of information within the video, thereby providing a technical effect that can improve user convenience.
[0218] One aspect of the present disclosure discloses a method for an electronic device (100) to provide video search results. The method for operating the electronic device (100) according to one embodiment of the present disclosure may include a step (S210) of receiving a search query regarding a video to be searched from a user. The method for operating the electronic device (100) according to one embodiment of the present disclosure may include a step (S220) of searching for at least one video matching a keyword extracted from the search query among a plurality of videos previously stored in a memory. The method for operating the electronic device (100) according to one embodiment of the present disclosure may include a step (S230) of recognizing a time period in which a search target object corresponding to the keyword is detected from at least one searched video, and identifying a region in which the search target object is located within the recognized time period, thereby obtaining spatio-temporal information regarding the search target object. An operating method of an electronic device (100) according to an embodiment of the present disclosure may include a step (S240) of generating a thumbnail video by extracting a plurality of image frames from at least one searched video based on section and area information. An operating method of an electronic device (100) according to an embodiment of the present disclosure may include a step (S250) of displaying the generated thumbnail video.
[0219] In one embodiment of the present disclosure, a content identifier including information about at least one of an object, an action, an action, a situation, or an event for the plurality of videos may be pre-extracted and stored together with the plurality of videos. The step (S220) of searching for at least one video matching the search query may include the step (S420) of converting a keyword extracted from the search query to correspond to a type of a content identifier for each of the plurality of videos; and the step (S430) of searching for at least one video having a content identifier matching the converted keyword by comparing the converted keyword with the content identifier.
[0220] In one embodiment of the present disclosure, the step of obtaining section and region information (S230) may include a step of performing object detection to identify a plurality of image frames in which a search target object is detected among all image frames constituting at least one searched video (S610). The step of obtaining section and region information (S230) may include a step of recognizing a time section including a start time and an end time of a timeline of the identified plurality of image frames (S620); and a step of obtaining position information on an object region in which the search target object exists from each of the plurality of image frames within the time section (S630). The step of obtaining section and region information (S230) may include a step of obtaining section and region information based on the position information on the time sections and object regions of the plurality of image frames (S640).
[0221] In one embodiment of the present disclosure, the step (S230) of obtaining the section and area information may include a step of obtaining a spatio-temporal tube for the search target object by tracking the search target object whose position and size change through frame-to-frame tracking for a plurality of consecutive image frames within a time interval.
[0222] In one embodiment of the present disclosure, the step of generating a thumbnail video (S240) may include a step of obtaining an object region of a search target object from each of a plurality of image frames based on section and area information (S910); and a step of transforming the object region based on size information of a thumbnail image for preview (S920). The step of generating a thumbnail video (S240) may include a step of generating a thumbnail video using images of the transformed object region (S930).
[0223] In one embodiment of the present disclosure, the step of transforming the object region (S920) may include a step of scaling the object region extracted from each of the plurality of image frames based on a preset width and height of a thumbnail image for preview. The step of generating the thumbnail video (S930) may include a step of generating the thumbnail video using images of the scaled object region.
[0224] In one embodiment of the present disclosure, the step of transforming the object region (S920) may include a step of obtaining a maximum value of the width and height of the object region (S1110); a step of setting a region of interest by scaling the maximum value of the width and height of the object region using information about a preset width and height of a thumbnail image for preview (S1120); and a step of aligning the center point of the object region and the center point of the region of interest (S1130). The step of generating a thumbnail video (S930) may include a step of generating a thumbnail video by cropping each of a plurality of image frames to the size of the aligned region of interest.
[0225] In one embodiment of the present disclosure, the operating method of the electronic device (100) may further include a step (S1050) of moving the region of interest into the image frame based on offset information of a region of the region of interest located outside the image frame, when a part of the region of interest is located outside the image frame according to the alignment result of the object region and the region of interest.
[0226] In one embodiment of the present disclosure, in the step (S240) of generating the thumbnail video, if a plurality of time intervals in which a search target object is detected are acquired based on the interval and area information, the electronic device (100) may select at least two time intervals among the plurality of time intervals to generate the thumbnail video.
[0227] In one embodiment of the present disclosure, the step of generating a thumbnail video (S240) may include a step of arranging a plurality of image frames corresponding to each of at least two selected time sections in an order determined by at least one of time, a size of an object area, an action and movement size of the search target object, a preset user preference, or search history information (S1320). The step of generating a thumbnail video (S240) may include a step of generating a thumbnail video using a plurality of image frames listed in an arrangement order (S1330).
[0228] Another aspect of the present disclosure discloses an electronic device (100) that provides video search results. According to one embodiment of the present disclosure, the electronic device (100) may include a user input unit (110) that receives a user input for entering a search query; at least one processor (120) including processing circuitry; a memory (130) that stores one or more instructions; and a display (140). The one or more instructions are individually or collectively executed by the at least one processor (120), thereby enabling the electronic device (100) to: search for at least one video matching a search query received through the user input unit (110) among a plurality of videos previously stored in the memory (130). By individually or collectively executing the one or more commands by at least one processor (120), the electronic device (100) can recognize a time period in which a search target object corresponding to a keyword extracted from a search query is detected from at least one searched video, and can obtain spatio-temporal information about the search target object by identifying a region in which the search target object is located within the recognized time period. By individually or collectively executing the one or more commands by at least one processor (120), the electronic device (100) can generate a thumbnail video by extracting a plurality of image frames from the at least one searched video based on the time period and region information. By individually or collectively executing the one or more commands by at least one processor (120), the electronic device (100) can display the thumbnail video on the display (140).
[0229] In one embodiment of the present disclosure, the one or more commands are individually or collectively executed by at least one processor (120), so that the electronic device (100) can perform object detection to identify a plurality of image frames in which a search target object is detected among all image frames constituting at least one searched video. The one or more commands are individually or collectively executed by at least one processor (120), so that the electronic device (100) can recognize a time section including a start time and an end time of a timeline of the identified plurality of image frames, obtain positional information about an object area in which a search target object exists from each of the plurality of image frames within the time section, and obtain section and area information based on the positional information about the time section and the object area of the plurality of image frames.
[0230] In one embodiment of the present disclosure, the one or more commands are individually or collectively executed by at least one processor (120), so that the electronic device (100) can track a search target object whose position and size change through frame-to-frame tracking for a plurality of consecutive image frames within a time interval, thereby obtaining a spatio-temporal tube for the search target object.
[0231] In one embodiment of the present disclosure, the one or more commands are individually or collectively executed by at least one processor (120), so that the electronic device (100) can obtain an object area of a search target object from each of a plurality of image frames based on section and area information, and transform the object area based on size information of a thumbnail image for preview. The one or more commands are individually or collectively executed by at least one processor (120), so that the electronic device (100) can generate a thumbnail video using images of the transformed object area.
[0232] In one embodiment of the present disclosure, the one or more commands are individually or collectively executed by at least one processor (120), so that the electronic device (100) can scale an object area extracted from each of a plurality of image frames based on a preset width and height of a thumbnail image for preview, and generate a thumbnail video using images of the scaled object area.
[0233] In one embodiment of the present disclosure, the electronic device (100) can set a region of interest by obtaining a maximum value of the width and height of an object region and scaling the maximum value of the width and height of the object region using information about a preset width and height of a thumbnail image for preview by individually or collectively executing one or more of the commands by at least one processor (120). The electronic device (100) can align a center point of the object region and a center point of the region of interest by individually or collectively executing one or more of the commands by at least one processor (120), and can generate a thumbnail video by cropping each of a plurality of image frames to the size of the aligned region of interest.
[0234] In one embodiment of the present disclosure, the one or more commands are individually or collectively executed by at least one processor (120), so that the electronic device (100) can move the region of interest into the image frame based on offset information of a region of the region of interest located outside the image frame, when a part of the region of interest is located outside the image frame according to the alignment result of the object region and the region of interest.
[0235] In one embodiment of the present disclosure, when the one or more commands are individually or collectively executed by at least one processor (120), the electronic device (100) can select at least two time intervals among the plurality of time intervals in which a search target object is detected based on interval and area information to generate a thumbnail video.
[0236] In one embodiment of the present disclosure, the one or more commands are individually or collectively executed by at least one processor (120), so that the electronic device (100) can arrange a plurality of image frames corresponding to each of at least two selected time intervals in an order determined by at least one of time, a size of an object area, an action and movement size of a search target object, a preset user preference, or search history information. The one or more commands are individually or collectively executed by at least one processor (120), so that the electronic device (100) can generate a thumbnail video using the plurality of image frames listed in an arrangement order.
[0237] The present disclosure provides a computer program product including a computer-readable storage medium. The storage medium may include instructions readable by an electronic device (100) so that the electronic device (100) performs the following operations: searching for at least one video matching a keyword extracted from a search query input by a user among a plurality of previously stored videos; recognizing a time period in which a search target object corresponding to the keyword is detected from the at least one searched video, and identifying a region in which the search target object is located within the recognized time period, thereby obtaining spatio-temporal information about the search target object; generating a thumbnail video by extracting a plurality of image frames from the at least one searched video based on the time and region information; and displaying the generated thumbnail video.
[0238] The program executed by the electronic device (100) described in the present disclosure may be implemented as hardware components, software components, and / or a combination of hardware components and software components. The program may be executed by any system capable of executing computer-readable instructions.
[0239] Software may include a computer program, code, instructions, or a combination of one or more of these, which may configure a processing device to do a desired thing or may independently or collectively command a processing device to do a desired thing.
[0240] Software may be implemented as a computer program containing instructions stored on a computer-readable storage medium. Examples of computer-readable storage media include magnetic storage media (e.g., read-only memory (ROM), random-access memory (RAM), floppy disks, hard disks, etc.) and optical readable media (e.g., CD-ROMs, DVDs (Digital Versatile Discs)). The computer-readable storage media may be distributed across network-connected computer systems, so that computer-readable code may be stored and executed in a distributed manner. The media may be readable by a computer, stored in a memory, and executed by a processor.
[0241] A computer-readable storage medium may be provided in the form of a non-transitory storage medium. Here, "non-transitory" simply means that the storage medium does not contain signals and is tangible, but does not distinguish between cases where data is stored semi-permanently or temporarily on the storage medium. For example, a "non-transitory storage medium" may include a buffer in which data is temporarily stored.
[0242] Additionally, programs according to the embodiments disclosed herein may be provided as part of a computer program product. The computer program product may be traded as a commodity between sellers and buyers.
[0243] A computer program product may include a software program, a computer-readable storage medium having the software program stored thereon. For example, the computer program product may be available from a manufacturer of an electronic device (100) or an electronic market (e.g., Samsung Galaxy Store). TM) may include a product in the form of a software program (e.g., a downloadable application) that is distributed electronically. For electronic distribution, at least a portion of the software program may be stored in a storage medium or temporarily created. In this case, the storage medium may be a storage medium of a server of a manufacturer of the electronic device (100), a server of an electronic market, or a relay server that temporarily stores the software program.
[0244] In a system comprising an electronic device (100) and / or a server, the computer program product may include a storage medium of the server or the storage medium of the electronic device (100). Alternatively, if there is a third device that is communicatively connected to the electronic device (100), the computer program product may include a storage medium of the third device. Alternatively, the computer program product may include a software program itself that is transmitted from the electronic device (100) to the third device, or from the third device to the electronic device.
[0245] In this case, either the electronic device (100) or the third device may execute the computer program product to perform the method according to the disclosed embodiments. Alternatively, at least one of the electronic device (100) and the third device may execute the computer program product to perform the method according to the disclosed embodiments in a distributed manner.
[0246] For example, the electronic device (100) may execute a computer program product stored in a memory (130, see FIG. 3) to control another electronic device that is in communication with the electronic device (100) to perform a method according to the disclosed embodiments.
[0247] As another example, a third device may execute a computer program product to control an electronic device in communication with the third device to perform a method according to the disclosed embodiment.
[0248] When a third device executes a computer program product, the third device may download the computer program product from the electronic device (100) and execute the downloaded computer program product. Alternatively, the third device may execute a computer program product provided in a pre-loaded state to perform the method according to the disclosed embodiments.
[0249] Although the embodiments described above have been described with limited examples and drawings, those skilled in the art will appreciate that various modifications and variations can be made based on the above description. For example, appropriate results can still be achieved even if the described techniques are performed in a different order than described, and / or components such as the described computer system or modules are combined or combined in a different manner than described, or are replaced or substituted with other components or equivalents.
Claims
1. In a method for an electronic device (100) to provide video search results, Step (S210) of receiving a search query from a user regarding a video to be searched; A step (S220) of searching for at least one video matching a keyword extracted from the search query among a plurality of videos stored in memory; A step (S230) of recognizing a time period in which a search target object corresponding to the keyword is detected from at least one searched video, and identifying a region in which the search target object is located within the recognized time period, thereby obtaining spatio-temporal information about the search target object; A step (S240) of generating a thumbnail video by extracting a plurality of image frames from at least one searched video based on the above section and area information; and Step of displaying the generated thumbnail video (S250); A method comprising:
2. In paragraph 1, A content identifier containing information about at least one of an object, action, behavior, situation, or event for the plurality of videos is pre-extracted and stored together with the plurality of videos, The step (S220) of searching for at least one video matching the above search query is: A step (S420) of converting keywords extracted from the search query to correspond to the type of content identifier for each of the plurality of videos; and A step (S430) of searching for at least one video having a content identifier that matches the converted keyword by comparing the converted keyword with the content identifier; A method comprising:
3. In any one of paragraphs 1 to 2, The step (S230) of acquiring the above section and area information is A step (S610) of performing object detection to identify a plurality of image frames in which the search target object is detected among the entire image frames constituting at least one of the searched videos; A step (S620) of recognizing the time section including the start time and the end time of the timeline of the identified plurality of image frames; A step (S630) of obtaining location information about an object area where the search target object exists from each of a plurality of image frames within the time interval; and A step (S640) of obtaining the section and area information based on the time section of the plurality of image frames and the location information regarding the object area; A method comprising:
4. In paragraph 3, The step (S230) of acquiring the above section and area information is A step of tracking the search target object whose position and size change through frame-to-frame tracking for the plurality of consecutive image frames within the time interval, thereby obtaining a spatio-temporal tube for the search target object; A method comprising:
5. In any one of paragraphs 1 to 4, The step of generating the above thumbnail video (S240) is: A step (S910) of obtaining an object area of the search target object from each of the plurality of image frames based on the section and area information; A step (S920) of transforming the object area based on the size information of the thumbnail image for preview; and Step (S930) of generating a thumbnail video using images of the above-mentioned converted object area; A method comprising:
6. In paragraph 5, The step of generating the above thumbnail video (S240) is: A method for generating a thumbnail video by selecting at least two time intervals among the plurality of time intervals in which the search target object is detected based on the above section and area information.
7. In paragraph 5, The step of generating the above thumbnail video (S240) is: A step (S1320) of arranging a plurality of image frames corresponding to each of the at least two selected time sections in an order determined by at least one of time, the size of the object area, the size of the action and movement of the object to be searched, preset user preference, or search history information; and A step (S1330) of generating the thumbnail video using the plurality of image frames listed in the arrangement order; A method comprising:
8. In an electronic device (100) that provides video search results, A user input unit (110) for receiving user input for entering a search query; At least one processor (120) comprising a processing circuit; A memory (130) storing one or more instructions; and display (140); Including, The electronic device (100) is configured such that the one or more instructions are individually or collectively executed by the at least one processor (120): Search for at least one video that matches the search query received through the user input unit (110) among the plurality of videos stored in the memory (130), Recognizing a time period in which a search target object corresponding to a keyword extracted from the search query is detected from at least one video searched for, and identifying a region in which the search target object is located within the recognized time period, thereby obtaining spatio-temporal information about the search target object. Based on the above section and area information, a thumbnail video is generated by extracting a plurality of image frames from at least one searched video, An electronic device (100) that displays the thumbnail video on the display (140).
9. In paragraph 8, The electronic device (100) is configured such that the one or more instructions are individually or collectively executed by the at least one processor (120): By performing object detection, a plurality of image frames in which the search target object is detected are identified among the entire image frames constituting at least one of the searched videos, and Recognize the time interval including the start and end points of the timeline of the identified plurality of image frames, Obtain location information about an object area where the search target object exists from each of a plurality of image frames within the above time interval, An electronic device (100) that obtains the section and area information based on the time section of the plurality of image frames and the position information regarding the object area.
10. In paragraph 9, The electronic device (100) is configured such that the one or more instructions are individually or collectively executed by the at least one processor (120): An electronic device (100) that tracks the search target object whose position and size change through frame-to-frame tracking for the plurality of consecutive image frames within the time interval, thereby obtaining a spatio-temporal tube for the search target object.
11. In any one of the clauses 8 to 10, The electronic device (100) is configured such that the one or more instructions are individually or collectively executed by the at least one processor (120): Obtaining an object area of the search target object from each of the plurality of image frames based on the above section and area information, Transform the object area based on the size information of the thumbnail image for preview, An electronic device (100) that generates the thumbnail video using images of the above-mentioned converted object area.
12. In paragraph 11, The electronic device (100) is configured such that the one or more instructions are individually or collectively executed by the at least one processor (120): Obtain the maximum values of the width and height of the above object area, By scaling the maximum values of the width and height of the object area using information about the preset width and height of the thumbnail image for preview, a region of interest is set, Align the center point of the above object area with the center point of the above area of interest, An electronic device (100) that generates the thumbnail video by cropping each of the plurality of image frames to the size of the aligned region of interest.
13. In paragraph 11, The electronic device (100) is configured such that the one or more instructions are individually or collectively executed by the at least one processor (120): An electronic device (100) that, when a plurality of time intervals in which the search target object is detected are acquired based on the above section and area information, selects at least two time intervals among the plurality of time intervals to generate the thumbnail video.
14. In paragraph 11, The electronic device (100) is configured such that the one or more instructions are individually or collectively executed by the at least one processor (120): Arrange a plurality of image frames corresponding to each of the at least two selected time intervals in an order determined by at least one of time, the size of the object area, the size of the action and movement of the object to be searched, preset user preferences, or search history information, An electronic device (100) that generates the thumbnail video using the plurality of image frames listed in the array order.
15. In a computer program product including a computer-readable storage medium, The above storage medium, An action of searching for at least one video that matches a keyword extracted from a search query input by a user among multiple previously stored videos; An operation of recognizing a time period in which a search target object corresponding to the keyword is detected from at least one searched video, and identifying a region in which the search target object is located within the recognized time period, thereby obtaining spatio-temporal information about the search target object; An operation of generating a thumbnail video by extracting a plurality of image frames from at least one searched video based on the above section and area information; and An action of displaying the generated thumbnail video; A computer program product comprising instructions readable by an electronic device (100) to cause the electronic device (100) to perform a task.
Citation Information
Patent Citations
Rectangular lumber block device
KR102424592B1
Mutual authentication method, appratus, and system thereof for secure and safe operation of urban air mobility based on wireless communication network
KR102636292B1
Spatio-temporal graphical user interface for querying videos
US20060256210A1
Temporal and spatial in-video marking, indexing, and searching
US20080046925A1
KR20210104979A