Scene analysis method and apparatus in a content streaming system

The method and apparatus in content streaming systems perform real-time scene analysis and ranking using AI models to address the lack of user-driven scene bookmarking and sharing, improving user engagement through efficient scene management.

JP2026510480APending Publication Date: 2026-04-07TVING CO LTD
View PDF 7 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2023-10-12
Publication Date
2026-04-07

AI Technical Summary

Technical Problem

Existing content streaming systems lack real-time scene analysis and ranking capabilities in response to user requests, particularly for bookmarking and sharing desired scenes.

Method used

A method and apparatus for performing real-time scene analysis in a content streaming system, utilizing artificial intelligence models to determine scene changes based on frame similarity and theme labels, and ranking scenes based on user interactions and popularity, allowing users to generate, store, and share bookmarks.

Benefits of technology

Enables real-time scene analysis and ranking in content streaming systems, allowing users to efficiently bookmark and share scenes based on their preferences and popularity, enhancing user engagement and content consumption.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026510480000001_ABST
    Figure 2026510480000001_ABST
Patent Text Reader

Abstract

This disclosure relates to a scene analysis method and apparatus in a content streaming system, the scene analysis method in the content streaming system may include the steps of: receiving a bookmark generation request from a user; analyzing a scene based on the user's bookmark generation request; and storing the analyzed scene in a scene library.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to a content streaming system, and more particularly, to a method and apparatus for analyzing scenes in a content streaming system.

Background Art

[0002] With the development of various technologies and changes in consumption trends, there have been significant changes in the way content is supplied and consumed. Due to the development of digital technology, computer technology, Internet / communication technology, etc., the boundaries between content types and production entities have become blurred, and significant changes have occurred in the content production and consumption patterns. Platforms have emerged that allow ordinary people to produce and distribute content. In addition, easy access to various content has been ensured, and the options for consumption methods have also diversified.

[0003] Among these various changes in the content industry is the OTT (Over-the-Top) service. The OTT service is a media platform based on the Internet and mobile communication. It goes beyond the framework of existing broadcast services and provides consumers with various content without the need for separate devices such as set-top boxes. The concept of OTT services started with providing movies and TV programs in the form of VOD (Video-on-Demand), but OTT services are still expanding, not only providing content created by OTT service providers but also expanding their scope to mobile platforms.

Summary of the Invention

Problems to be Solved by the Invention

[0004] An object of the present disclosure is to provide a method and apparatus for performing real-time scene analysis in a content streaming system in response to a bookmark request.

[0005] This disclosure aims to provide a method and apparatus for ranking scenes stored on a server in a content streaming system.

[0006] This disclosure aims to provide a method and apparatus for ranking scenes similar to scenes stored in a scene library within a content streaming system.

[0007] The technical problems addressed by this disclosure are not limited to those mentioned above, and other technical problems not mentioned herein will be clearly understood by a person with ordinary skill in the art to which this disclosure pertains from the following description. [Means for solving the problem]

[0008] According to one embodiment of the present disclosure, a scene analysis method in a content streaming system includes the steps of receiving a bookmark generation request from a user, analyzing a scene based on the user's bookmark generation request, and storing the analyzed scene in a scene library, wherein the step of analyzing the scene may include the step of determining the change time of the scene.

[0009] According to one embodiment of the present disclosure, the change point of the scene may be determined based on the similarity between frames.

[0010] According to one embodiment of the present disclosure, the similarity between frames may be determined further based on the similarity of pixels or groups of pixels between frames.

[0011] According to one embodiment of the present disclosure, the step of analyzing the scene may further include the step of analyzing the scene using a trained artificial intelligence (AI) model, the AI ​​model may be trained using at least one of images or scene change time information.

[0012] According to one embodiment of the present disclosure, the AI ​​model may be trained using at least one theme label corresponding to each frame.

[0013] According to one embodiment of the present disclosure, when a first frame having a first theme label changes to a second frame having a second theme label in order of playback time, the frame having the second theme label is determined to be the frame at the time of the scene change.

[0014] According to one embodiment of the present disclosure, the theme label may be assigned based on at least one of the following: audio data, a particular actor, a particular action, a particular original soundtrack (OST), a particular background music (BGM), or a particular location.

[0015] According to one embodiment of the present disclosure, the method may further include the steps of: analyzing a first theme label of a scene stored in a scene library; analyzing a second theme label of a scene stored in a server; assigning a ranking score to the scene stored in the server by comparing the first theme label and the second theme label; summing the ranking scores assigned to the scene stored in the server; and displaying the scene stored in the server in descending order of the summed ranking scores.

[0016] According to one embodiment of the present disclosure, the method further includes the steps of receiving a popular section request from a terminal and transmitting popular section information based on the popular section request to the terminal, wherein the bookmark generation request may be generated based on the popular section information, and the popular section information may be information relating to a popular section determined based on at least one of the cumulative number of scene views or the number of scene saves.

[0017] According to one embodiment of the present disclosure, the method may further include the steps of: determining at least one representative scene based on the number of times a scene is saved, the number of scene hits, scene feedback information, etc.; identifying a representative scene corresponding to a bookmark generation request when the scene is not analyzed based on the bookmark generation request; and storing the identified representative scene in the scene library.

[0018] According to one embodiment of the present disclosure, the step of determining the at least one representative scene may include determining the representative scene as the overlapping similar scene if the number of overlapping similar scenes among the similar scenes stored by the bookmark generation request is equal to or greater than a predetermined threshold.

[0019] According to one embodiment of the present disclosure, the step of storing the analyzed scene in the scene library may include the steps of transmitting information about the analyzed scene to a terminal, receiving a storage request from the terminal based on user input indicating a modification of at least one of the start and end points of the scene, and storing the scene in the scene library based on the storage request, wherein the storage request includes information indicating at least one of the modified start and end points of the analyzed scene.

[0020] According to one embodiment of the present disclosure, the method may further include the steps of receiving a request to access a scene library from a second user terminal that has received a share approval signal based on a first user terminal, and transmitting to the second user terminal, in response to the request to access the scene library, information relating to the scene library associated with the user of the first user terminal.

[0021] According to an embodiment of the present disclosure, a scene memory method in a content streaming system may include: sending a user's bookmark generation request to a server; receiving information about a scene analyzed based on the bookmark generation request from the server; and sending a request to store the scene in a scene library to the server based on the information about the scene.

[0022] According to an embodiment of the present disclosure, the information about the scene may include 1) Information relating to the start time of the scene and 2) the difference between the end time and start time of the scene, in relation to the start time of the scene. and, based on the information about the scene, the step of sending a request to store the scene in the scene library may include: receiving user input associated with modification of at least one of a start time point and an end time point of the scene based on the information about the scene; and sending a request to store in the scene library to the server based on the user input.

[0023] The features briefly summarized above about the present disclosure are exemplary aspects of the detailed description of the present disclosure to be described later, and do not limit the scope of the present invention.

Advantages of the Invention

[0024] According to the present disclosure, in a content streaming system, a scene stored in a scene library can be effectively analyzed. <00000�5>

[0025] According to the present disclosure, in a content streaming system, scene analysis can be performed in real time in response to a user's bookmark generation request.

[0026] According to the present disclosure, a scene desired by a user can be stored in a library.

[0027] According to the present disclosure, scene ranking in a content streaming system can be calculated.

[0028] The effects that can be achieved by the present disclosure are not limited to the effects specifically described above, and other advantages not mentioned in this specification will also be clearly understood by those skilled in the art from the following detailed description.

Brief Description of the Drawings

[0029] [Figure 1] It is a diagram showing a content streaming system according to an embodiment of the present disclosure.

[0030] [Figure 2] It is a diagram showing the structure of a client device according to an embodiment of the present disclosure.

[0031] [Figure 3] It is a diagram showing the structure of a server according to an embodiment of the present disclosure.

[0032] [Figure 4] It is a diagram showing the concept of a content streaming service according to an embodiment of the present disclosure.

[0033] [Figure 5] It is a diagram showing the procedure for receiving popularity interval information according to an embodiment of the present disclosure.

[0034] [Figure 6A] It is a diagram showing popularity interval information according to an embodiment of the present disclosure.

[0035] [Figure 6B] It is a diagram showing a modified screen of an analysis scene according to an embodiment of the present disclosure.

[0036] <​​​​​​​This figure shows a flowchart of a scene analysis method according to one embodiment of the present disclosure.

[0038] [Figure 9] This figure shows a flowchart of a scene analysis method according to one embodiment of the present disclosure.

[0039] [Figure 10] This figure shows an example of storage in a scene library according to one embodiment of the present disclosure.

[0040] [Figure 11] This figure shows an example of real-time scene analysis according to one embodiment of the present disclosure.

[0041] [Figure 12] This figure shows an example of storing already extracted scenes according to one embodiment of the present disclosure.

[0042] [Figure 13] This figure shows an example of extracting a representative scene according to one embodiment of the present disclosure.

[0043] [Figure 14] This figure shows a flowchart of a scene ranking method according to one embodiment of the present disclosure.

[0044] [Figure 15] This figure shows the procedure for sharing a scene library according to one embodiment of the present disclosure. [Modes for carrying out the invention]

[0045] Hereinafter, embodiments of the present disclosure will be described in detail with reference to the accompanying drawings so that those skilled in the art can easily implement the present disclosure. However, the present disclosure can be implemented in a variety of different forms and is not limited to the embodiments described herein.

[0046] In describing embodiments of this disclosure, detailed descriptions of known configurations or functions will be omitted if they would obscure the essence of this disclosure. In drawings, parts not relevant to the description of this disclosure will be omitted, and similar parts will be given the same reference numerals.

[0047] The functional blocks shown in the drawings and described below are merely possible embodiments. Other functional blocks may be used in other embodiments, without deviating from the ideas and scope of the detailed description. Furthermore, while one or more functional blocks of this disclosure are shown as individual blocks, one or more of the functional blocks of this disclosure may be a combination of various hardware and software configurations that perform the same function.

[0048] Furthermore, the expression "contains a certain component" is an "open" expression, merely indicating the existence of that component and should not be understood as excluding additional components. Consequently, when it is mentioned that a component is "connected" or "linked" to another component, it should be understood that it may be directly connected or linked to the other component, but other components may also exist in between.

[0049] Furthermore, a singular expression for a subject may be understood as a plural expression unless the context clearly indicates otherwise. In this disclosure, expressions such as "A or B" or "at least one of A and / or B" may be understood to include all possible combinations of the items listed together. Expressions such as "first," "second," and "third" may modify subjects regardless of order or importance and are used solely to distinguish one subject from other subjects of the same kind.

[0050] Furthermore, in this disclosure, "configured to..." may be understood, depending on the context, to be technically equivalent to any one of the following expressions from a hardware or software perspective: "suitable for...", "capable of...", "modified to...", "made to...", "capable of...", and "designed to...". These expressions may be interchangeable.

[0051] This disclosure pertains to a content streaming system, Analyze and remember the scene. This invention provides methods and apparatus. Specifically, This document describes a technology that stores scenes desired by users in a scene library, allows users to share these scenes with other users, and displays scenes that are suitable for the user's preferences based on the scenes stored in the scene library. Here, the term "video player" may mean a module that performs the function of playing video. Figure 1 is a diagram showing a content streaming system according to one embodiment of the present disclosure. Figure 1 shows a system for providing content-related services, such as content streaming and content-related information provision, and entities belonging to the system. Hereinafter in the present disclosure, various content-related services may be referred to as "content services" or other terms with equivalent technical meaning.

[0052] Referring to Figure 1, the content streaming system may include client devices 110 and servers 120. Here, client devices 110 are shown as a set of three client devices 110-1 to 110-3, but the content streaming system may include two or fewer client devices, or four or more client devices. Also, although one server 120 is shown, the content streaming system may include multiple servers that share various functions and interact with each other.

[0053] The client device 110 receives and displays content. After accessing the server 120 via a network, the client device 110 may receive content streamed from the server 120. In other words, the client device 110 is hardware on which client software or applications designed to use the content services provided by the server 120 are installed, and can interact with the server 120 through the installed software or applications. The client device 110 can be embodied as various types of devices. For example, the client device 110 may be one of the following: a portable device that is movable, a device that is movable but is usually fixed in place during use, or a device that is permanently installed in a specific location.

[0054] Specifically, the client device 110 can be embodied in at least one of the following forms: a smartphone 110-1, a desktop computer 110-2, a tablet PC, a laptop PC, a netbook computer, a workstation, a server, a personal data assistant (PDA), a portable multimedia player (PMP), a camera, or a wearable device. Here, the wearable device may be embodied in at least one of the following forms: accessory type (e.g., watch, ring, bracelet, necklace, glasses, contact lenses, head-mounted device (HMD), clothing type, body attachment type (e.g., skin pad or tattoo), or bio-implantable circuit). The client device 110 is a home appliance and may be embodied in at least one of the following forms: television 110-3, digital video disc (DVD) player, audio system, refrigerator, air conditioner, vacuum cleaner, oven, microwave oven, washing machine, or air purifier.

[0055] Server 120 performs various functions for providing content services. In other words, Server 120 can use various functions to provide content streaming and various content-related services to client devices 110. Specifically, Server 120 can digitize content for streaming and transmit it to client devices 110 over the network. To this end, Server 120 can perform at least one of the following: content encoding, data segmentation, transmission scheduling, or streaming transmission. Furthermore, for the convenience of content use, Server 120 can further perform at least one of the following functions: providing content guides, managing user accounts, analyzing user preferences, or recommending content based on preferences. Multiple functions from the various functions described above may be provided, and for this purpose, Server 120 may be embodied as multiple servers.

[0056] The client device 110 and the server 120 exchange information over the network, and content services may be provided to the client device 110 based on the exchanged information. In this case, the network may be a single network or a combination of various types of networks. The network may be understood as a configuration in which different types of networks are connected depending on the region. For example, the network may include at least one of a wireless network or a wired network. Specifically, the network may include a cellular network based on at least one of the following: 6th generation mobile communication system (6G), 5th generation mobile communication system (5G), Long Term Evolution (LTE), LTE Advance (LTE-A), Code Division Multiple Access (CDMA), Wideband CDMA (WCDMA), and Universal Mobile Telecommunications System (UMTS), Wireless Broadband (WiMAX), or Global System for Mobile Communications (GSM). Furthermore, the network may include a local area network based on at least one of the following: wireless local area network (WLAN), Bluetooth, Zigbee, near field communication (NFC), or ultra-wideband (UWB). In addition, the network may include wired networks such as the Internet or Ethernet.

[0057] Figure 2 shows the structure of a client device according to one embodiment of the present disclosure. Figure 2 shows the block structure of the client device (for example, the client device 110 in Figure 1).

[0058] Referring to Figure 2, the client device includes a display 202, an input unit 204, a communication unit 206, a sensing unit 208, an audio input / output unit 210, a camera module 212, a memory 214, a power supply unit 216, an external connection terminal 218, and a processor 220. However, depending on the type of device, at least one of the components shown in Figure 2 may be omitted.

[0059] The display 202 outputs visually recognizable information such as images and graphics. Therefore, the display 202 may include a panel and circuitry for controlling the panel. For example, the panel may include at least one of the following: a liquid crystal display (LCD), a light-emitting diode (LED), a light-emitting polymer display (LPD), an organic light-emitting diode (OLED), an active matrix organic light-emitting diode (AMOLED), or a flexible LED (FLED).

[0060] The input unit 204 receives input generated by the user. The input unit 204 may include various types of input sensing units. For example, the input unit 204 may include at least one of a physical button, a keypad, or a touchpad. Alternatively, the input unit 204 may include a touch panel. If the input unit 204 includes a touch panel, the input unit 204 and the display 202 may be realized as a single module.

[0061] The communication unit 206 provides an interface that enables client devices to form a network with other devices and send and receive data over the network. To this end, the communication unit 206 may include circuits for physically processing signals (e.g., encoders / decoders, modulators / demodulators, radio frequency (RF) front-ends, etc.) and a protocol stack for processing data according to a communication standard (e.g., a modem). According to various embodiments, the communication unit 206 may include multiple modules to support multiple different communication standards.

[0062] The sensing unit 208 collects sensing data, including data relating to the state of the client device or the surrounding environment. For example, the sensing unit 208 may measure physical values ​​or changes in values ​​related to the operating state or orientation of the client device and generate an electrical signal representing the measurement result. The sensing unit 208 may also measure physical values ​​or changes in values ​​of the surrounding environment of the client device and generate an electrical signal representing the measurement result. To this end, the sensing unit 208 may include at least one sensor and a circuit for controlling at least one sensor. Specifically, the sensing unit 208 may include at least one of the following: a gyro sensor, a magnetic sensor, an acceleration sensor, a grip sensor, a proximity sensor, a color sensor, a biosensor, a barometric pressure sensor, a temperature sensor, a humidity sensor, an illuminance sensor, or an ultraviolet (UV) sensor, an electronic nose (e-nose) sensor, a gesture sensor, an electromyography (EMG) sensor, an electroencephalogram (EEG) sensor, an electrocardiogram (ECG) sensor, an infrared (IR) sensor, an iris sensor, or a fingerprint sensor.

[0063] The audio input / output unit 210 outputs sound in accordance with electrical signals generated based on audio data and detects external sounds. In other words, the audio input / output unit 210 can convert sound and electrical signals to each other. For this purpose, the audio input / output unit 210 may include at least one of a speaker, a microphone, or a circuit to control them.

[0064] The camera module 212 collects data for generating images and videos. To this end, the camera module 212 may include at least one of a lens, a lens drive circuit, an image sensor, a flash, or an image processing circuit. The camera module 212 can focus light through the lens and generate data representing the color and brightness values ​​of the light using the image sensor.

[0065] Memory 214 may store the operating system, programs, applications, commands, configuration information, etc., necessary for operating the client device. Memory 214 may store data temporarily or permanently. Memory 214 may include volatile memory, non-volatile memory, or a combination of volatile and non-volatile memory.

[0066] The power supply unit 216 supplies the power necessary for the operation of the client device's components. For this purpose, the power supply unit 216 may include a converter circuit that converts the power into the amount of power required by each component. The power supply unit 216 may depend on an external power source or may include a battery. If it includes a battery, the power supply unit 216 may further include a charging circuit. The charging circuit may support wired or wireless charging.

[0067] The external connection terminal 218 is a physical connection unit for connecting the client device to other devices. For example, the external connection terminal 218 may include at least one of various standard terminals, such as a universal serial bus (USB) terminal, an audio terminal, a high-definition multimedia interface (HDMI) terminal, a recommended standard-232 (RS-232) terminal, an infrared terminal, an optical terminal, or a power terminal.

[0068] The processor 220 controls the overall operation of the client device. The processor 220 can control the operation of other components and perform various functions using those components. For example, the processor 220 can also request content data from the server via the communication unit 206 and receive content data. The processor 220 can also decode the received content data to restore the content. Furthermore, the processor 220 can output the content received from the server via the display 202 and the audio input / output unit 210. In addition, the processor 220 can control the state of content playback based on information input or sensed by at least one of the input unit 204, communication unit 206, sensing unit 208, audio input / output unit 210, camera module 212, power supply unit 216, and external connection terminal 218. For this purpose, the processor 220 may include at least one of at least one processor, at least one microprocessor, or at least one digital signal processor (DSP). In particular, the processor 220 can control other components and perform necessary operations so that the client device operates according to the various embodiments described below.

[0069] In the client device structure described with reference to Figure 2, all components are shown as being connected to the processor 220. Although not shown in Figure 2, at least some of the components may be connected via a bus. In this case, direct data exchange may occur between some components under the control of the processor 220.

[0070] Figure 3 shows the structure of a server according to one embodiment of the present disclosure. Figure 3 illustrates the block structure of the server (server 120 in Figure 1).

[0071] Referring to Figure 3, the server includes a communication unit 302, memory 304, storage 306, and a processor 308. However, according to various embodiments, at least one of the components shown in Figure 3 may be omitted.

[0072] The communication unit 302 provides an interface for communication between the server and other devices. To this end, the communication unit 302 may include circuits for generating and analyzing physical signals for communication. The interface provided by the communication unit 302 may support wired or wireless communication.

[0073] Memory 304 contains various information, Instructions and / or can store information and load computer programs, instructions, etc. stored in storage 306. Memory 304 can temporarily store data and instructions for the operation of the server and may include random access memory (RAM). Alternatively, memory 304 may include various storage media.

[0074] The storage 306 may non-temporarily store an operating system for running the server, programs for performing the server's functions, configuration information for the server's operation, and so on. For example, the storage 306 may include at least one of non-volatile memory such as read-only memory (ROM), eraseable programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), flash memory; hard disk; removable disk; solid state drive (SSD); or any form of computer-readable recording medium widely known in the art to which this disclosure belongs.

[0075] The processor 308 controls the overall operation of the server. The processor 308 can control the operation of other components and use those components to perform various functions. The processor 308 may include at least one of a central processing unit (CPU), a microprocessor unit (MPU), a microcontroller unit (MCU), or a processor of a form well known in the art to which this disclosure belongs. In particular, the processor 308 can control other components so that the server operates according to the various embodiments described below and performs the necessary operations.

[0076] This was explained with reference to Figure 3. server In this structure, it is illustrated that all components are connected to the processor 308. Although not shown in Figure 3, at least some of the components may be connected via a bus. In this case, direct data exchange may occur between some of the components under the control of the processor 308.

[0077] Figure 4 is a diagram illustrating the concept of a content streaming service according to one embodiment of the present disclosure. Figure 4 is a schematic diagram of some functions related to content streaming, and content streaming services according to various embodiments may have various other functions in addition to those shown in Figure 4.

[0078] Referring to Figure 4, control data and content data can be transmitted and received between client 410 and server 420. Specifically, control data can be transmitted from client 410 to server 420, control data can be transmitted from server 420 to client 410, and content data can be transmitted from server 420 to client 410.

[0079] Server 420 stores user information 422a, content information 422b, and content database (DB) 422c. User information 422a may include user account information, user service usage history information, and information about user preferences. Content information 422b may include a list of available content, content guide information, content metadata, and content consumption history information. Content DB 422c may contain content stored in data format. In addition, server 420 may store other information necessary to provide services.

[0080] The control data transmitted from client 410 to server 420 may include information about user login, information about user content selection, and information about user control of content. For this purpose, client 410 may generate and transmit control data from user input by user input processing operation 401. The control data from client 410 is processed by control / management operation 403 and used to provide content. For example, control / management operation 403 may select control data and / or content based on the control data from client 401. Furthermore, control / management operation 403 may analyze the user's consumption history and behavior to determine preferences, and recommend content according to the determined preferences.

[0081] The procedure for providing content to a user will be described below with reference to Figure 4. First, the client 410 generates control data containing login information (e.g., ID and password) entered by the user via the user input processing operation 401, and transmits the control data. The server 420 determines whether the user is valid by searching for the login information contained in the control data from the client 410 in the user information 422a, and determines the scope of content and services permitted according to the user's authority. However, if login is not required, or if limited services that can be provided without login are supported, the transmission and processing of login information may be omitted.

[0082] Next, the server 410 extracts content guide information from content information 422b by control / management operation 403 and sends control data containing the content guide information to the client 410. The client 410 outputs the content guide information contained in the control data to confirm the user's selection. The user's selection is sent to the server 410 as control data by user input processing operation 401. Information regarding the user's selection is processed by control / management operation 403 and used to select the content to be streamed. The server 420 searches for the selected content from the content DB 422, compresses and divides the retrieved content by encoding operation 407, and sends the content data. The content data may be pre-compressed and stored by encoding operation 407. Here, encoding operation 407 may include not only the operation of compressing the original content image, but also the operation of decoding the content data generated by compression and then recompressing it. In this case, compression may be performed based on the resolution, bitrate, and frames per second of the content image. If the data is already compressed and stored, the compression operation is omitted. profitServer 420 may perform segmentation on the content data. The content data may be restored by a decoding operation 409 and provided to the user by a playback operation 411. At this time, at least one of various video codecs or various audio codecs may be used for compression. For example, various video codecs include at least one of MPEG-2 (Moving Picture Experts Group-2), H.264 / Advanced Video Coding (AVC), H.265 / High Efficiency Video Coding (HEVC), H.266 / Versatile Video Coding (VVC), Video Processor 8 (VP8), Video Processor 9 (VP9), Alliance for Open Media (AOMedia) Video 1 (AV1), DivX, Xvid, VC-1, or Daala.

[0083] Audio codecs may include MP3 (MPEG 1 Audio Layer 3), AC3 (Dolby Digital AC-3), Enhanced AC-3 (E-AC3), AAC (Advanced Audio Coding, MPEG 2 Audio), Free Lossless Audio Codec (FLAC), High Efficiency Advanced Audio Coding (HE-AAC), OGG Vorbis, OPUS, and others.

[0084] By compressing content images according to various image resolutions, bitrates, and frames per second, multiple content data sets can be pre-generated. 410 This can measure throughput (or bandwidth) and determine the bitrate based on the measured throughput (or bandwidth).

[0085] Client 410 may receive information about multiple content data from server 420. The received information may include information representing the bitrate, resolution, frames per second, and location of the multiple content data.

[0086] Client 410 may determine at least one content data based on the bitrate, and may determine the playback content data corresponding to the playable resolution and frames per second, and its position, from among the at least one content data based on the capability information of Client 410. In this case, the capability information may include, but is not limited to, the client's maximum supported resolution and maximum supported frames per second.

[0087] Client 410 may send a content request to server 420 based on the position of the played content data. Server 420 may send content data corresponding to the content request to client 410 based on the received content request.

[0088] In another embodiment, the client 410 may receive user input relating to at least one of the image resolution or the number of frames per second, determine playback content data and its position in response to the user input, and send a content request to the server 420.

[0089] This disclosure is made at the user's request. Based onThis disclosure relates to real-time scene analysis technology. According to various embodiments of this disclosure, a user can generate bookmarks while viewing an image, share the generated bookmarks with other users, and use bookmarks generated by other users. Here, a bookmark may be data relating to a specific section (scene) within an image. Specifically, a bookmark may be data relating to the start and / or end points of a specific section within an image. A bookmark may also further include data indicating the image (e.g., data relating to an address (location), such as a URL). Here, the start and end points may be indicated using either a time value or a frame number, but this disclosure is not limited thereto. Also, according to various embodiments of this disclosure, a representative scene may mean a scene that has already been analyzed and stored in server 120. A representative scene may be selected based on the number of times a scene has been saved by multiple users, the number of hits, feedback information (likes, saves, shares, etc.), but this disclosure is not limited thereto. Furthermore, according to various embodiments of this disclosure, the server 120 can analyze scenes, store the analyzed scenes, and calculate a ranking of the scenes stored in the server 120. Hereinafter, this disclosure will describe various embodiments with reference to the drawings.

[0090] Figure 5 shows a procedure for receiving popular segment information according to one embodiment of the present disclosure. Server 512 in Figure 5 can be understood as a server of a content streaming system (for example, Server 120 in Figure 1). Similarly, the popular segment analysis server 513 can be understood as a server of a content streaming system (for example, Server 120 in Figure 1) or as a different server. Furthermore, terminal 511 in Figure 5 can be understood as a terminal of a content streaming system (for example, Terminal 110 in Figure 1).

[0091] Referring to Figure 5, in order to receive popular section information according to one embodiment of the present disclosure, terminal 511 may request popular section information. More specifically, in step S521, when the user requests a popular section for the image being viewed, terminal 511 may send popular section request information to server 512. Here, a popular section means a scene with a high number of views and / or saves. Referring to Figure 6, popular sections are expressed using a graph of the cumulative number of views for each scene, a graph of the number of saves for each scene, or a graph 601 that considers both the cumulative number of views and the number of saves for each scene. expressed However, this disclosure is not limited to these. Furthermore, in this disclosure, “scene” may be understood to mean the same thing as “section” or “image section.”

[0092] Next, in step S522, server 512 may transmit popular interval request information to popular interval analysis server 513. After popular interval analysis server 513 receives the popular interval request information, in step S523, popular interval analysis server 513 performs popular interval analysis Perform Popular section information may be transmitted to server 512. Here, popular section information may be information about representative scenes present in the image. Specifically, information about representative scenes may include the start and / or end times of the representative scenes. Here, the start and end times may be indicated using either a time value or a frame number. Next, in step S524, Popularity section analysis Server 512, having received popular section information from server 513, may transmit the popular section information to terminal 511. Next, in step S525, terminal 511, having received the popular section information from server 512, transmits the popular section information based on the received popular section information. display It is possible.

[0093] According to other embodiments of this disclosure, the server 512 and the popular interval analysis server 513 in Figure 5 may be the same server. In this case, the popular interval analysis server 513 may receive popular interval request information directly from the terminal 511. The popular interval analysis server 513 may also transmit popular interval information directly to the terminal 511.

[0094] Figure 6A shows popular section information according to one embodiment of the present disclosure. The terminal may, at the user's request (or automatically), provide information about popular sections as a graph of cumulative views, a graph of scene saves, or a graph 601 that considers both cumulative views and scene saves. By looking at the graph 601 thus provided, the user may select and view popular sections. Here, the graph of cumulative views, the graph of saves, or the graph 601 that considers both cumulative views and scene saves may provide information about sections viewed by many users using the server 120 and / or sections saved by many users. The graphs according to the present disclosure may, but are not limited to, be linked graphs and may be represented by other types of graphs (e.g., bar graphs, line graphs, etc.).

[0095] Figure 6B shows an analysis according to one embodiment of the present disclosure. scene This figure shows the editing screen. Referring to Figure 6B, the terminal may, at the user's request (or automatically), provide information about popular intervals as graph 601. When the user clicks on a specific part of graph 601 that they want to remember, the terminal may display the scene containing that part. Here, the scene containing that part may be a scene analyzed by the popular interval analysis server 513. Specifically, the terminal may display frame 611 (hereinafter, the first scene) that constitutes the scene. Here, the terminal may display frames before and after the first scene 611, and even frame 612 that includes those frames 611. This allows the user to modify the first scene 611. For example, the user may add and remember frames before and after the first scene 611, or delete and remember some frames. Here, the addition and deletion of frames may be done by modifying the start and / or end times of the first scene 611. The start and / or end times of the first scene 611 may be modified by user operation. For example, dragging the boundaries of a scene on the screen using a mouse cursor or touch input can modify the start and / or end points, and this disclosure is not limited to this.

[0096] Figure 7 shows a scene analysis procedure based on a bookmark generation request according to one embodiment of the present disclosure. Server 512 in Figure 7 can be understood as a server of a content streaming system (for example, Server 120 in Figure 1). Similarly, the popularity segment analysis server 513 can be understood as a server of a content streaming system (for example, Server 120 in Figure 1) or as another server.

[0097] Referring to Figure 7, in step S721, terminal 511 may send bookmark request information (e.g., current time information) to server 512. Referring to Figure 6A, bookmark request information may be generated when the user clicks the bookmark button 602 displayed on the terminal 511's display. The bookmark request information may include current time information and / or information indicating the position of the image the user is viewing.

[0098] Next, in step S722, server 512 may transmit bookmark request information to popularity segment analysis server 513. In step S723, the popularity segment analysis server 513, having received bookmark request information from server 512, may perform scene analysis. That is, the popularity segment analysis server 513 may perform scene analysis using the current time information contained in the bookmark request information. Here, scene analysis is performed to extract a scene using the bookmark request information and store that scene in the scene library. In other words, as an operation performed to store a scene desired by the user, scene analysis includes operations to extract a scene from the relevant image content based on the current time information in the bookmark request information (e.g., frame analysis, analysis of relationships between frames, etc.). According to one embodiment, scene analysis may be performed using the similarity between frames. Furthermore, according to other embodiments of this disclosure, scene analysis may be performed using a trained artificial intelligence (AI) model, but this disclosure is not limited thereto. A more detailed description of the scene analysis method will be given with reference to Figure 11.

[0099] In step S724, the popular interval analysis server 513, which has completed the scene analysis, recommends Bookmark The information may be transmitted to the server 512. In step S725, the server 512, having received the recommended bookmark information, may transmit the recommended bookmark information to the terminal 511. Here, the recommended bookmark information may include information about the intervals analyzed by the popular interval analysis server 513. Specifically, the recommended bookmark information may include information about the start and / or end times of certain intervals present in the image. For example, the recommended bookmark information may include information about 1) the start time of the interval and 2) the duration (the difference between the start and end times of the interval). Here, the start and end times can be indicated using either a time value or a frame number.

[0100] In step S726, terminal 511, having received the recommended bookmark information, may confirm and / or modify the recommended bookmark section. Specifically, terminal 511 may confirm the recommended bookmark information and then request server 512 to store the information. In this case, in step S728, server 512 may store the bookmark information in the personal library in response to terminal 511's storage request. At this time, image data related to the section may be stored. Alternatively, instead of image data related to the section, information about the start and / or end points of the image (i.e., bookmark information) may be stored in the personal library. In this case, since only information about the points in time is stored, the storage space of server 512 can be used efficiently.

[0101] According to other embodiments of the present disclosure, terminal 511 may, after confirming the recommended bookmark information, modify the section according to the recommended bookmark information. For example, terminal 511 may modify the section by changing the start and / or end points according to the recommended bookmark information. In this case, after modifying the section, terminal 511 may request server 512 to store the modified section. Upon receiving the storage request, server 512 may store the modified section in the personal library in response to the storage request (S728). At this time, image data relating to the section may be stored. Alternatively, instead of image data relating to the section, information regarding the start and / or end points of the image (i.e., bookmark information) may be stored in the personal library. In this case, since only information regarding the points in time is stored, the storage space of server 512 can be used efficiently.

[0102] According to yet another embodiment of the present disclosure, after performing scene analysis (S723), the popular section analysis server 513 can immediately store the analyzed scene in the personal library without sending recommended bookmark information (scene analysis bookmark information) to the server 512.

[0103] According to other embodiments of this disclosure, the server 512 and the popular section analysis server 513 in Figure 7 may be the same server. In this case, the popular section analysis server 513 may receive bookmark request information directly from the terminal 511. The popular section analysis server 513 may also send recommended bookmark information directly to the terminal 511. Figure 8 is a flowchart of a scene analysis method according to one embodiment of this disclosure. The scene analysis procedure in Figure 8 may be performed by a server of the content streaming system (for example, server 120 in Figure 1).

[0104] Referring to Figure 8, in step S801, the server 120 may receive a user request. Here, the user request may be a bookmark generation request. Specifically, if the user requests that a bookmark be generated at a specific point in time (for example, the current time) of the image currently being viewed, the server 120 may receive bookmark request information that includes information about that specific point in time.

[0105] In step S802, the server 120 may perform scene analysis. Here, scene analysis may be performed using the similarity between frames. Here, the similarity between frames may be determined based on the similarity of pixels present in each frame. In addition, according to other embodiments of the present disclosure, scene analysis may be performed using a trained artificial intelligence (AI) algorithm. For example, scene analysis may be performed by the Scale Invariant Feature Transform (SIFT) algorithm. Here, SIFT is an algorithm that extracts invariant feature points and descriptors according to the size and rotation of an image, and the similarity between frames may be determined by comparing the descriptors of the SIFT feature points. Alternatively, the Speed ​​Up Robust Features (SURF) algorithm may be used as an alternative to the SIFT algorithm. Or, the Oriented and Rotated Binary Robust Independent Elementary Features (ORB) algorithm may be used as an alternative to the SIFT algorithm. Alternatively, the FAST (features from accelerated segment test) algorithm can be used as an alternative to the SIFT algorithm. Or, a first algorithm for extracting feature points and a second algorithm for extracting descriptors can be used. Here, the first and second algorithms may be, but are not limited to, one of the algorithms described above. The first and second algorithms may be different.

[0106]

[0107] As another example, the server may extract feature points using a feature point extraction deep neural network (DNN), generate multidimensional (e.g., 128-dimensional) descriptors for the extracted feature points (using the algorithm described above), extract matching points between frames through comparison of the descriptors, and detect a scene change if the ratio of matching points is below a predetermined threshold. Here, the DNN may be a neural network that has a portion of a frame as input data and the probabilities of feature points as output data.

[0108] As a result of the scene analysis according to this disclosure, a scene composed of a set of consecutive frames within an image can be extracted. A more detailed scene analysis method will be explained with reference to Figure 11.

[0109] In step S803, the server 120 may transmit scene information extracted as a result of scene analysis. Here, the scene information may include information about a specific scene in the image analyzed by the server 120. Specifically, the scene information may include information about the start and / or end times of a specific scene present in the image. In this disclosure, “scene information” may be understood to mean the same thing as “recommended bookmark information” or “scene analysis bookmark information”.

[0110] In step S804, the server 120 may store the extracted scenes in the scene library. Specifically, the server 120 may store the scenes extracted in step S802 in the scene library. By storing the analyzed scenes in the scene library, the server 120 can collect information about the user's preferences and tastes. Therefore, the server 120 can recommend images that are suitable for the user's preferences. Here, the extracted scenes may be stored in the scene library regardless of the user's request for storage. That is, once the server 120 has completed the scene analysis, it may immediately store the analyzed scenes in the scene library. In this case, the user can later modify the analyzed scenes, and this disclosure is not limited to this.

[0111] Figure 9 is a flowchart showing a scene analysis method according to one embodiment of the present disclosure. The scene analysis procedure in Figure 9 may be performed by a server of a content streaming system (for example, server 120 in Figure 1).

[0112] Referring to Figure 9, in step S901, the server 120 may receive a user request. Here, the user request may be a bookmark generation request. Specifically, if the user requests that a bookmark be generated at a specific point in time in the image currently being viewed, the server 120 may receive bookmark request information that includes information about that specific point in time.

[0113] In step S902, the server 120 may perform scene analysis. Here, scene analysis may be performed using the similarity between frames. In addition, according to other embodiments of the present disclosure, scene analysis may be performed using a trained artificial intelligence (AI) model, but the present disclosure is not limited thereto. As a result of scene analysis according to the present disclosure, a scene composed of a set of some consecutive frames in the image may be extracted. A more detailed method of scene analysis will be explained with reference to Figure 11.

[0114] In step S903, the server 120 may transmit scene information extracted as a result of scene analysis. Here, the scene information may include information about a specific scene in the image analyzed by the server 120. Specifically, the scene information may include information about the start and / or end times of a specific scene present in the image. In this disclosure, “scene information” may be understood to mean the same thing as “recommended bookmark information” or “scene analysis bookmark information”.

[0115] In step S904, the server 120 may receive a request to store the extracted scene. Specifically, the server 120 may receive information from the user regarding whether or not to store the scene extracted in step S902. At this time, the user may request to store the extracted scene after making modifications to it. For example, the user may modify the start and / or end points of the extracted scene.

[0116] In step S905, the server 120 may store the extracted scenes in the scene library. Specifically, the server 120 may store the scenes extracted in step S902 in the scene library in response to a user's request to store them. By storing the extracted scenes in the scene library in response to a user's request to store them, the server 120 can collect information about the user's preferences and tastes. This allows the server 120 to recommend images that are suitable for the user's preferences.

[0117] Figure 10 shows an example of storage in a scene library according to one embodiment of the present disclosure. Specifically, a user can click a bookmark button 1010 displayed on the image player's display at a specific point in time they wish to remember while viewing an image. This allows the server 120 to analyze the scene using the information at the time the user clicked the bookmark button 1010 (hereinafter, the first point in time). The server 120 can store the scene extracted as a result of the image analysis in the scene library 1020 regardless of the user's storage request. According to another embodiment of the present disclosure, the server 120 can store the scene extracted as a result of the image analysis in the scene library 1020 in response to the user's storage request.

[0118] Scenes stored for each user can be categorized and displayed by theme (based on theme labels) or by title. Users can assign a score to each scene stored in their library. Assign This may result in scenes being displayed in order of score. However, this disclosure is not limited to this, and scenes may also be displayed in order of hits. Furthermore, when a user selects a theme, the scenes included in that theme may be played sequentially. Each scene may be stored on the device in an image file format (e.g., GIF) upon user request.

[0119] Scene Library 1020 may include not only images stored by each user ("My Library") but also "Real-Time Popular Images," etc. Real-Time Popular Images, etc. may be displayed in Scene Library 1020 based on rankings using user feedback. Scene Library 1020 may also include "Scenes for Each Theme." Images for each theme may be displayed in Scene Library 1020 according to rankings based on user preferences and tastes (for example, theme labels of images stored by the user in the library). Here, the preview scene displayed for "Real-Time Popular Images" and "Images for Each Theme" may be a single representative scene determined by considering various factors. When a user selects such a representative scene, scenes similar to that scene (i.e., scenes from different sections within the same image) may be displayed according to rankings, but this disclosure is not limited to this, and only the representative scene may be displayed.

[0120] Scene Library 1020 may include images from popular creators ("Popular Creator Images"). The scenes displayed in "Popular Creator Images" may be determined based on factors such as the total number of hits / feedbacks in the Scene Library and the level of interest in the creator (number of followers).

[0121] The preview images provided in Scene Library 1020 are from the start. and End time Inside It can be dynamically constructed using only I-frames.

[0122] When a scene stored in the scene library 1020 is played back, indicators showing the start and end points of the scene may be displayed, along with an indicator showing the section outside the scene. When the section outside the scene is selected, the user can view the video of that section. At this time, the viewing range may be automatically changed to the entire image range. In another embodiment, when the user views the relevant image, the scene portion stored in the user's own library may be separately displayed by an indicator.

[0123] Scenes in scene library 1020 can be searched using a search function. In this case, scenes can be searched based on the similarity between keywords and scene metadata, and the scene metadata can be predetermined based on theme labels assigned to the scenes. Figure 11 shows an example of real-time scene analysis according to one embodiment of the present disclosure. The scene analysis in Figure 11 can be performed by a server of the content streaming system (for example, server 120 in Figure 1).

[0124] Referring to Figure 11, the server 120 may receive bookmark request information (S1110). Here, bookmark request information may be generated when a user clicks a bookmark button. The bookmark request information may include information about a first time point. After receiving the bookmark request information, the server 120 may use the information in the bookmark request information to identify the frame 1101 of the first time point and the frames before and after the first time point (hereinafter, first frames) 1102 and 1103. Next, the server 120 may compare each frame 1101, 1102, and 1103. The comparison of each frame 1101, 1102, and 1103 may be performed based on the similarity between the frames. Here, the similarity between frames may be determined based on the pixels of each frame or based on the similarity of a group of pixels. Here, a pixel group may include multiple pixels; for example, a pixel group may be a K × K block (where K is an integer greater than 1).

[0125] If the similarity score, which is the result of comparing the frame at the first time point 1101 with the first frame, is greater than a predetermined threshold, then frames 1101, 1102, and 1103 can be determined to be similar to each other. In this case, the server 120 can identify the second frames 1104 and 1105 adjacent to the first frame. After identifying the second frames, the server 120 can compare frame 1101 at the first time point with the second frame. If the similarity score, which is the result of comparing the frame at the first time point 1101 with the second frame, is less than a predetermined threshold, then frames 1101, 1104, and 1105 can be determined to be dissimilar to each other. Here, the server 120 can determine that the scene has changed based on the time when the frames were determined to be dissimilar, and can store information indicating the scene change time points 1106 and 1107 (or the frame immediately before / after that time point) as bookmark information. According to this disclosure, the determination of each scene change point before and after the first point in time may be performed independently, but this disclosure is not limited to this. Furthermore, according to this disclosure, scene analysis may be performed in real time. In addition, server 120 uses a predetermined threshold according to this disclosure. any This disclosure may be limited to the above.

[0126] According to other embodiments of this disclosure, the server 120 may perform scene analysis using popularity intervals. Specifically, the server 120 may use popularity intervals 601 represented by a graph, Peak point Scene analysis can be performed based on this. Here, as an example, Peak point Scene analysis using this method may involve the use of specific mathematical formulas, but this disclosure is not limited to these.

[0127] According to other embodiments of the present disclosure, the server 120 has a predetermined time before and after the bookmark time. continuation time based on Scene analysis can be performed. That is, if scene analysis is difficult, server 120 will either... any Based on the unit time determined, the scene define This is possible. For example, if scene analysis based on the first point in time is difficult, server 120 can do the following: Scene R seconds before and after the first time point Defined as This is possible. Here, R represents a real number.

[0128] According to other embodiments of this disclosure, scene analysis may be performed using a trained artificial intelligence (AI) model. That is, the server 120 may detect scene change points using the AI ​​model. Here, the AI ​​model may be trained using images and information indicating scene change points within those images. In other words, specific images and information indicating scene change points within those specific images may be used as training data to train the AI ​​model. Using the first algorithm thus employed, the server 120 can perform accurate scene analysis. Here, scene Information indicating the point of change can be shown using either a time value or a frame number.

[0129] on the other handThe AI ​​model can be trained using theme labels assigned to each image. That is, the server 120 can recognize the scene theme in each frame of the image and detect the point of scene change. Here, the theme label may be a classification label relating to a specific actor, a specific action (fight, kiss, performance, etc.), a specific original soundtrack (OST), a specific background music (BGM), a specific location, etc., but this disclosure is not limited to these. According to this disclosure, the AI ​​model can be trained using theme labels assigned to each frame. Using the second algorithm used in this way, the server 120 can assign theme labels to each frame of the image and analyze the scene using the assigned theme labels. For example, if the frame at the first time point corresponds to a specific theme, the server 120 can assign a theme label to that frame. The server 120 can then determine the theme labels of the frames before and after the first time point. This allows the server 120 to identify the theme label for each frame, and if it determines that the theme label changes in a particular frame, it can determine that point in time as a scene change point and analyze the scene accordingly.

[0130] According to one embodiment of the present disclosure, the server 120 can perform scene analysis using audio data. For more accurate scene analysis, the server 120 can determine the scene change point using the audio data of the scene in question and perform scene analysis. Specifically, the server 120 can determine the point at which the audio data played at the first point in time is cut off or modified as the scene change point.

[0131] According to other embodiments of this disclosure, the server 120 may perform scene analysis using specific actors, specific actions (such as fighting, kissing, or playing music), specific OSTs, etc. For more accurate scene analysis, the server 120 may determine scene change points by having an AI model recognize when a specific actor appears or when a specific action or OST is played, and perform scene analysis accordingly.

[0132] According to one embodiment of this disclosure, the scene analysis by the first algorithm is interrupted due to an error. fail In this case, server 120 may use the second algorithm and change the order in which the algorithms are used. Here, an error is when server 120 is not at a scene change point but is at a scene change point. do This means when making a determination. do However, this disclosure is not limited thereto. Also, in Figure 11, scene analysis according to the similarity between frames is fail In this case, server 120 may use at least one of the first algorithm, the second algorithm, audio data, a specific actor, a specific action, and a specific OST, and the scene analysis method according to this disclosure is not limited thereto.

[0133] Figure 12 shows an embodiment of the present disclosure. First extraction done This is a diagram showing how scenes are memorized. Figure 12 destination The memory of the analyzed scenes can be stored by a server in the content streaming system (for example, server 120 in Figure 1).

[0134] Referring to Figure 12, the user can save a bookmark displayed on the image player's screen at a specific point in time they want to remember while viewing the image. button 1201 may be clicked. Here, server 120 may identify information 1202 regarding the time when the bookmark button was pressed (hereinafter referred to as the second time point). The second time point is destinationIf it is included in the representative scene 1203 analyzed by the server, the server 120 can store the representative scene in the scene library 1204 without performing scene analysis. Here, a representative scene is, destination This may refer to scenes that have been analyzed and stored on server 120. Representative scenes may be selected based on the number of times a scene has been saved by multiple users, the number of hits, feedback information (likes, saves, shares, etc.), etc., but this disclosure is not limited to these. For more detailed information on how to extract representative scenes, please refer to Figure 13 and see below.

[0135] Figure 13 shows an example of representative scene extraction according to one embodiment of the present disclosure. The representative scene extraction in Figure 13 may be performed by a server of the content streaming system (for example, server 120 in Figure 1).

[0136] Referring to Figure 13, the server 120 can identify a scene that a user remembers through a bookmark generation request using a single image. Here, a scene that a user remembers through a bookmark generation request may be called a similar scene. If similar scenes 1301 are concentrated at a particular point in time, the server 120 may select and extract a representative scene 1302 at that point in time. Here, in order to select the representative scene 1302, the server 120 may use the number of overlapping similar scenes. That is, the server 120 may select a scene as a representative scene if the number of overlapping similar scenes at a particular point in time is greater than or equal to a predetermined threshold.

[0137] Figure 14 is a flowchart illustrating a scene ranking according to one embodiment of the present disclosure. The scene ranking method in Figure 14 may be performed by a server of a content streaming system (for example, server 120 in Figure 1).

[0138] Referring to Figure 14, in step S1401, the server 120 may analyze the theme labels (hereinafter referred to as the first theme labels) of scenes stored in the scene library. Here, the scenes stored in the scene library may be scenes that were stored through scene analysis as a result of a user's bookmark generation request. Therefore, the server 120 may recommend scenes of a similar type to those stored in the scene library by the user and may rank those scenes.

[0139] In step S1402, the server 120 may analyze the theme label (hereinafter referred to as the second theme label) of the scene stored in the server 120. Here, the scene stored in the server 120 may be a scene stored by another user using the content streaming system according to this disclosure, but this disclosure is not limited to this.

[0140] In step S1403, the server 120 may assign a ranking score. Specifically, the server 120 may identify a first theme label and a second theme label. Next, the server 120 may compare the first theme label and the second theme label analyzed in steps S1401 and S1402. If the first theme label and the second theme label are the same, the server 120 may assign a ranking score to the scene. A single scene may have multiple theme labels, and the more theme labels that overlap with scenes stored in the user's scene library, the higher the ranking score that scene may be assigned. For example, The scene being compared and the scene stored in the user's scene library If the same actor appears in the same scene, server 120 may assign a ranking score of 1 to that scene. In If the same background music is used, server 120 may assign a ranking score of 1 to the scene. Here, Granted Ranking score It is fine if it differs depending on the theme.However, this disclosure is not limited to this. Furthermore, ranking scores may be assigned based on whether theme labels are similar rather than whether they are identical. In this case, ranking scores may be assigned based on weights corresponding to the similarity of theme labels.

[0141] In step S1404, the server 120 may aggregate the ranking scores. Specifically, the ranking scores assigned to a particular scene in step S1403 may be aggregated. For example, if a particular scene stored in the server 120 and a scene stored in the user's scene library have three identical theme labels, the server 120 may derive a ranking score of 3 by aggregating the ranking scores assigned to the particular scene. Here, the ranking scores may differ for each theme label, and this disclosure is not limited thereto.

[0142] In step S1405, the server 120 may calculate a scene ranking. Specifically, it may calculate a scene ranking using the ranking scores totaled in step S1404. The server 120 may calculate a scene ranking by comparing the ranking scores assigned to each scene. A higher ranking score can result in a higher ranking for that scene. Based on this, the server 120 may use the ranking calculation results to display the scenes in the user's scene library in ranking order. Through this ranking calculation, the server 120 can calculate a scene ranking that suits each user's preferences.

[0143] In other embodiments of this disclosure, the server 120 may calculate rankings using user feedback in real time. Here, user feedback means the number of scenes stored, the number of likes, the number of saves, the number of hits, the number of shared scenes, etc. Specifically, the server 120 may use user feedback to calculate a ranking of all scenes stored in the server 120 and display the scenes in the scene library in the order of the ranking thus calculated. For example, each piece of feedback may be weighted, and a ranking score for the scenes may be calculated by a weighted sum, and a ranking may be calculated based on the calculated ranking score.

[0144] In other embodiments of the present disclosure, the server 120 may calculate rankings according to each theme of a scene. Specifically, the server 120 may classify themes using an AI model and calculate a ranking for each theme. The ranking may be calculated by assigning ranking scores based on theme labels and / or by using user feedback, and the present disclosure is not limited thereto.

[0145] Figure 15 shows a procedure for sharing a scene library according to one embodiment of the present disclosure. Server 512 in Figure 15 can be understood as a server of a content streaming system (e.g., Server 120 in Figure 1). The first user terminal and the second user terminal can be understood as client devices of a content streaming system (e.g., Client device 110 in Figure 1).

[0146] Referring to Figure 15, in step S1521, the first user terminal 1511 may send a share approval signal to the second user terminal 1512. Here, the share approval signal may mean a signal that includes an intention to approve the second user terminal 1512 viewing the scenes stored in the first user terminal 1511's scene library. As a result, the second user terminal 1512, having received the share approval signal, may request the server 512 to access the first user terminal's scene library (S1522). In response, the server 512, having received the scene library access request, may provide the scene library to the second user terminal 1512 (S1523). Here, the scene library provided by the server 512 may be the scene library that exists in the first user terminal 1511. The second user terminal 1512, having received the scene library, may display the scene library so that the user can view it (S1524). Through this process, users can share scene libraries with each other. Here, a server user authentication procedure may be performed, and only authenticated users may be able to share the scene library, but this disclosure is not limited to this.

[0147] According to other embodiments of the present disclosure, the first user terminal 1511 may transmit a share approval signal to the server 512, and the server 512 may transmit a share approval signal to the second user terminal 1512.

[0148] The exemplary methods in this disclosure are represented as a series of actions for clarity of explanation, but this is not intended to limit the order in which the steps are performed, and each step may be performed simultaneously or in a different order, as is necessary. To implement the methods according to this disclosure, the illustrated steps may include further steps, or the remaining steps may include some of the steps, or even further steps may include some of the steps.

[0149] The various embodiments of this disclosure are not intended to enumerate all possible combinations, but rather to illustrate representative aspects of this disclosure. The matters described in the various embodiments may be applied independently or in combination of two or more.

[0150] Furthermore, various embodiments of this disclosure may be implemented by hardware, firmware, software, or a combination thereof. In the case of hardware implementation, embodiments may be implemented by one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), general-purpose processors, controllers, microcontrollers, microprocessors, and the like.

[0151] The scope of this disclosure includes software or machine-executable commands (e.g., operating systems, applications, firmware, programs, etc.) that cause an apparatus or computer to perform operations according to the methods of various embodiments, and non-temporary computer-readable media on which such software or commands are stored and executed on the apparatus or computer.

Claims

1. A scene analysis method in a content streaming system, The steps include receiving a user's request to generate a bookmark, The steps include: analyzing the scene based on the user's bookmark generation request; The steps include storing the analyzed scene in the scene library, Methods that include...

2. The step of analyzing the aforementioned scene includes the step of determining the point in time when the scene changes. The method according to claim 1, wherein the change point of the scene is determined based on the similarity between frames.

3. The method according to claim 2, wherein the similarity between the frames is determined further based on the similarity of pixels or groups of pixels between the frames.

4. The step of analyzing the scene further includes the step of analyzing the scene using a trained artificial intelligence (AI) model. The method according to claim 1, wherein the AI ​​model is trained using at least one of images or scene change time information.

5. The method according to claim 4, wherein the artificial intelligence (AI) model is trained using at least one theme label corresponding to each frame.

6. The method according to claim 4, wherein when a first frame having a first theme label changes to a second frame having a second theme label in order of playback time, the frame having the second theme label is determined to be the frame at the time of the scene change.

7. The method according to claim 5, wherein the theme label is assigned based on at least one of audio data, a specific actor, a specific action, a specific original soundtrack (OST), specific background music (BGM), or a specific location.

8. The steps include: analyzing the first theme label of the scene stored in the scene library, The steps include: analyzing the second theme label of the scene stored on the server, The steps include: comparing the first theme label and the second theme label to assign a ranking score to the scenes stored in the server; The steps include: summing up the ranking scores assigned to the scenes stored in the server; The steps include: displaying the scenes stored on the server in descending order of the combined ranking scores; The method according to claim 5, further comprising:

9. The steps include receiving a request for a popular section from the terminal, The steps include: transmitting popular section information based on the popular section request to the terminal; It further includes, The aforementioned bookmark generation request is generated based on the aforementioned popular section information, The method according to claim 1, wherein the popular section information is information about popular sections determined based on at least one of the cumulative number of scene views or the number of scene saves.

10. A step of determining at least one representative scene based on the number of scene saves, the number of scene hits, and scene feedback information, If the scene is not analyzed based on the bookmark generation request, the steps include identifying a representative scene corresponding to the bookmark generation request, The steps include storing the identified representative scene in the scene library, The method according to claim 1, further comprising:

11. The step of determining at least one representative scene is: The method according to claim 10, further comprising the step of determining the representative scene as the overlapping similar scene if the number of overlapping similar scenes among the similar scenes stored by a bookmark generation request is equal to or greater than a predetermined threshold.

12. The step of storing the analyzed scene in the scene library is: The steps include transmitting the information regarding the analyzed scene to a terminal, The steps include receiving a storage request from the terminal based on user input indicating a modification of at least one of the start and end points of the aforementioned scene, The step includes storing the scene based on the storage request in the scene library, The method according to claim 1, wherein the memory request includes information indicating at least one of the modified start and end points of the analyzed scene.

13. The steps include receiving a request to access the scene library from a second user terminal that has received a share approval signal based on the first user terminal, In response to a request to access the scene library, the second user terminal receives information about the scene library associated with the user of the first user terminal. The method according to claim 1, further comprising:

14. A method for storing scenes in a content streaming system, The steps include sending a user's bookmark generation request to the server, The steps include receiving information about the scene analyzed based on the aforementioned bookmark generation request from the server, The steps include sending a request to the server to store the scene in the scene library based on the information relating to the scene, Methods that include...

15. The information relating to the scene includes, in relation to the scene change point, 1) the start time of the scene, and 2) information regarding the difference between the end time and the start time of the scene. The step of sending a request to store the scene in the scene library based on the information relating to the scene is: The steps include receiving user input associated with modifying at least one of the start and end points of the scene, based on the information relating to the scene; The method according to claim 14, comprising the step of sending a storage request to the scene library based on the user input to the server.

Citation Information

Patent Citations

  • streaming video bookmark

    JP2004526372A

  • Video Scene Detection

    JP2015536094A

  • System and method for generating media bookmarks

    US20100042642A1

  • Bookmarking segments of content

    US20120209841A1

  • Systems and Methods for Generating a Summary Storyboard from a Plurality of Image Frames

    US20190130192A1