Interactive video playback system based on user operation
Patent Information
- Application Number
- KR1020250106833
- Authority / Receiving Office
- KR · KR
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2025-08-04
- Publication Date
- 2026-08-11
- Estimated Expiration
- 2045-08-04
Smart Images

Figure 112025088406658-PAT00001_ABST
Abstract
Description
Technology Field
[0001] The present invention relates to an interactive video playback system based on user operation, and more specifically, to an interactive video playback system based on user operation that can dynamically control a user-customized video playback flow based on user touch or gesture input without the installation of a separate application. Background Technology
[0002] User input-based interactive video content control technology is a technology that controls video playback by changing it according to user inputs such as touches or gestures. Through this, users can go beyond simply watching videos and directly select and manipulate the flow or scenes of the video via input.
[0003] However, existing user input-based interactive video content control technologies have been limited in that many systems require the installation of separate applications or account logins, which hinders user accessibility and makes immediate content experience difficult; input methods are restricted or gesture recognition precision is low, resulting in cases where various user input intentions are not accurately reflected; responses to input are limited to simple playback and pause, lacking content branching or customized progression, making it difficult to provide an immersive user experience; and session management within the system is unstable or relies excessively on cookies or storage, which limits the ability to continuously track or utilize user input records or behavioral data.
[0004] Therefore, there is an urgent need to develop technology that can resolve the aforementioned existing problems and control the playback method of dynamic video based on user inputs such as touch or gestures. The problem to be solved
[0005] One aspect of the present invention provides an interactive video playback system based on user operation, which allows a user to immediately access content in a web browser through QR code scanning or NFC tagging without installing a separate application or logging in.
[0006] In addition, it provides an interactive video playback system based on user operations that interact with video content through the recognition of various forms of gesture input (Tap, Hold, Drag, Rotate, Pinch, Spread).
[0007] In addition, an interactive video playback system based on user operation is provided, configured so that content branches according to the user's input method and range.
[0008] In addition, it provides an interactive video playback system based on user operation that adjusts the video playback range on a frame-by-frame basis according to various physical parameters such as the position, speed, and direction of gesture input.
[0009] In addition, the system provides an interactive video playback system based on user operations that creates a unique session based on IP address and browser information and temporarily stores user input data without cookies.
[0010] In addition, the system provides an interactive video playback system based on user operation, which automatically executes subsequent responses such as rewind, action clip playback, and pause when user gesture input is interrupted.
[0011] In addition, an interactive video playback system based on user operation is provided, which enables automatic repeat playback or branching playback to a subsequent video after all video content has been played.
[0012] In addition, it provides an interactive video playback system based on user operations that collects and analyzes user behavior data such as input location, success rate, and dropout rate. means of solving the problem
[0013] The present invention provides a user operation-based interactive video playback system comprising: an input unit that receives data from the outside into the system according to the concept of the present invention; an output unit that outputs result data to the outside; a communication unit that transmits and receives data between the inside and outside of the system; a storage unit that stores data generated by the system; a control unit that controls the operation of the system; and a memory unit that stores data resulting from the execution of various programs, wherein the memory unit comprises: a content entry unit that allows a user to enter content through a user terminal; a user identification unit that identifies individual activities of the user; an initial screen output unit that provides an initial screen of the content to the user; a gesture recognition unit that recognizes the user's gesture; a playback video control unit that plays the video of the content according to the user's gesture; and a user interface unit that collects interaction data between the user and the system and provides it to the user. Effects of the invention
[0014] According to one aspect of the present invention, accessibility is greatly improved because the user can immediately access content in a web browser by scanning a QR code or NFC tagging without installing a separate application or logging in.
[0015] In addition, it can recognize various forms of gesture input (Tap, Hold, Drag, Rotate, Pinch, Spread), allowing users to interact with video content in an intuitive and natural way.
[0016] In addition, since the content can be configured to branch based on the user's input method and range, personalized content delivery is possible.
[0017] In addition, the video playback range can be adjusted on a frame-by-frame basis according to various physical parameters such as the position, speed, and direction of the gesture input, enabling detailed and precise content response.
[0018] In addition, the system creates a unique session based on IP address and browser information and temporarily stores user input data without cookies, thereby simultaneously satisfying privacy protection and lightweight processing performance.
[0019] In addition, if the user's gesture input is interrupted, the system automatically executes subsequent responses such as rewinding, playing action clips, or pausing to maintain a smooth content flow.
[0020] In addition, after the video content has finished playing, automatic repeat playback or branching playback to a subsequent video is possible, allowing the continuity and immersion of the content to be maintained.
[0021] In addition, since user behavior data such as input location, success rate, and bounce rate can be collected and analyzed, it is useful for future UI improvements and content organization optimization. Brief explanation of the drawing
[0022] FIG. 1 is a block diagram of an interactive video playback system based on user operation according to an embodiment of the present invention. FIG. 2 is a diagram showing a user generating an initial screen by scanning an NFC tag or QR code using a personal mobile terminal according to an embodiment of the present invention. FIG. 3 is a diagram showing the process of a user scanning a QR code using a personal mobile terminal to generate an initial screen according to an embodiment of the present invention. FIG. 4 is a diagram showing six user gestures for controlling an image according to an embodiment of the present invention. FIG. 5 is a diagram showing that a video is played according to conditions set according to six gestures of a user according to one embodiment of the present invention. FIG. 6 is a diagram showing a method in which a video is played differently when a gesture input is interrupted according to an embodiment of the present invention. FIG. 7 is a diagram showing a method of playback in which an image according to one embodiment of the present invention is played after all the entire frames set have been played. Specific details for implementing the invention
[0023] The embodiments described in this specification and the configurations illustrated in the drawings are merely preferred examples of the disclosed invention, and various modifications that may replace the embodiments and drawings of this specification may exist at the time of filing this application.
[0024] Additionally, the same reference numerals or symbols presented in each drawing of this specification represent parts or components that perform substantially the same function.
[0025] Furthermore, the terms used in this specification are for describing embodiments and are not intended to limit or / or restrict the disclosed invention. Singular expressions include plural expressions unless the context clearly indicates otherwise. In this specification, terms such as "comprising" or "having" are intended to indicate the existence of the features, numbers, steps, actions, components, parts, or combinations thereof described in the specification, and do not preclude the existence or addition of one or more other features, numbers, steps, actions, components, parts, or combinations thereof.
[0026] Additionally, terms including ordinal numbers, such as “first,” “second,” etc., as used herein may be used to describe various components, but said components are not limited by said terms, and said terms are used solely for the purpose of distinguishing one component from another. For example, without departing from the scope of the present invention, the first component may be named the second component, and similarly, the second component may be named the first component. The term “and / or” includes a combination of a plurality of related described items or any one of a plurality of related described items.
[0027] In addition, terms such as "~part," "~unit," "~block," "~part," and "~module" may refer to a unit that processes at least one function or operation. For example, the above terms may refer to at least one piece of hardware such as an FPGA (field-programmable gate array) or ASIC (application specific integrated circuit), at least one piece of software stored in memory, or at least one process processed by a processor.
[0028] Hereinafter, embodiments according to the present invention will be described in detail with reference to the attached drawings. However, the following drawings attached to this specification are intended to illustrate preferred embodiments of the present invention and serve to further enhance understanding of the technical concept of the present invention together with the aforementioned description; therefore, the present invention should not be interpreted as being limited only to the matters described in such drawings.
[0029] FIG. 1 is a block diagram of an interactive video playback system based on user operation according to an embodiment of the present invention.
[0030] Referring to FIG. 1, an interactive video playback system (100, hereinafter referred to as the "system") based on user operation is composed of an input unit (110), an output unit (120), a communication unit (130), a storage unit (140), a control unit (150), and a memory unit (160). The input unit (110) collects operation signals such as touch, click, and gesture from the user and transmits them to the system, and the output unit (120) visually outputs video content and a user interface to the screen. The communication unit (130) is connected to an external server via an NFC tag or QR code to call a content URL and transmit and receive data, and the storage unit (140) stores user settings, content history, etc. The control unit (150) integrally controls the operation of all components and adjusts the content playback flow and interface output according to user input. The memory unit (160) includes a content entry unit (170), a user identification unit (180), an initial screen output unit (190), a gesture recognition unit (200), a playback video control unit (210), and a user interface unit (220), thereby realizing a user-customized video playback experience such as content loading, user recognition, gesture interpretation, and interaction UI control.
[0031] More specifically, the content entry section (170) is a web interface entry section designed to allow the user to directly access browser-based interactive content by scanning a QR code or NFC (Near Field Communication) tagging. The content can be accessed immediately without the need for a separate application installation or login procedure, and the content is automatically loaded through an HTML5-based web player, maximizing the convenience and accessibility of the user experience. The user is connected to the content page via a URL (Uniform Resource Locator). For example, if the user recognizes a QR code or NFC tag attached to a specific product, exhibit, or poster, or accesses a URL link provided through various media such as a webpage, text message, email, or social media, the URL is called and connected to the content page. At this time, an initial screen is displayed immediately, contributing to the formation of immersion before the user begins interacting with the content.
[0032] The user identification unit (180) is a component that creates a unique session based on the user's terminal IP address and browser information (User-Agent) to identify and manage individual user activities. The user identification unit (180) is maintained on a browser-by-browser basis and is automatically deleted when the page is refreshed or closed. In addition, the number of touch inputs and gesture logs can be preserved for a certain period of time without using cookies or local storage (client-side data storage provided by the web browser), satisfying both requirements of privacy protection and lightweight content operation. Input data collected per session is subsequently used for response processing and user behavior analysis.
[0033] The initial screen output unit (190) displays a screen that is exposed to the user before the main video is played, thereby serving to induce prior awareness of the flow of the video content and immersion. On the initial screen, the user can naturally grasp the atmosphere of the content through subtle movements of the character (e.g., blinking eyes, breathing, etc.) or still cuts (screens extracted in the form of still images of specific scenes from the video). Additionally, the initial screen output unit (190) functions as a preparation section prior to user gesture input, and the content creator configures the initial scene so that the visual transitions of the video continue smoothly. Consequently, the initial screen functions as a core interface that attracts the user's attention and creates a sense of immersion.
[0034] The gesture recognition unit (200) detects actions input by a user on a touchscreen in real time (such as input coordinates, direction of movement, speed of movement, and duration of the user screen touch), classifies them into predefined gesture types (Tap, Hold, Drag, Rotate, Pinch, Spread), and converts them into system commands. It precisely recognizes gestures by analyzing physical characteristics such as input coordinates, direction of movement, speed, and duration. Additionally, each gesture is connected to a pre-set trigger point to induce playback of a specific video clip (Clip, a single independent video unit or a short, playable video segment) or frame segment. Through this, completely different subsequent reactions can be induced depending on the input method even within the same video content, thereby expanding the range of video interaction to suit the user.
[0035] The playback video control unit (210) precisely maps the user's gestures to the playback frames of the video content to control the playback section of the video according to the input in real time. The content creator can pre-define playback start frames, end frames, frame branching conditions, etc., for each gesture type. In addition, for complex gestures such as Drag, Rotate, Pinch, and Spread, it enables dynamic playback range control based on movement distance, movement speed, or movement direction, and enables the implementation of various video story branches depending on the input method even at the same screen location.
[0036] Additionally, the playback video control unit (210) includes a function to automatically perform a subsequent processing operation by determining the current state of the content when the user stops gesture input midway. The subsequent processing operation includes rewind (reverse playback), automatic playback of action clips, and a pause state switching operation.
[0037] Additionally, the playback video control unit (210) determines the action to be performed after all frames of the video have been played, and includes functions such as rewind (reverse playback), autoplay (automatic video switching), and scenario branching. Through this, subsequent video can be automatically played or the screen can naturally move to other branched video content according to conditions set by the video creator.
[0038] Additionally, the user interface unit (220) performs the role of collecting and analyzing interaction data between the user and video content in real time. It collects and analyzes indicators including the input location of user gestures, gesture success rate, and video content branching drop rate. Furthermore, it utilizes the collected user behavior data for content video renewal, UI improvement, and analysis of repeated viewing patterns, and performs a user-customized video content recommendation function in the long term.
[0039] FIG. 2 is a drawing showing a user generating an initial screen by scanning an NFC tag or QR code using a personal mobile terminal according to an embodiment of the present invention, and FIG. 3 is a drawing showing a process of a user generating an initial screen by scanning a QR code using a personal mobile terminal according to an embodiment of the present invention.
[0040] Referring to FIGS. 2 and 3, the system provides a web player-based interface designed to allow users to immediately access interactive content in a web browser environment without being required to install a separate application or log in to an account. The access method is very intuitive, and the content URL is automatically called when the user scans a QR code attached to a product package, display, poster, etc., or tags an NFC tag embedded in a specific object with a smartphone. The link is opened through a browser, and the system is designed so that the content is automatically loaded without any separate confirmation process for the user.
[0041] The content is executed on an HTML5-based web player (the 5th major version of HyperText Markup Language), and upon connection, a unique session is automatically created based on the user's IP address and User-Agent information (an HTTP header that transmits environment information as a string to the server when a user accesses a website). This session is maintained on a browser session basis; if the user refreshes the page or closes the browser, the session is immediately terminated, and the number of user touches and gesture input data collected during that session are also reset. Since data is not stored separately via cookies (small amounts of data that a website stores in the user's browser) or local storage, it satisfies both objectives of privacy protection and lightweight content execution.
[0042] The screen that is first displayed upon initial entry is called the 'initial screen' and serves to provide visual immersion to the user before the main video begins. The initial screen may consist of (1) a form in which simple repetitive actions such as a character breathing or blinking are automatically played (e.g., upon initial entry, an initial action such as a character blinking may be automatically played once without user input), and (2) a still image immediately before the main video begins. This initial composition is designed to allow the user to grasp the flow and atmosphere of the content in advance before inputting on the screen, and to feel a natural sense of connection during actual interaction.
[0043] FIG. 4 is a diagram showing six user gestures for controlling an image according to an embodiment of the present invention.
[0044] Referring to FIG. 4, the system is sophisticatedly configured to recognize various types of touch gestures input by a user on the screen and to differentially control video content according to the input type. Through presets, the content creator designates a specific point or area within the content as a touch active point and sets a video segment linked to the gesture type corresponding to it. This provides a structure in which multiple content flows can unfold according to various input methods even within a single piece of content.
[0045] Supported gestures are classified into a total of six types: Tap, Drag, Hold, Rotate, Pinch, and Spread. The Tap gesture is the most basic input method, involving a single touch on a specific point, and acts as a trigger to play the entire video or a single clip (an independent video) following the set point. The Hold gesture plays content only while the user keeps their hand on the screen; playback stops or switches to a separate action the moment the input is interrupted. The Drag gesture detects the direction and speed of the user's finger movement relative to the start and end points, adjusting the content playback duration in real-time. The Rotate gesture changes the flow of the video depending on the direction the user rotates within a set circular area, either clockwise or counterclockwise. The Pinch gesture involves pressing two points simultaneously to bring them inward, zooming out the screen. The Spread gesture involves pressing two points simultaneously to spread them outward, zooming in the screen.
[0046] Meanwhile, a rub gesture may be additionally included in the six gestures. When the system recognizes the user's rub gesture, it can precisely control the playback method of video content based on the rub direction, speed, and distance traveled. In other words, the system recognizes the user's rub gestures in the left-right, up-down, or diagonal directions and precisely controls the playback method of video content based on the corresponding speed and distance traveled.
[0047] FIG. 5 is a diagram showing that a video is played according to conditions set according to six gestures of a user according to one embodiment of the present invention.
[0048] Referring to FIG. 5, each gesture input induces the playback of a specific video clip or frame range within the content according to input points and a frame mapping structure (a method of aligning or connecting individual frames or data blocks along a time axis) predefined by the content creator. The Tap gesture is generally used to play the entire segment of the main video in a single flow after the initial screen and is suitable for advertising content or simple playback purposes. On the other hand, the Hold gesture dynamically changes the number of frames played depending on the user's input holding time (e.g., plays as many video frames as set during content creation while the user maintains the touch state) and is suitable for stop or loop-based interaction content. Additionally, the Drag, Rotate, Pinch, and Spread gestures compare and map in real-time the input range of movement actions after screen touch predefined by the content creator with the physical input range (distance, direction, etc.) of the movement actions after screen touch actually performed by the user, and dynamically play the corresponding video frame segment; this is suitable for content where the video playback range is adjusted in real-time according to physical parameters such as input range, speed, and direction. For example, when a drag gesture is input, the system can be configured to play different video clips depending on the direction of the drag—up, down, left, or right—from the same starting point, or the range of frames played can be set based on the drag length, thereby increasing or decreasing the number of frames played. As a result, users can assume the role of participants who directly manipulate the progression of the content, rather than merely being passive viewers.
[0049] Furthermore, during content creation, a multi-responsive interface can be designed by varying only the input method at the same on-screen location through trigger structures defined differently for each gesture. For example, a tap gesture at the same location initiates simple playback, a drag gesture unfolds a branched storyline, and a hold gesture triggers a temporary information layer. This granular responsiveness enhances content immersion and enables user-action-based video storytelling.
[0050] In addition, in a user interactive content environment, various touch gestures occur according to the user's intent, and in certain situations, different types of gestures frequently overlap and are input. If the system applies only a simple single input interpretation method to such overlapping inputs, malfunctions or unnatural content responses may occur. Therefore, by embedding an overlapping input response algorithm in the gesture recognition unit (200), precise video content control that matches the user's intent is enabled even in complex gesture situations.
[0051] The gesture recognition unit (200) first detects whether multiple gestures are input in a temporal or spatially overlapping manner, and then quantitatively analyzes the characteristic values (coordinates, direction of movement, speed, duration, etc.) of each gesture. Based on the analysis results, the system is configured to either integrate them into a single command according to the priority between the gestures or to branch them into multiple commands according to an interpretation algorithm. For example, if a Spread or Pinch gesture for zooming in or out of content and a Rotate gesture for switching video clips are input simultaneously, a specific gesture may be selected according to a pre-set priority table and integrated into a single command. On the other hand, if a combination of mutually non-conflicting gestures is input (e.g., Tap and Drag), the gesture recognition unit (200) separates each into independent commands and executes them in parallel.
[0052] Meanwhile, the gesture recognition unit (200) can detect when a user touches a specific point or a specific coordinate area on the screen, or performs a drag gesture to that point. At this time, the gesture recognition unit (200) determines whether the input location corresponds to a 'gesture trigger point' predefined by the content creator. If the location is mapped, the first clip (the corresponding gesture clip) linked to that location is played immediately once, and then a basic action clip (e.g., character's eye blinking, facial expression change, etc.) is played continuously to maintain a natural flow of content for visual consistency. This action can simultaneously guarantee the repeatability of the interaction and immersion through a structural cyclicity of user input → feedback → return to the initial state.
[0053] Furthermore, the gesture recognition unit (200) internally accumulates and records repetitive touch inputs for the same location, and based on this, can precisely track the number of inputs. In this process, the gesture recognition unit (200) can define different content responses for each number of inputs, and can be configured to repeat the existing action once again upon the first re-touch (second touch). Subsequently, if the user's repetitive input exceeds a specific threshold (e.g., when there is a third or more consecutive touch), the gesture recognition unit (200) automatically plays a predefined hidden content (Hidden Clip). The said hidden content is not easily exposed in the normal usage flow and has the meaning of an expandable content accessible only through the user's active exploration or repetitive interaction.
[0054] Additionally, the gesture recognition unit (200) can be configured to diversify the content flow according to user input by applying one or more gestures to a single video content in parallel or on a conditional basis. In particular, completely different video content can be played depending on the direction in which the user drags their finger from the same touch starting point. For example, if the user drags to the right on the same image, a product information video may be played, and if they drag to the left, a user review video may be played. This is made possible by the gesture recognition unit (200) precisely determining the input direction (vector direction) and pre-mapping content branching conditions for each direction.
[0055] In addition, for drag gestures, the user's finger movement speed (speed parameter) is measured in real time to dynamically adjust the video playback range. The playback frame range or video type can be changed according to the input speed, such as by playing only the key summary section of the entire video when the user drags quickly, and playing the detailed video when dragging slowly.
[0056] Additionally, regarding the rub gesture, if the user rapidly rubs the timeline active area at the bottom of the screen in the left-right direction, the gesture recognition unit (200) recognizes that the preset speed threshold has been exceeded and can immediately perform jump playback to the starting point of the 'next chapter'. Conversely, if the same area is rubbed slowly, it switches to scrub playback mode, allowing forward and backward navigation frame by frame, so the user can accurately find the desired scene.
[0057] Additionally, in Match Highlight mode, if a user performs a diagonal swipe (Rub) gesture, pre-set highlight clips corresponding to that diagonal direction can be invoked. For example, swiping diagonally toward the top right immediately switches to a 'scoring scenes' clip, while swiping diagonally toward the bottom left switches to a 'defense scenes' clip, allowing users to intuitively replay key matches.
[0058] In addition, when a user swipes slowly in the left-right direction on the science experiment video playback screen, the scrub speed (a parameter that adjusts the amount of frame movement on the timeline in real time according to the speed at which the user swipes the screen) is set slowly, allowing the complex reaction process to be observed slowly frame by frame; conversely, when swiping quickly in the up-down direction, the entire experiment process is played back at high speed, allowing the overall flow of the science experiment to be checked at a glance.
[0059] FIG. 6 is a diagram showing a method in which a video is played differently when a gesture input is interrupted according to an embodiment of the present invention.
[0060] Referring to Fig. 6, the system includes response processing logic that determines the content state at a given point in time in real time when gesture input is interrupted, and determines the direction of subsequent development. In the case of the Tap gesture, since unidirectional playback and the number of frames corresponding to the gesture input are set to a minimum, playback automatically ends upon completion without separate follow-up processing after user intervention ends. On the other hand, for gestures that require continuous input, such as Hold, Drag, Rotate, Pinch, and Spread gestures, three follow-up processing methods are considered sequentially. These are rewind, separate action clip playback, and pause methods.
[0061] The rewind method reverses the content from the point where the user stopped input, naturally transitioning to the initial screen or the start frame. This method encourages the content to play repeatedly without being artificially interrupted and is suitable for repetitive experience interfaces. The second method plays a separate action clip (e.g., a responsive feedback video) predefined by the content creator and returns to the initial screen after the clip has finished playing; this method visually reflects the results of the user's actions. The third method, the pause method, stops content playback at the point where input stops and transitions to a standby state. Subsequently, it allows the content to resume playback upon the user's re-input or to switch to different content based on new gesture input. For example, if the user inputs the remainder of a previously performed gesture while paused, the video resumes playback from the point where it was stopped; if another gesture is input while paused, the paused state is released and the video action corresponding to the newly input gesture is played; and if there is no input for a certain period of time (e.g., 10 seconds) while the video is paused, it automatically returns to the initial screen.
[0062] FIG. 7 is a diagram showing a method of playback in which an image according to one embodiment of the present invention is played after all the entire frames set have been played.
[0063] Referring to FIG. 7, the post-playback routine, which is automatically performed after the entire frame of the video content has finished playing, is consistently controlled by the natural return and repetition of the video flow, scenario branching, and subsequent content transition. In particular, to enable the system to actively adjust the playback direction of the content while maintaining user immersion, multiple control sequences such as rewind, repeat playback, auto-play, and scene transition processing are structured.
[0064] First, when the entire frame of the video content has finished playing, the system prioritizes performing a rewind operation. In the case of the Tap gesture, rewind is not applied because the number of playback frames is limited; instead, the content creator maintains a sense of visual continuity by composing the last frame with the same scene as the initial screen. On the other hand, for gestures such as Drag, Hold, Rotate, Pinch, and Spread, the system reverses playback from the last frame of the input to the first frame of the video, naturally returning the user to the beginning of the content. This process is designed to maintain a smooth visual flow without any temporary pause.
[0065] After the rewind is completed, the system returns to the user input stage and performs the same gesture-based content playback once again, or repeats the content according to a predefined repetition count condition (N times) and then automatically plays subsequent content according to an autoplay condition. The autoplay condition is divided into two methods. The first method is a method of automatically playing additional content that naturally follows the last frame of the existing video, and the second method is a method of constructing a non-linear content structure (branching narrative) by switching and playing a new video clip independent of the existing content. Both methods are designed so that the content flow can be maintained without user input, and automatic switching is performed according to the user's input pattern and system setting conditions.
[0066] In addition, fade-out and fade-in transition effects are applied at the end of the video to minimize visual discontinuity between content and to induce user gesture input on newly loaded content. These scene transition effects enable an immersive connection to new scenes, and the framework is designed to allow the application of the same or extended gesture rules to subsequent content.
[0067] Fade-out and fade-in transition effects for video content are broadly classified into two types. The first method involves returning to an initial screen set within the same video as the content just played and fading in. In this case, the content re-enters the user gesture input stage, which is the initial point of interaction, and the system is configured to perform one repeated playback following the same flow. This method is suitable for content types intended for repeated user engagement, such as repetitive learning, product demonstrations, and advertisements, and is effective for implementing reset-based content loop structures. The second method involves transitioning to new video content by fading the screen in to a separate, independent scene pre-set by the system or the content creator. In this case, the flow of the content branches off from a single narrative or interaction structure and naturally leads into a video sequence with a completely different composition. This method is suitable for series-type content, branched storytelling, or episodic development structures, and is configured to dynamically select and connect subsequent scenes based on user input history or content playback conditions.
[0068] To explain the separate additional video autoplay method in more detail, it is as follows. The separate additional video autoplay method can be classified into the following three types. The first method automatically plays subsequent video content that naturally takes over from the last frame of the currently played video content following user input, including gestures such as Tap, Drag, Hold, Rotate, Pinch, and Spread (a method that visually takes the last frame of the currently played video content and plays it as the first frame of the subsequent video content). This maintains the continuity of the flow of time and provides the user with the perception that a single scene flow continues without interruption, making it suitable for narrative content where the story unfolds sequentially. It is structured in an edited form so that the final scene of the video connects seamlessly, both visually and thematically, to the introductory scene of the next content, and the content expands automatically under system-led initiative without additional user input. The second method is a structure that automatically connects and plays a separate video clip as new content, completely independent of the previously played video, after the same gesture input. This method is suitable for non-linear narrative structures where the content path switches based on user choices, branching conditions, or repetition counts, rather than continuity with the existing story. Each video clip is selectively called according to predefined trigger conditions, and the system can provide multiple development directions within a single content sequence by automatically loading and playing the next clip based on this. The third method is a structure in which the current video scene is visually terminated through a fade-out effect at the end of video playback, and then transitions to a new video scene designated by the system using a fade-in method. A new gesture input condition predefined by the content creator is set in the video scene, and the user directly manipulates the content flow of the next video by inputting a new gesture on that scene.The above method can induce clear transitions between content units while minimizing visual dissonance, and is suitable for scene-based structural content design.
[0069] Specific embodiments have been illustrated and described above. However, the invention is not limited to the embodiments described above, and those skilled in the art may make various modifications without departing from the essence of the technical concept of the invention as described in the following claims. Explanation of the symbols
[0070] 100: Interactive video playback system based on user operation 110: Input section 120: Output section 130: Communications Department 140: Storage section 150: Control unit 160: Memory section 170: Content Entry 180: User Identifier 190: Initial screen output section 200: Gesture recognition unit 210: Playback video control unit 220: User Interface Section
Claims
Claim 1 A user operation-based interactive video playback system comprising: an input unit that receives data from the outside into the system; an output unit that outputs result data to the outside; a communication unit that transmits and receives data between the inside and outside of the system; a storage unit that stores data generated by the system; a control unit that controls the operation of the system; and a memory unit that stores data resulting from the execution of various programs, wherein the memory unit comprises: a content entry unit that allows a user to enter content through a user terminal; a user identification unit that identifies individual activities of the user; an initial screen output unit that provides an initial screen of the content to the user; a gesture recognition unit that recognizes the user's gesture; a playback video control unit that plays the video of the content according to the user's gesture; and a user interface unit that collects interaction data between the user and the system and provides it to the user, wherein the user identification unit preserves the user's input data in session units for a certain period of time without using cookies or local storage. Claim 2 In claim 1, the content entry unit is a user operation-based interactive video playback system that allows the user to automatically access the content by scanning a QR code attached to a product, exhibit, or poster. Claim 3 In claim 1, a user operation-based interactive video playback system that creates a unique session based on IP address and browser information collected from the user terminal by the user identification unit. Claim 4 In claim 1, the user identification unit preserves the user's input data on a session basis for a certain period of time without using cookies or local storage, in a user operation-based interactive video playback system. Claim 5 A user operation-based interactive video playback system comprising: an input unit that receives data from the outside into the system; an output unit that outputs result data to the outside; a communication unit that transmits and receives data between the inside and outside of the system; a storage unit that stores data generated by the system; a control unit that controls the operation of the system; and a memory unit that stores data resulting from the execution of various programs, wherein the memory unit comprises: a content entry unit that allows a user to enter content through a user terminal; a user identification unit that identifies individual activities of the user; an initial screen output unit that provides an initial screen of the content to the user; a gesture recognition unit that recognizes the user's gesture; a playback video control unit that plays the video of the content according to the user's gesture; and a user interface unit that collects interaction data between the user and the system and provides it to the user, wherein the user identification unit is configured such that the session is automatically terminated when the browser is closed or the page is refreshed. Claim 6 A user operation-based interactive video playback system comprising: an input unit that receives data from the outside into the system; an output unit that outputs result data to the outside; a communication unit that transmits and receives data between the inside and outside of the system; a storage unit that stores data generated by the system; a control unit that controls the operation of the system; and a memory unit that stores data resulting from the execution of various programs, wherein the memory unit comprises: a content entry unit that allows a user to enter content through a user terminal; a user identification unit that identifies individual activities of the user; an initial screen output unit that provides an initial screen of the content to the user; a gesture recognition unit that recognizes the user's gesture; a playback video control unit that plays the video of the content according to the user's gesture; and a user interface unit that collects interaction data between the user and the system and provides it to the user, wherein the initial screen output unit automatically plays a repetitive motion animation prior to the playback of the main video to induce immersion in the user. Claim 7 In claim 1, the initial screen output unit outputs a still cut associated with the video to allow the user to recognize the atmosphere and flow of the content in advance, thereby creating a user operation-based interactive video playback system. Claim 8 A user operation-based interactive video playback system comprising: an input unit that receives data from the outside into the system; an output unit that outputs result data to the outside; a communication unit that transmits and receives data between the inside and outside of the system; a storage unit that stores data generated in the system; a control unit that controls the operation of the system; and a memory unit that stores the execution of various programs and the data resulting therefrom, wherein the memory unit comprises: a content entry unit that allows a user to enter content through a user terminal; a user identification unit that identifies individual activities of the user; an initial screen output unit that provides an initial screen of the content to the user; a gesture recognition unit that recognizes a gesture of the user; a playback video control unit that plays a video of the content according to the gesture of the user; and a user interface unit that collects interaction data between the user and the system and provides it to the user, wherein the initial screen output unit provides an initial scene to operate as a preparation section until the user inputs a gesture, and subsequently allows a smooth visual transition to the content. Claim 9 In claim 1, the gesture recognition unit analyzes the input coordinates, movement direction, movement speed, and duration of a user screen touch to recognize the user's gesture in real time, thereby forming a user operation-based interactive video playback system. Claim 10 In claim 1, the gesture recognition unit classifies six gesture types—Tap, Hold, Drag, Rotate, Pinch, and Spread—and generates a control command corresponding to each gesture, in a user operation-based interactive video playback system. Claim 11 In claim 1, the gesture recognition unit is linked with a trigger point designated by a content creator to call a specific video clip or frame segment, thereby forming a user operation-based interactive video playback system. Claim 12 A user operation-based interactive video playback system comprising: an input unit that receives data from the outside into the system; an output unit that outputs result data to the outside; a communication unit that transmits and receives data between the inside and outside of the system; a storage unit that stores data generated by the system; a control unit that controls the operation of the system; and a memory unit that stores data resulting from the execution of various programs, wherein the memory unit comprises: a content entry unit that allows the user to enter content through a user terminal; a user identification unit that identifies the individual activities of the user; an initial screen output unit that provides the user with an initial screen of the content; a gesture recognition unit that recognizes the user's gesture; a playback video control unit that plays the video of the content according to the user's gesture; and a user interface unit that collects interaction data between the user and the system and provides it to the user, wherein the gesture recognition unit, when multiple gestures are input in a superposition, processes them as a single command or processes them in parallel as multiple commands according to a pre-set priority table. Claim 13 In claim 1, the playback video control unit is a user operation-based interactive video playback system that dynamically adjusts the range of video playback or transition effects of the content according to the user's gesture. Claim 14 A user operation-based interactive video playback system comprising: an input unit that receives data from the outside into the system; an output unit that outputs result data to the outside; a communication unit that transmits and receives data between the inside and outside of the system; a storage unit that stores data generated by the system; a control unit that controls the operation of the system; and a memory unit that stores the execution of various programs and the data resulting therefrom, wherein the memory unit comprises: a content entry unit that allows a user to enter content through a user terminal; a user identification unit that identifies individual activities of the user; an initial screen output unit that provides an initial screen of the content to the user; a gesture recognition unit that recognizes a gesture of the user; a playback video control unit that plays the video of the content according to the gesture of the user; and a user interface unit that collects interaction data between the user and the system and provides it to the user, wherein when the user's gesture input is interrupted, the playback video control unit selects and performs one of the following methods: a method of reverse-playing the video content to the initial screen or start frame based on the point in time when the user stopped input; a method of playing video content predefined by a content creator; or a method of stopping the playback of the video content and switching to a standby state at the point in time when the user stopped input. Claim 15 A user operation-based interactive video playback system comprising: an input unit that receives data from the outside into the system; an output unit that outputs result data to the outside; a communication unit that transmits and receives data between the inside and outside of the system; a storage unit that stores data generated in the system; a control unit that controls the operation of the system; and a memory unit that stores data resulting from the execution of various programs, wherein the memory unit comprises: a content entry unit that allows a user to enter content through a user terminal; a user identification unit that identifies individual activities of the user; an initial screen output unit that provides an initial screen of the content to the user; a gesture recognition unit that recognizes the user's gesture; a playback video control unit that plays the video of the content according to the user's gesture; and a user interface unit that collects interaction data between the user and the system and provides it to the user, wherein the playback video control unit selects and performs one of the following methods: sequentially playing the video content in reverse from the last frame to the first frame after playing all the frames of the video content set; visually taking the last frame of the video content and playing it as the first frame of the subsequent video content; or playing a new video content that is completely different from the video content. Claim 16 In claim 1, the playback video control unit applies fade-in and fade-out effects when switching subsequent content to maintain visual smoothness, thereby forming a user operation-based interactive video playback system.
Citation Information
Patent Citations
Image display apparatus and method for operationg the same
KR101691795B1
Animation-thumbnail browser system
KR1020100058218A
Method and system for controlling play of multimeida content
KR1020180096857A
Methods for controlling and interacting with a three-dimensional environment.
KR1020250075620A