Highlight video providing system based on interaction between wearable device, user terminal and video server
Patent Information
- Application Number
- KR1020260009695
- Authority / Receiving Office
- KR · KR
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2026-01-19
- Publication Date
- 2026-09-29
- Estimated Expiration
- 2046-01-19
Smart Images

Figure 112026006959775-PAT00005_ABST
Abstract
Description
Technology Field
[0001] The present invention relates to a highlight video providing system comprising a wearable device, a user terminal, and a video server, which generates at least one highlight video through exercise video based on interaction between devices and provides it to the user. Background Technology
[0002] Conventional sports video services often require users to manually search for desired scenes by retrospectively reviewing the entire recorded footage, which makes it difficult to quickly secure highlight moments and reduces user convenience. In particular, there is a frequent lack of input methods to immediately display specific scenes during exercise, or even when input is available, the synchronization between that input and the recorded video is limited, leading to interruptions in the highlight generation process.
[0003] Conventional automatic highlight generation technologies often rely on batch analysis of the entire video to detect events, which can lead to increased server computation load and network traffic, as well as processing delays. Furthermore, if videos that are excessively short or long, or of low quality in terms of resolution, frame rate, or clarity, are processed in the same manner, the quality of the highlight results will be inconsistent, resulting in unnecessary resource consumption.
[0004] Furthermore, conventional technology often selects highlight sections using uniform thresholds or single criteria, despite the fact that user movements and the movement characteristics of exercise tools (e.g., balls, equipment) differ depending on the sport, making optimization for each sport difficult. For instance, while the selection criteria for highlight candidate frames need to differ between sports where the movement of exercise tools is prominent, such as soccer, and sports where user body movements are relatively important, such as Pilates, sufficient technical means to systematically reflect this have not been provided. The problem to be solved
[0005] According to the present invention, through the interaction between the pressure input of a wearable device, a user terminal, and a video server, the user can more intuitively search for a video at a desired point in time, and by determining conditions based on the metadata of the target video at the user terminal and then performing a highlight request, user convenience and the speed of obtaining highlights can be improved. As a result, the user can quickly select a desired video among a plurality of exercise videos and receive a highlight of a certain length (e.g., 15 seconds) from the server.
[0006] According to the present invention, whether to request a highlight can be controlled based on metadata including the total playback time, motion time, resolution, frame rate, sharpness, and image quality score of a target image, thereby reducing unnecessary server processing for images of low quality or those that do not meet the conditions. Accordingly, server computational load and network traffic are reduced, and processing delays are alleviated, resulting in improved stability and responsiveness of the entire system.
[0007] According to the present invention, a video server calculates the amount of movement of a user and the amount of movement of an exercise tool (or exercise object) on a frame-by-frame basis, and selects frames in which the amount exceeding the overall average is greater than or equal to a threshold as target frames. Furthermore, by setting the relationship between a first value and a second value differently according to the type of exercise set by the user (e.g., soccer, Pilates), it is possible to identify highlights optimized for the characteristics of the sport. For example, by selecting highlight candidate frames in a manner that relatively emphasizes the movement characteristics of the exercise tool (e.g., soccer ball) in the case of soccer, and relatively emphasizes the user's body movements in the case of Pilates, highlights of consistent quality can be provided for each sport.
[0008] The technical problems of the present invention are not limited to those mentioned above, and other unmentioned technical problems will be clearly understood by those skilled in the art from the description below. means of solving the problem
[0009] A highlight video providing system according to one embodiment of the present invention, comprising a wearable device, a user terminal, and a video server, and generating at least one highlight video through an exercise video based on interaction between devices and providing it to a user, may be configured such that the wearable device transmits a highlight video generation signal to the user terminal in response to receiving a predetermined first pressure input from the user, the user terminal provides the user with a video list including a thumbnail of at least one exercise video received from the video server in response to receiving the highlight video generation signal, and when the user terminal receives a selection input for a target exercise video from the user, determines whether target metadata corresponding to the target exercise video satisfies a predetermined condition including a video length condition and a video quality condition, and if the target exercise video satisfies the predetermined condition, the user terminal transmits a highlight video request signal including the target exercise video to the video server, and the video server generates at least one highlight video corresponding to 15 seconds based on the target exercise video and provides it to the user terminal.
[0010] According to one embodiment, the highlight video providing system may be configured such that the video server identifies, among a plurality of exercise videos registered in correspondence with the user, an image in which a second pressure input distinct from the first pressure input is acquired from the wearable device during shooting as the at least one exercise video, and the video server identifies, for the at least one exercise video, the sum of the playback times of the sections in which the user's first movement and the exercise tool's second movement are each greater than or equal to the first movement amount and the second movement amount, respectively, as the exercise time, and then stores metadata including the identification result by matching it to the at least one exercise video.
[0011] According to one embodiment, the highlight video providing system may be configured such that the user terminal checks the user grade corresponding to the wearable device, and if the user terminal's user grade corresponds to the premium grade, transmits the highlight video request signal to the video server even if the target metadata does not satisfy the predetermined condition, and if the user terminal's user grade corresponds to the general grade, checks the total playback time and exercise time of the target exercise video based on the metadata, and determines that the target metadata satisfies the predetermined condition only if the total playback time of the target exercise video is greater than or equal to the first time and the ratio of the exercise time to the total playback time is greater than or equal to the threshold ratio.
[0012] According to one embodiment, the highlight video providing system may be configured such that the user terminal checks quality parameters including resolution, frame rate, and image clarity of the target motion video based on the metadata, and then checks whether a first condition is satisfied regarding whether each of the quality parameters is greater than or equal to a set reference value; if the quality parameters satisfy the first condition, the user terminal checks whether a second condition is satisfied regarding whether an image quality score calculated based on the resolution, frame rate, and image clarity is greater than or equal to a threshold score; and the user terminal determines that the target metadata satisfies the predetermined condition only when the image quality score is greater than or equal to the threshold score.
[0013] According to one embodiment, the highlight video providing system is configured such that the video server divides the target exercise video into a plurality of frames, checks the first average movement amount of the user and the second average movement amount of the exercise tool during the entire playback time of the target exercise video, and the first partial movement amount of the user and the second partial movement amount of the exercise tool for each of the plurality of frames, and the video server identifies a frame among the plurality of frames in which the value obtained by subtracting the first average movement amount from the first partial movement amount is greater than or equal to a positive first value, and the value obtained by subtracting the second average movement amount from the second partial movement amount is greater than or equal to a positive second value as a target frame, and the video server identifies at least one highlight video based on the target frame, and the video server determines the first value and the second value based on an exercise type set by the user corresponding to the target exercise video, and the video server is configured to set the first value smaller than the second value if the exercise type corresponds to soccer, and to set the first value larger than the second value if the exercise type corresponds to Pilates.
[0014] The above highlight video providing system is,
[0015] The above user terminal is configured to calculate the image quality score based on the following mathematical formula 1, and
[0016] [Mathematical Formula 1]
[0017]
[0018] SQ is the image quality score, K is a positive constant for determining the scale of the score, a, b, and c are positive exponential coefficients for determining the sensitivity of each term, p is a positive coefficient for determining the intensity of the blur penalty, and m is a small positive constant for ensuring numerical stability of the logarithmic calculation when B is close to 0,
[0019] R is a value corresponding to resolution, F is a value corresponding to frame rate, S is a value corresponding to image sharpness index, R_0, F_0, and S_0 are normalized reference values corresponding to reference resolution, reference frame rate, and reference sharpness, respectively, B is a blur index representing the degree of motion blur or blurring, B0 is a reference blur index, and J is a jitter index representing shaking or stabilization quality, corresponding to a value such as the variability of global movement between frames or stabilization residual.
[0020] The above-described highlight video providing system is configured such that the user terminal monitors the playback status of the highlight video to check the viewing duration, and if it is confirmed that the viewing duration has exceeded a threshold time, virtual points proportional to the total playback length of the target exercise video are awarded to the user's account information based on a set winning probability.
[0021] The above-described highlight video providing system is configured such that the video server calculates the total sum of target virtual assets consumed as a third user accessing the highlight video via a shared link uses a paid function, and distributes reward virtual assets corresponding to a predetermined ratio of the total sum to the user's account information.
[0022] The above highlight video providing system may be configured such that the video server sets a video time point corresponding to the occurrence time of the second pressure input as a reference time point, searches for a peak section in which the time change pattern of the user's first partial movement amount and the exercise tool's second partial movement amount has maximum correlation within a predetermined time range before and after the reference time point, determines the start time point and end time point of the highlight section to include the peak section, and generates a highlight video by normalizing the determined highlight section to a length of 15 seconds.
[0023] The highlight video providing system described above may be configured such that the video server identifies a cluster of motion videos including two or more videos determined to be the same motion session among at least one motion video, selects the highest quality section by comparing the video quality scores of the two or more motion videos with respect to a reference section based on the time of occurrence of the second pressure input, and synthesizes the highest quality section into a single 15-second highlight video by time-axis alignment.
[0024] The highlight video providing system is configured such that the video server identifies a highlight section and generates explanatory metadata indicating the basis for the identification of the highlight section, and provides this metadata to the user terminal, wherein the explanatory metadata is configured to include at least one of the time of occurrence of the second pressure input, the time distribution of the target frame, the magnitude of the excess amount of movement, and the result of applying the motion type-based threshold. Effects of the invention
[0025] The effects of the highlight image providing system according to the embodiments of the present invention are described as follows.
[0026] According to the present invention, through the interaction between the pressure input of a wearable device, a user terminal, and a video server, the user can more intuitively search for a video at a desired point in time, and by determining conditions based on the metadata of the target video at the user terminal and then performing a highlight request, user convenience and the speed of obtaining highlights can be improved. As a result, the user can quickly select a desired video among a plurality of exercise videos and receive a highlight of a certain length (e.g., 15 seconds) from the server.
[0027] According to the present invention, whether to request a highlight can be controlled based on metadata including the total playback time, motion time, resolution, frame rate, sharpness, and image quality score of a target image, thereby reducing unnecessary server processing for images of low quality or those that do not meet the conditions. Accordingly, server computational load and network traffic are reduced, and processing delays are alleviated, resulting in improved stability and responsiveness of the entire system.
[0028] According to the present invention, a video server calculates the amount of movement of a user and the amount of movement of an exercise tool (or exercise object) on a frame-by-frame basis, and selects frames in which the amount exceeding the overall average is greater than or equal to a threshold as target frames. Furthermore, by setting the relationship between a first value and a second value differently according to the type of exercise set by the user (e.g., soccer, Pilates), it is possible to identify highlights optimized for the characteristics of the sport. For example, by selecting highlight candidate frames in a manner that relatively emphasizes the movement characteristics of the exercise tool (e.g., soccer ball) in the case of soccer, and relatively emphasizes the user's body movements in the case of Pilates, highlights of consistent quality can be provided for each sport.
[0029] In addition, various effects that can be identified directly or indirectly through this document may be provided. Brief explanation of the drawing
[0030] FIG. 1 is a block diagram showing the components of a user terminal according to one embodiment of the present invention. FIG. 2 is a block diagram showing the components of a highlight video providing system including a wearable device, a user terminal, and a video server according to an embodiment of the present invention. FIG. 3 is a flowchart of the operation of a highlight image providing system according to one embodiment of the present invention. FIG. 4 is a flowchart of the operation of a highlight image providing system according to one embodiment of the present invention. In relation to the description of the drawings, the same or similar reference numerals may be used for identical or similar components. Specific details for implementing the invention
[0031] Hereinafter, some embodiments of the present invention will be described in detail with reference to exemplary drawings. It should be noted that in assigning reference numerals to the components of each drawing, the same components are given the same reference numeral whenever possible, even if they are shown in different drawings. Furthermore, in describing the embodiments of the present invention, if it is determined that a detailed description of related known components or functions would hinder understanding of the embodiments of the present invention, such detailed description is omitted.
[0032] In describing the components of the embodiments of the present invention, terms such as first, second, A, B, (a), (b), etc., may be used. These terms are intended merely to distinguish the components from other components, and the essence, order, or sequence of the components is not limited by the terms. Furthermore, unless otherwise defined, all terms used herein, including technical or scientific terms, have the same meaning as generally understood by those skilled in the art to which the present invention pertains. Terms such as those defined in commonly used dictionaries should be interpreted as having a meaning consistent with their meaning in the context of the relevant technology, and should not be interpreted in an ideal or overly formal sense unless explicitly defined in this application.
[0033] Hereinafter, embodiments of the present invention will be described in detail with reference to FIGS. 1 to 4.
[0035] FIG. 1 is a block diagram showing the components of a user terminal according to one embodiment of the present invention.
[0036] According to one embodiment, the user terminal (100) may include a memory (110), a processor (120), a communication interface (130), and / or a display device (140). The configuration of the user terminal (100) illustrated in FIG. 1 is exemplary and the embodiments of the present invention are not limited thereto. For example, the user terminal (100) may further include components not illustrated in FIG. 1 (e.g., a web crawler, a user interface, an input device, a notification unit, a sensor unit, or at least one of any combination thereof).
[0037] According to one embodiment, the memory (110) may store instructions or data. For example, the memory (110) may store one or more instructions that cause the user terminal (100) to perform various operations when executed by the processor (120).
[0038] For example, the memory (110) may be implemented as a single chipset with the processor (120). The processor (120) may include at least one of a communication processor or a modem.
[0039] For example, the memory (110) can store various information related to the user terminal (100). For example, the memory (110) can store information regarding the operation history of the processor (120). For example, the memory (110) can store input data acquired by the user terminal (100), output data output by the user terminal (100), data acquired from an external server and / or user terminal, etc.
[0040] For example, the memory (110) may include multiple storage devices of different types. For example, the memory (110) may include volatile and / or non-volatile storage media. For example, the memory (110) may include at least one of RAM (random-access memory), ROM (read only memory), eMMC (Embedded Multi-Media Card), or any combination thereof.
[0041] The steps of the method or algorithm described in connection with the embodiments disclosed in this specification may be directly implemented in hardware, software modules, or a combination of both, executed by the processor (120). The software modules may reside in a storage medium (i.e., memory (110)) such as RAM memory, flash memory, ROM memory, EPROM memory, EEPROM memory, registers, a hard disk, a removable disk, or a CD-ROM.
[0042] For example, the memory (110) is coupled to a processor (120), and the processor (120) can read information from a storage medium and write information to a storage medium. Alternatively, the memory (110) may be integrated with the processor (120). The memory (110) and the processor (120) may reside within an application-specific integrated circuit (ASIC). The ASIC may reside within a user terminal. Alternatively, the memory (110) and the processor (120) may reside as separate components within the user terminal.
[0043] According to one embodiment, the processor (120) may be operatively connected to the memory (110), the communication interface (130), and / or the display device (140). For example, the processor (120) may control the operation of the memory (110), the communication interface (130), and / or the display device (140).
[0044] According to one embodiment, the communication interface (130) may support the establishment of a direct (e.g., wired) communication channel or a wireless communication channel between a user terminal (100) and an external device (e.g., a wearable device (210) of FIG. 2, a video server (220), and a database, etc.), and the performance of communication through the established communication channel. The communication interface (130) may include one or more communication processors that operate independently of the processor (120) (e.g., an application processor) and support direct (e.g., wired) communication or wireless communication. According to one embodiment, the communication interface (130) may include a wireless communication module (e.g., a cellular communication module, a short-range wireless communication module, or a GNSS (global navigation satellite system) communication module) or a wired communication module (e.g., a LAN (local area network) communication module, or a power line communication module). The corresponding communication module among these communication modules can communicate with an external electronic device through a first network (e.g., a short-range communication network such as Bluetooth, WiFi (wireless fidelity) direct, or IrDA (infrared data association)) or a second network (e.g., a legacy cellular network, a 5G network, a next-generation communication network, the Internet, or a long-range communication network such as a computer network (e.g., LAN or WAN). These various types of communication modules may be integrated into a single component (e.g., a single chip) or implemented as multiple separate components (e.g., multiple chips). The wireless communication module can identify or authenticate a user terminal (100) within a communication network, such as the first network or the second network, using subscriber information (e.g., International Mobile Subscriber Identifier (IMSI)) stored in a subscriber identification module.
[0045] According to one embodiment, the display device (140) may include at least one output device that provides various information and a user interface to the user.
[0046] For example, the display device (140) may include a display device, an audio output device, a virtual reality output device, etc.
[0047] For example, the display device (140) can provide the administrator with various types of user interfaces (e.g., images) described in the present disclosure visually and / or audibly.
[0048] The components of the user terminal (100) illustrated in FIG. 1 are exemplary, and the embodiments of the present disclosure are not limited thereto.
[0049] At least some of the embodiments of the present disclosure may be implemented as artificial intelligence (AI) through the processor (120) and memory (110) of the user terminal (100). The processor (120) may be composed of one or more processors, and the one or more processors may be general-purpose processors such as a CPU, AP, DSP (digital signal processor), etc., graphics-dedicated processors such as a GPU, VPU (vision processing unit), or artificial intelligence-dedicated processors such as an NPU. The one or more processors may be controlled to process input data according to predefined operation rules or artificial intelligence models stored in the memory (110). Alternatively, if the one or more processors are artificial intelligence-dedicated processors, the artificial intelligence-dedicated processors may be designed with a hardware structure specialized for processing a specific artificial intelligence model.
[0050] The predefined operation rules or artificial intelligence model are characterized by being created through learning. Here, being created through learning means that a predefined operation rules or artificial intelligence model configured to perform a desired characteristic (or purpose) is created by a basic artificial intelligence model being trained using a number of learning data by a learning algorithm. Such learning may be performed on the user terminal (100) itself where the artificial intelligence according to the present disclosure is performed, or it may be performed through a separate server and / or system. Examples of learning algorithms include supervised learning, unsupervised learning, semi-supervised learning, or reinforcement learning, but are not limited to the examples described above.
[0051] An artificial intelligence model may be composed of multiple neural network layers. Each of the multiple neural network layers has multiple weight values and can perform neural network operations through operations between the results of previous layers and the multiple weights. The multiple weights possessed by the multiple neural network layers can be optimized based on the learning results of the artificial intelligence model. For example, the multiple weights can be updated so that the loss value or cost value obtained from the artificial intelligence model during the learning process is reduced or minimized. Artificial neural networks may include, but are not limited to, deep neural networks (DNN), convolutional neural networks (CNN), recurrent neural networks (RNN), restricted Boltzmann machines (RBM), deep belief networks (DBN), bidirectional recurrent deep neural networks (BRDNN), or deep Q-networks.
[0052] The artificial intelligence model can be implemented as an artificial intelligence model based on the relationship between the training input data and the training output data, using exercise performance history information collected from a user terminal and text data related to the exercise type, performance time, exercise intensity, and repetition pattern included in the exercise performance history information as training input data, and at least one highlight video segment among multiple highlight video candidates stored on a video server that corresponds to the text data and is evaluated as having high user preference as training output data, and the artificial intelligence model can be configured to predict a highlight video segment that matches user preference for a new exercise video.
[0053] The artificial intelligence model may be implemented as a large language model or a multilayer neural network model that learns patterns between the input data and the output data, using time-series data related to the viewing history of highlight videos accumulated and stored in correspondence with user identification information, the frequency of highlight video selection, and whether highlight video playback is completed as input data for training, and information on the creation time or video segment of highlight videos judged to have high user satisfaction based on the time-series data as output data for training, and the artificial intelligence model may be configured to automatically determine the length, start time, or included segment of highlight videos to be generated in the future.
[0055] FIG. 2 is a block diagram showing the components of a highlight video providing system including a wearable device, a user terminal, and a video server according to an embodiment of the present invention.
[0056] According to one embodiment, the highlight video providing system may include a user terminal (100), a wearable device (210), and a video server (220).
[0057] According to one embodiment, the user terminal (100) is a control center device of a highlight video provision system and may be configured to control a user interface in response to a highlight video generation signal transmitted from a wearable device (210) and to provide a list of exercise videos and highlight videos through communication with a video server (220). The user terminal (100) may be implemented as various types of electronic devices such as a smartphone, tablet, laptop, wearable-linked terminal, or vehicle infotainment terminal, and may be configured to operate in an application or web-based environment.
[0058] When a user terminal (100) receives a signal corresponding to a first pressure input received from a wearable device (210), it may be configured to display a video list including a thumbnail, shooting time, sport, shooting location, or session identification information regarding at least one exercise video received from a video server (220). At this time, the user terminal (100) may be configured to check access rights stored corresponding to a user account and filter and provide the video list on a session basis or for a certain period of time.
[0059] The user terminal (100) may be configured to determine whether a predetermined condition, including a video length condition and a video quality condition, is satisfied by checking the target metadata corresponding to the target exercise video when the user selects a target exercise video from a video list. The user terminal (100) may perform a condition determination using information such as total playback time, exercise time, resolution, frame rate, clarity, or quality score, and may be configured to control whether to transmit a highlight video request signal according to the result of the condition determination.
[0060] The user terminal (100) can be configured to check the user grade and apply different condition judgment policies to the premium grade and the general grade. For example, in the case of the premium grade, it can be configured to transmit a highlight video request signal even if some conditions are not satisfied, and in the case of the general grade, it can be configured to transmit a request signal only when the ratio of total playback time to exercise time is above a threshold ratio. Accordingly, the user terminal (100) can function as a control node that maintains a balance between server resource efficiency and user experience.
[0061] The user terminal (100) may be configured to collect interaction data, such as the user's playback start time, viewing duration, whether to rewatch, whether to share, or whether to download, after the highlight video is provided, and to transmit it to the video server (220) or store it internally. The user terminal (100) may be configured to evaluate the quality of highlight provision based on the interaction data and to use it as feedback to update a condition judgment threshold or a quality score calculation policy.
[0062] According to one embodiment, the wearable device (210) is an electronic device that can be worn by a user and may be configured to provide a trigger for generating a highlight video based on user input that occurs during exercise. The wearable device (210) may be implemented as a wristband, a watch-type device, a necklace-type device, or a clothing-attached device, and may be configured to transmit a signal via short-range wireless communication regardless of whether it is paired with a terminal.
[0063] The wearable device (210) may be configured to transmit a highlight image generation signal to a user terminal (100) in response to receiving a first pressure input from a user. The wearable device (210) may be configured to provide tactile feedback for the pressure input, and may be configured to distinguish between the first pressure input and the second pressure input by distinguishing different input types according to the number of pressure inputs, duration, or input interval.
[0064] The wearable device (210) may be configured to transmit timestamp information or session identification information, including time information when a pressure input occurs, to a user terminal (100) or a video server (220). Through this, the video server (220) can efficiently search for a scene corresponding to a specific point in time during shooting, and the user terminal (100) can control the flow of video list provision or highlight request in alignment with the user's intent.
[0065] The wearable device (210) may be configured to apply an input detection threshold pressure, an input holding time, or a minimum interval between consecutive inputs to reduce unintended inputs caused by shock or malfunction during movement. The wearable device (210) may be configured to identify an incorrect input pattern and invalidate or process the corresponding input with a low priority, and may be configured to transmit the result of such incorrect input processing to a user terminal (100).
[0066] The wearable device (210) may be configured to periodically transmit status data including device status information, such as battery status, communication status, or waterproof status, or when an event occurs. The user terminal (100) may be configured to calculate input reliability or output user guidance messages by referring to the status data of the wearable device (210), thereby improving the usability of the system as a whole.
[0067] According to one embodiment, the video server (220) may be implemented as a server device that stores a plurality of exercise videos and performs the function of generating and providing highlight videos. The video server (220) may be implemented as a single server, a cloud-based computing device, an edge server, or part of a distributed processing system, and may be configured to provide exercise video data and highlight videos in a streaming or download form in response to a request from a user terminal (100).
[0068] The video server (220) may be configured to identify at least one video among a plurality of registered exercise videos corresponding to a user, in which a second pressure input is obtained from a wearable device (210) during shooting. The video server (220) may match an input event with a video segment using shooting session identification information, input time information, or shooting location information, and may prioritize configuring a candidate for highlight generation target for a video in which an input event exists.
[0069] A video server (220) may be configured to identify the sum of the playback times of the intervals in which the user's first movement and the second movement of the exercise tool or exercise tool (or, exercise object) are each greater than or equal to the first movement amount and the second movement amount, respectively, as the exercise time for at least one exercise video, and to store metadata including the result by matching it to the video. The video server (220) may calculate the exercise time by utilizing the position change of the exercise tool (or, exercise object) on a frame-by-frame basis, the movement amount between frames, or whether the movement amount in consecutive frames is maintained.
[0070] The video server (220) may be configured to divide a target exercise video into multiple frames, check the user's first average movement amount and the exercise tool's second average movement amount and the first and second partial movement amounts per frame during the entire playback time, and select frames as target frames in which the value obtained by subtracting the average movement amount from the partial movement amount is greater than or equal to the first value and the second value, respectively. The video server (220) may be configured to generate a highlight video by identifying a highlight section that includes the time index of the target frame or a surrounding frame section.
[0071] The video server (220) can be configured to determine a first value and a second value based on the exercise type set by the user, and the threshold values can be set differently by reflecting the relative importance of the user's movement and the movement of the exercise tool for each exercise type. For example, in sports where the movement of the exercise tool (e.g., soccer ball) is prominent, such as soccer, the contribution of the second value is set relatively large, and in sports where the user's body movement is central, such as Pilates, the contribution of the first value is set relatively large, thereby ensuring consistent highlight quality for each sport.
[0073] FIG. 3 is a flowchart of a highlight image providing system according to one embodiment of the present invention.
[0074] According to one embodiment, the components of a highlight video providing system may perform the operations disclosed in FIG. 3. For example, at least some of the components included in the highlight video providing system (e.g., the user terminal (100), wearable device (210), and video server (220) of FIG. 2 may be configured to perform the operations of FIG. 3.
[0075] In the following embodiments, the operations S310 to S350 may be performed sequentially, but are not necessarily performed sequentially. For example, the order of each operation may be changed, and at least two operations may be performed in parallel. Additionally, content corresponding to or overlapping with the above description in relation to FIG. 3 may be briefly explained or omitted.
[0076] According to one embodiment, the wearable device may transmit a highlight image generation signal to a user terminal in response to receiving a predetermined first pressure input from a user (S310).
[0077] The wearable device can be implemented in a wrist-worn or necklace-type form so that the user can easily operate it while exercising, and the first pressure input can be defined as an input in which the user presses the input part of the wearable device with a pressure greater than a certain amount.
[0078] A wearable device may include a mechanical switch, a pressure sensor, a capacitive sensor, or a piezoelectric element to detect a press on an input part, and may generate trigger information indicating that an event has occurred at the time of input detection.
[0079] The highlight video generation signal may be composed of minimal information so that the user terminal can immediately perform a subsequent action. For example, the highlight video generation signal may be configured to include at least one of user identification information, wearable device identification information, input time information, input type information, and shooting session identification information.
[0080] Wearable devices can transmit signals using short-range wireless communication. For example, wearable devices can transmit signals based on Bluetooth Low Energy, Low Power Wide Area Communication, or relay devices installed within the facility.
[0081] The wearable device may be configured to provide vibration feedback or light indicators so that the user can intuitively confirm that the input has been properly recognized, which can improve the user experience while increasing the reliability of the input and stabilizing the system operation.
[0082] According to one embodiment, in response to receiving a highlight video generation signal, a user terminal may provide a video list to the user that includes at least one thumbnail of a motion video received from a video server (S320).
[0083] When a user terminal receives a highlight video generation signal from a wearable device, it can send a video list request to a video server based on user identification information or session identification information included in the highlight video generation signal.
[0084] The video server can generate a list of exercise videos corresponding to the user or session and provide it to the user terminal, and the user terminal can receive and display them on the screen.
[0085] The video list may include at least one of the following: a video thumbnail, shooting date and time, type of exercise, video length, whether it is multi-angle, video quality summary information, or an indication of whether a wearable input event exists, allowing the user to quickly identify the desired video.
[0086] For example, the thumbnail may be a representative frame image that the video server has generated and stored in advance, or it may be generated immediately using a specific frame of the video received by the user terminal via streaming.
[0087] For example, the user terminal can be configured to verify access rights using an authentication token in communication with the video server, and even in the case of an unstable network environment, download only minimal data (e.g., thumbnails and metadata) first, and then request detailed information step by step according to the user's choice.
[0088] According to one embodiment, when a user terminal receives a selection input for a target motion video from a user, the processor (120) can determine whether the target metadata corresponding to the target motion video satisfies a predetermined condition including a video length condition and a video quality condition (S330).
[0089] Target metadata may include information stored on the video server or received by the user terminal from the server and stored in the cache. For example, target metadata may include items such as the total playback time, exercise time, resolution, frame rate, image clarity, encoding information, number of shooting angles, or quality score of the target exercise video.
[0090] The video length condition may be implemented simply by determining whether the total playback time exceeds a specific threshold, or the total playback time and exercise time may be utilized separately. For example, even if the total playback time is sufficient, if the actual exercise time is short, the video may have low value as a highlight; therefore, the user terminal may be configured to determine whether the video length condition is satisfied based on the ratio of exercise time to the total playback time. For instance, the exercise time may be the result calculated by the video server analyzing the movement of the user and / or exercise tool on a frame-by-frame basis, and such calculated result may be included in the target metadata.
[0091] Video quality conditions can be implemented by checking whether quality parameters such as resolution, frame rate, and image clarity are above a reference value, and additionally, can be configured to determine whether a quality score is above a threshold score by calculating a quality score that combines multiple parameters.
[0092] The user terminal may recognize the user tapping to select a specific video from the video list as a selection input, or the user long-pressing a thumbnail of a specific video as a selection input.
[0093] The user terminal can reduce unnecessary requests to the video server by performing condition determination internally, and the condition threshold can be determined by referring to a configuration file or server policy value so that it can be changed according to the user class, network status, or facility policy.
[0094] This condition judgment logic can function as a technical means to efficiently utilize server resources while maintaining the quality of highlight generation above a certain level.
[0095] According to one embodiment, when the target exercise video satisfies a predetermined condition, the user terminal can transmit a highlight video request signal including the target exercise video to the video server (S340).
[0096] The highlight video request signal may be configured to include the minimum identification information required for the video server to generate a highlight. For example, the highlight video request signal may include at least one of a target exercise video identifier, a user identifier, a request time, a request type, a desired highlight generation mode, or wearable input event time information.
[0097] For example, the highlight generation mode can be defined as whether to generate highlights in a single segment, in multiple segments, or around a specific event point, and the user can specify this through a simple option selection on the interface of the user terminal. For example, when a user performs a first pressure input to capture a specific moment during a tennis rally and then selects the corresponding session video from the video list, the user terminal confirms that the video length and quality conditions have been met through metadata and then sends a request signal to the server, and the server can generate highlights in a specified manner according to the request.
[0098] Prior to receiving selection input for a target exercise video from the user, the user terminal may display, in an identifiable form, exercise videos from a video list provided by the video server that have a history of a second pressure input occurring, and the user selects a target exercise video from the list. Here, the exercise videos included in the video list are not videos currently being recorded in real time, but may be limited to videos among a plurality of exercise videos already stored in the video server corresponding to user accounts or user identification information in which a second pressure input was recorded during recording.
[0099] A user terminal can be configured to check information such as total playback time, exercise time, resolution, frame rate, sharpness, or quality score included in the target metadata for a target exercise video selected by the user, determine whether a predetermined condition including video length condition and video quality condition is satisfied, and then transmit a highlight video request signal to a video server only if the condition is satisfied.
[0100] The highlight video request signal may include at least one of a target exercise video identifier, user identification information, and shooting session identification information so that the video server can query the exact same target exercise video from the storage, and may be configured to reduce the search range of the video server by including together the occurrence time information of the second pressure input or input event identification information if they are already stored as metadata.
[0101] In this way, when a user terminal performs a precondition determination based on the metadata of a selected target motion video and then transmits a request signal, the video server does not perform unnecessary highlight generation processing for videos that do not meet the conditions, thereby achieving the technical effect of reducing server resources and network traffic.
[0102] The user terminal may assign a request identifier to prevent network delays or duplicate requests when transmitting a request, and provide a progress status indicator based on an acceptance response or processing status information received from the server.
[0103] If the conditions are not met, the user terminal may not transmit a request signal and instead provide the user with the reason, which helps the user understand and reduces unnecessary server calls. In this case, the guidance message may include reasons such as failure to meet video length conditions, failure to meet quality conditions, or network restrictions, allowing a person skilled in the art to configure clear branching logic during implementation.
[0104] According to one embodiment, the processor (120) can provide at least one highlight video corresponding to 15 seconds based on a target motion video to a user terminal (S350).
[0105] When the video server receives a highlight video request signal from a user terminal, it can load a target exercise video file or stream, determine a highlight section, and then encode the section to generate a 15-second highlight video.
[0106] A 15-second length is an example of an implementation that ensures the generated highlight is provided to the user in a short and focused form. To ensure the video server fits exactly 15 seconds, it can cut the start and end points by aligning frames, and if there is audio, it can also synchronize and include the audio stream.
[0107] The highlight video may be provided to the user terminal in a streaming format or as a downloadable file, and the user terminal can play or save it.
[0108] The video server can provide the generated highlight video to the user terminal by adding metadata such as thumbnails, creation timestamps, creation rules, and quality information, and the user can perform subsequent actions such as replaying, sharing, or creating additional clips on the highlight playback screen.
[0109] The video server can query a selected target exercise video among multiple stored exercise videos using a target exercise video identifier included in the request signal, and load the file or stream source of the corresponding video to perform highlight generation processing; in this case, the processing target may be an existing exercise video already stored on the server, rather than a real-time video.
[0110] To determine the highlight section, the video server may set a certain range before and after the video viewpoint corresponding to the time of occurrence of the second pressure input as a candidate section, or it may identify the highlight section by calculating the amount of movement of the user and the exercise tool on a frame-by-frame basis and setting candidate sections centered on frames where the amount of excess relative to the overall average is greater than or equal to a threshold value.
[0111] If the video server contains exercise time information in the metadata already matched and stored with the target exercise video, it can prioritize selecting a section where sufficient exercise time is secured, and the highlight section can be configured to include key scenes of the actual exercise performance.
[0112] The video server can generate a highlight video by cutting the identified highlight segment by aligning the start and end times to the frame boundaries, adjusting the segment length so that the length of the resulting video is 15 seconds, and then encoding it. The generated highlight video can be provided to a user terminal in a streaming manner or as a downloadable file.
[0113] The video server may provide at least one of the creation time, segment information, and quality information of the highlight video as additional metadata during the provision process, and the user terminal may use this to configure the playback screen of the highlight video or subsequently control the flow of additional highlight requests in a consistent manner.
[0114] By generating highlights by combining the stored target motion video with the second pressure input event recorded during shooting, highlights are generated centered on scenes reflecting the user's intent, thereby improving the perceived quality of the result and enabling stable delivery while avoiding the uncertainty of real-time processing.
[0115] From an implementation perspective, the server can enhance responsiveness by using a work queue to handle concurrent requests from multiple users or by caching pre-generated highlight candidates to provide them quickly. Additionally, it can be configured to apply retry logic or alternative generation modes in the event of a highlight generation failure. This server processing flow increases the stability of highlight delivery and supports the technical effectiveness of delivering results within the timeframe expected by users.
[0116] The process of the video server generating (or identifying) highlight images will be described in more detail later in the description of Figure 4 below.
[0117] Additionally or generally, the video server may be configured to identify, among a plurality of exercise videos registered in correspondence with a user, an image in which a second pressure input distinct from the first pressure input is acquired from a wearable device during shooting as at least one exercise video, and for at least one exercise video, identify the sum of the playback times of the intervals in which the user's first movement and the exercise tool's second movement are each greater than or equal to the first movement amount and the second movement amount, respectively, as the exercise time, and then match and store metadata including the identification result to at least one exercise video.
[0118] The following description is based on an embodiment in which, in response to a user, a video in which a second pressure input occurs during recording is selected from among a plurality of previously registered exercise videos on a video server, and the result of the selection is subsequently utilized for selecting a target for highlight generation and determining conditions based on metadata.
[0119] First, the video server can store exercise videos matched to user accounts or user identification information on a per-shooting session basis, and each video can record the shooting start time, shooting end time, shooting device identification information, and time information when an input event occurred during shooting.
[0120] Here, the second pressure input is an input that is distinguished from the first pressure input in terms of input type, and can be defined as an input that has a longer pressing duration, has a specific pattern of consecutive pressing counts (e.g., 3 consecutive times), or satisfies a pressure threshold different from the first pressure input.
[0121] The video server receives an input event record transmitted from a wearable device or a user terminal, determines whether the input type information included in the input event record corresponds to a second pressure input, and can identify the video corresponding to the shooting session in which the input event occurred as the second pressure input acquired video. Furthermore, the meaning of the video server identifying the video in which the second pressure input is acquired goes beyond merely confirming the existence of a video file and may include the meaning of matching and storing the time when the second pressure input occurred with the position on the video time axis that includes that time together. Through this, in a subsequent step, the occurrence of the second pressure input can be indicated in the video list or the scene surrounding the time of the input occurrence can be quickly searched.
[0122] The video server may perform an operation of aligning the time information of an input event and a video to identify at least one motion video based on a second pressure input. For example, when a wearable device generates an event record containing time information that detects the second pressure input and transmits it to a user terminal or a video server, the video server may compare the time information of the event record with the shooting time information of the video file to determine whether the video corresponds to the same shooting session. Since a time synchronization error may exist, the video server may search for a video session containing the event time within a certain allowable error range, or prevent incorrect session matching by referencing the shooting device identification information and the user identification information together. Additionally, considering the case where the second pressure input occurs multiple times in a single shooting session, the video server may store a list of the time points of each input occurrence as supplementary information for the corresponding video. Subsequently, when the user terminal displays a video list, the video may be configured to indicate that the video contains the second pressure input, or to prioritize the exposure of the video in which the latest input or a specific pattern of input occurred among the multiple inputs. This identification process is configured based on the technical premise that highlight videos are generated from footage already stored on a server rather than real-time footage, and it provides the effect of prioritizing videos containing scenes intentionally marked by the user during shooting.
[0123] Subsequently, the video server can calculate and store an indicator called exercise time for each exercise video separately from the total playback time, and the exercise time can be defined to reflect the length of the section where the exercise was actually performed, rather than being equal to the total length of the video.
[0124] The first movement of the user may be defined as a change in the position of the user's body or a part of the body, and the second movement of the exercise tool may be defined as a change in the position of the tool used depending on the sport, such as a ball, racket, club, band, dumbbell, etc.
[0125] The video server can estimate the positions of the user and the exercise tool frame by frame based on frame-unit data of the video, or calculate the amount of movement through the change between frames. For example, the video server can calculate the center coordinates or contour box of an object corresponding to the user for each frame, calculate the magnitude of the coordinate difference between consecutive frames as the user's amount of movement, and calculate the amount of movement for the exercise tool in the same manner. Sections where frames with a calculated amount of movement exceeding a threshold of first or second amount exist consecutively can be identified as exercise performance sections, and the playback times of these sections can be summed to calculate the exercise time.
[0126] The video server can identify reference points for the first movement and the second movement based on the type of movement input by the user immediately before capturing each video. For example, if the type of movement for a specific movement video is input as soccer, the video server can determine the movement of the user's foot and the soccer ball as the first movement and the second movement, respectively, for that movement video.
[0127] The video server may apply additional segment determination rules to enhance the reliability of exercise time calculation. For example, since a large amount of movement calculated in only a single frame may be due to camera shake or momentary noise, the video server can be configured to recognize a segment as a valid exercise segment only if the movement threshold is maintained for a certain number of consecutive frames. Furthermore, as the simultaneity of user movement and exercise tool movement may vary depending on the sport, the video server can be configured to select only segments where both conditions are met simultaneously, or to merge segments within a temporally close range where both conditions are met to determine them as a single exercise segment. For instance, in soccer, considering the moment of sudden ball movement together with the moment accompanied by user movement can more accurately reflect valid scenes; in Pilates, where exercise tools may be absent or movement may be limited, threshold settings or tool movement definitions can be applied flexibly to ensure that the user's movement condition becomes the core element. These rules aim to calculate exercise time in a stable and reproducible manner so that the metadata item can be directly used for condition determination on the user terminal.
[0128] The video server can organize the auxiliary information used in the calculation, along with the calculated exercise time, into metadata and store it by matching it with the target video. For example, the metadata may include the total playback time, exercise time, exercise time ratio, a list of the start and end times of the exercise segment, setting values for the first and second movement amounts, and time information where the second pressure input occurred. When storing this metadata in a database, the video server can generate matching information to create a one-to-one link with the target exercise video identifier, and be configured to return the corresponding metadata when a user terminal requests a video list or queries the target metadata. As a result, the user terminal can quickly determine conditions such as the ratio of exercise time to total playback time locally, and the video server can generate highlights centered on key segments containing actual exercise while reducing requests for unnecessary highlight generation. Furthermore, by implementing the system to prioritize only videos where the second pressure input exists as targets for exercise time calculation, the server's computational load can be further reduced, allowing for the effect of performing intensive analysis on candidate videos that reflect the user's intent.
[0129] Additionally or generally, the user terminal can be configured to check the user grade corresponding to the wearable device, and if the user grade corresponds to the premium grade, transmit a highlight video request signal to the video server even if the target metadata does not satisfy a predetermined condition, and if the user grade corresponds to the general grade, check the total playback time and exercise time of the target exercise video based on the target metadata, and then determine that the target metadata satisfies a predetermined condition only if the total playback time of the target exercise video is greater than or equal to the first time and the ratio of exercise time to total playback time is greater than or equal to the threshold ratio.
[0130] A user grade is a grade assigned based on the subscription status, payment status, or service usage policy of a user account, and may include a premium grade and a general grade. A user terminal may be configured to check the user grade at at least one of the following times: when receiving a highlight generation signal or providing a video list, or when the user selects a target exercise video.
[0131] The user terminal may cache the grade information linked to the user account in local storage and retrieve it immediately, or the user terminal may send user identification information and an authentication token to a video server or a separate authentication server to retrieve the current grade and receive and verify the response.
[0132] Since the user level verification step functions not merely for display purposes but as a control criterion that branches subsequent condition judgment policies, the user terminal can be configured to check the level's validity period, policy version, or set of allowed features per level along with the level verification result, allowing behavior to vary even within the same level depending on policy changes. For example, the Premium level may more easily allow highlight creation requests and relax limits on the number of creations, while the General level may adopt a policy that strictly applies preconditions to protect server resources; furthermore, the user terminal can select different thresholds or filter rules to apply in subsequent steps based on the level verification result. Additionally, the user terminal can display the level verification result in the user interface to help users understand why highlight creation is possible or impossible for a specific video, which has the effect of enhancing service reliability.
[0133] When a target exercise video is selected, the user terminal typically determines certain conditions, including video length and video quality conditions; however, for the premium grade, it is configured to transmit a request signal to the server even if the result of the condition determination is not met, thereby enhancing the user's perceived convenience and service value.
[0134] Examples of non-compliance with conditions here may include i) cases where the total playback time is short and it is determined that the highlight generation efficiency is low for the general grade, ii) cases where the resolution or frame rate is lower than the standard value and the request is blocked due to concerns about quality degradation for the general grade, or iii) cases where the ratio of exercise time is low and it is determined that there is a lack of meaningful exercise segments for the general grade.
[0135] The meaning of transmitting a request even though the conditions are not met in the premium grade does not need to be limited to the user terminal omitting the condition judgment itself; an embodiment may also be implemented in which the condition judgment is performed but the result is not used as a blocking criterion and is transmitted to the video server so that the video server applies an alternative generation mode or correction procedure.
[0136] For example, even if the user terminal confirms that the quality parameter is below a threshold value, if it is of the premium grade, it may include a quality non-compliance flag when transmitting the request signal, and the video server may be configured to ensure result quality by performing correction procedures such as video stabilization, image quality correction, or relaxing the criteria for selecting highlight sections for the request containing the flag.
[0137] In addition, to prevent the system from causing unlimited server load even when conditions are relaxed in the Premium tier, user terminals and / or video servers may apply policies such as limits on the number of generation attempts per period, limits on concurrent processing, or delayed processing during server congestion to the Premium tier; these policies can be implemented in a manner that ensures system stability while maintaining the Premium tier experience.
[0138] “Total playback time” can be directly included in the metadata as the total playable length defined by the video file or stream, and “movement time” can be included in the metadata as the sum of the playback times of actual movement performance segments calculated by the video server through frame-by-frame analysis or movement-based segment identification.
[0139] When a general-grade user selects a target exercise video, the user terminal can first check the total playback time to determine if it is longer than a first time, and control the process so that a request for highlight generation is not performed if it is shorter than the first time. In this case, the first time can be implemented as a minimum length threshold set according to the service operation policy. For example, since performing highlight generation for a video that is too short may result in a result with low significance or waste server resources, this is intended to block it in advance.
[0140] The user terminal can check the exercise time and calculate the ratio of exercise time to total playback time to determine if it exceeds a threshold ratio; this ratio functions as an indicator to determine whether the exercise video is merely long or if it sufficiently includes actual exercise performance. For example, a video in which the user keeps the camera on and includes long preparation or rest periods may have a long total playback time but a low ratio of exercise time; therefore, in the general rating category, requests for highlight creation may be blocked, or the user may be encouraged to perform editing first. As another example, if a soccer practice video contains a long duration of actual play, the exercise time ratio is calculated to be high, making it highly likely to satisfy the threshold ratio. Only in this case can the user terminal determine that the target metadata satisfies specific conditions and be configured to send a signal requesting a highlight video to the server.
[0141] Since the user terminal directly references the total playback time and exercise time from the metadata, the computational complexity is low, and by utilizing exercise time information that has been calculated and stored in advance on the video server side, the computational burden on the user terminal can be reduced while ensuring objectivity in judgment.
[0142] Rather than setting the threshold ratio or first time value as a fixed value, the user terminal can dynamically adjust it in conjunction with network conditions, server congestion, or video quality levels. For example, a balance between system stability and user experience can be achieved by applying a stricter threshold to reduce requests when the network is unstable or server congestion is high, and relaxing the threshold for usability when the network is stable.
[0143] Pre-filtering for such general grades can be more effective when combined with the premise that highlight generation is performed on existing exercise videos stored on the server rather than on real-time video, and server resource usage can be systematically controlled by determining whether the target exercise video selected by the user is actually suitable for highlight generation at a pre-request stage.
[0144] Additionally or generally, the user terminal may be configured to check quality parameters including resolution, frame rate, and image clarity of a target motion video based on metadata, check whether a first condition regarding whether each quality parameter is greater than or equal to a set reference value is satisfied, and if the quality parameters satisfy the first condition, check whether a second condition regarding whether an image quality score calculated based on resolution, frame rate, and image clarity is greater than or equal to a threshold score is satisfied, and determine that the target metadata satisfies a predetermined condition only if the image quality score is greater than or equal to the threshold score.
[0145] A user terminal can perform a quality judgment procedure that suppresses the excessive generation of highlight requests and simultaneously ensures consistent quality of highlight results by using target metadata to determine image quality conditions in stages. For example, the user terminal is configured to first verify quality parameters such as resolution, frame rate, and image sharpness using individual thresholds, then calculate an image quality score by combining multiple parameters only for images that pass the first stage, and verify them a second time; and finally determine that the target metadata satisfies a predetermined condition only when the result is equal to or greater than the threshold score. This dual-gate structure can simultaneously secure user experience and system resource efficiency by first excluding defective images that can be quickly filtered out by the terminal and then performing a comprehensive judgment only on images in boundary areas.
[0146] “Resolution” is information regarding the frame size of an image, which can be expressed, for example, as the number of horizontal pixels and the number of vertical pixels, and the user terminal can check the resolution from the encoding header information stored in the target metadata or from the metadata field provided by the video server.
[0147] “Frame rate” refers to the number of frames per second and can be identified by the average frame rate or fixed frame rate information included in the metadata; in the case of variable frame rate video, the effective frame rate can be calculated as a correction value by sampling the actual frame interval of a certain section.
[0148] “Image clarity” is not merely a subjective expression, but can be defined as an objective indicator that can be implemented by a person skilled in the art. For example, an image server may calculate an edge intensity-based clarity indicator for sample frames in advance and store it as metadata, and a user terminal may check the value. As another embodiment, the user terminal may receive a sample frame image from a server or use a portion of a frame of a low-resolution preview stream to simply calculate the clarity indicator. In this case as well, the clarity can be quantified in the form of a distribution of high-frequency components within the frame or contour density.
[0149] The “first condition” corresponds to a condition for a step of verifying whether each parameter is greater than or equal to a reference value. For example, it may be configured such that the resolution is greater than or equal to a predetermined resolution, the frame rate is greater than or equal to a predetermined frame rate, and the clarity is greater than or equal to a predetermined clarity index. The user terminal may be configured to immediately determine that the quality condition is not met without proceeding to the second condition step if even one of these falls below the standard.
[0150] The “Video Quality Score” can be implemented by calculating a weighted combination of resolution, frame rate, and sharpness after normalizing each of them; the normalization includes processing logic to align each parameter to the same comparison scale. For example, resolution can be converted as a ratio relative to a reference resolution, the frame rate as a ratio relative to a reference frame rate, and sharpness as a ratio relative to a reference sharpness index. The weighted combination can reflect cases where importance varies depending on the sport or shooting environment. For instance, a policy can be configured to increase the weight of the frame rate in sports with vigorous movement (e.g., soccer) and increase the weight of sharpness in sports centered on static posture (e.g., Pilates). The user terminal can receive weight tables or policy values from the video server and apply them to the calculation of the Video Quality Score.
[0151] Since the “second condition” is implemented through a single threshold comparison, it places a low computational burden on the user terminal and effectively reduces quality fluctuations in the highlight generation results by further filtering out videos with borderline quality even among those that passed the first condition. For example, a video that barely meets the criteria for resolution and frame rate but has low clarity resulting in poor actual viewing quality may receive a low score in the overall score and be blocked at the request stage; conversely, a video that has relatively low parameters but maintains overall balance and ensures viewing quality may exceed the threshold in the score and allow the request.
[0152] If either the first condition or the second condition is not satisfied, the user terminal determines that the target metadata does not satisfy the predetermined conditions and may be configured to suppress the transmission of the highlight video request signal accordingly. Conversely, if both the first and second conditions are satisfied, the user terminal determines that the target metadata satisfies the predetermined conditions and proceeds with the highlight request. In this case, the final decision made by the user terminal is not limited to merely setting an internal flag, but may be used as a trigger to directly control subsequent operations, such as guidance text on the user interface, whether the request button is activated, or whether quality grade information included in the request signal transmitted to the video server is added.
[0153] For example, if the quality conditions are not met, the user terminal may disable the highlight generation button or display a notice that highlight generation is restricted due to low quality, and if the quality conditions are met, it may proceed with the request immediately or provide the user with a highlight generation progress screen.
[0154] When quality conditions are repeatedly not met, the user terminal can provide guidance to the user by estimating causes such as insufficient illumination, potential shaking, or potential frame rate degradation to induce improvement of the shooting environment; since such guidance can be provided based on substandard items among the quality parameters recorded in the metadata, feasibility of implementation and clarity of explanation are ensured.
[0155] Consequently, the above dual judgment structure functions as a quality control mechanism during the input stage of highlight generation, providing technical effects that improve the average quality of highlight results provided to users and reduce unnecessary server-side processing.
[0157] FIG. 4 is a flowchart of a highlight image providing system according to one embodiment of the present invention.
[0158] According to one embodiment, the components of a highlight video providing system may perform the operations disclosed in FIG. 4. For example, at least some of the components included in the highlight video providing system (e.g., the user terminal (100), wearable device (210), and video server (220) of FIG. 2 may be configured to perform the operations of FIG. 4.
[0159] In the following embodiments, the operations S410 to S430 may be performed sequentially, but are not necessarily performed sequentially. For example, the order of each operation may be changed, and at least two operations may be performed in parallel. Additionally, content corresponding to or overlapping with the above description in relation to FIG. 4 may be briefly explained or omitted.
[0160] In particular, the actions according to FIG. 4 can be understood as actions performed by the video server (220) of FIG. 2. FIG. 4 relates to an embodiment in which the video server performs frame-by-frame analysis of movement amount for a selected target exercise video, selects frames in which the excess amount relative to the average movement amount of the entire video is above a certain standard, identifies highlight sections based on the selected frames, determines the standard value differently according to the exercise type set by the user, and reflects the relative size relationship between the two standard values when the exercise type is soccer or Pilates. The following description is based on a scenario performed by the video server when the length and quality conditions are passed as metadata.
[0161] According to one embodiment, the video server (220) can divide the target exercise video into multiple frames and then check the first average movement amount of the user and the second average movement amount of the exercise tool during the entire playback time of the target exercise video, and the first partial movement amount of the user and the second partial movement amount of the exercise tool (or, exercise object) for each of the multiple frames (S410).
[0162] The video server can load the target exercise video from the storage and sequentially obtain frames in chronological order through decoding, and frame splitting can be performed adaptively according to the frame rate per second.
[0163] The user’s “first movement amount” may be defined as the amount of change in position of the user’s body or part of the body, and the “second movement amount” of the exercise tool may be defined as the amount of change in position of the tool corresponding to the type of exercise, such as a ball, racket, club, band, dumbbell, etc.
[0164] The video server estimates the positions of the user and the exercise tool in each frame in the form of coordinates or areas, and can calculate the movement amount per frame using the change in position between consecutive frames. Subsequently, the video server can calculate a representative value of the user's movement amounts calculated over the entire video segment as the first average movement amount, and a representative value of the exercise tool's movement amounts as the second average movement amount, and the average can be implemented as an arithmetic mean.
[0165] The video server may define and store the amount of movement in a corresponding section for each frame or a certain group of frames as a first partial movement amount and a second partial movement amount, or use them for immediate comparison operations. In this case, the partial movement amount may be defined as the amount of movement in a single frame, or it may be defined as the accumulated amount of movement over a short window period to mitigate frame rate fluctuations or sudden spikes. The result data resulting from this video processing operation logic is used as basic data to calculate the excess amount relative to the average in a subsequent step.
[0166] According to one embodiment, a video server may be configured to calculate a first partial movement amount by estimating the coordinates of a user's body joints or key points in each of multiple frames and accumulating or averaging the coordinate change amounts between consecutive frames, in order to calculate a movement amount corresponding to a user's first movement included in a target exercise video. The video server may calculate a scalar value representing the user's center of gravity movement amount or posture change amount by using the movement amount for at least one of the multiple joint coordinates or by combining the movement amounts of multiple joints. Additionally, the video server may be configured to calculate the movement amount using only coordinates where the reliability value per joint is greater than or equal to a threshold value in order to exclude or correct joint estimation results with low reliability. Accordingly, the user's first movement is quantified as a change in actual body motion rather than a simple movement of an object within the screen, thereby improving the precision of exercise time calculation and highlight section identification.
[0167] According to one embodiment, the video server may be configured to identify a contour area or bounding box corresponding to the user in each frame to calculate the first movement of the user, and to define the amount of change between frames of the center coordinates or bottom ground point coordinates of the bounding box as the first partial movement amount. Since a change in the size of the bounding box may include a change in distance from the camera, the video server may be configured to consider the amount of change in the bounding box area or aspect ratio as an auxiliary term to reduce misjudgment of the movement amount. Additionally, the video server may be configured to perform inter-frame associative tracking to maintain identification of the same user, and to suspend the calculation of the movement amount or interpolate with the movement amount of an adjacent section in sections where tracking is interrupted. This method has a low computational burden, allowing for rapid application even in large-scale video processing environments, and can be utilized as a basic indicator for calculating motion time.
[0168] According to one embodiment, a video server may be configured to detect a feature point or marker corresponding to the exercise tool in each frame and calculate the amount of change in the coordinates of the feature point between consecutive frames as the second partial movement amount in order to calculate the amount of movement corresponding to the second movement of the exercise tool. If the exercise tool is a ball, the coordinates of the center of the ball may be used; if it is a racket or club, the coordinates of the feature point of the head portion may be used; and if it is a device such as a band or ring, the amount of movement of a specific corner of the device or the center of the device may be used. Considering cases where the exercise tool is obscured or detection reliability is low, the video server may be configured to select the feature point with the highest reliability among multiple candidate feature points, or to combine the amounts of movement of multiple feature points to define a single second partial movement amount. Accordingly, the second movement of the exercise tool is stably calculated as a separate signal distinct from the user's movement, thereby improving the accuracy of exercise time and highlight candidate frame selection.
[0169] According to one embodiment, the video server may be configured to calculate light flow information between consecutive frames and calculate the amount of movement based on the magnitude distribution of the light flow vectors in order to calculate the first movement of the user and the second movement of the motion tool. Since the light flow of the entire screen may include camera shake, the video server may be configured to estimate and remove the global light flow component of the background area, and then calculate the first partial amount of movement or the second partial amount of movement using only the local light flow component corresponding to the user area or the motion tool area. Additionally, the video server may define the amount of movement as the ratio of pixels whose light flow vector magnitude is above a certain threshold or the average value of the light flow magnitude, and such an amount of movement can be calculated relatively robustly even in low-resolution video where object detection is difficult. Consequently, the light flow-based amount of movement can act as a supplementary means in environments where object estimation is unstable, thereby improving the stability of motion time calculation and highlight section identification.
[0170] Since the optical flow method can increase server computational load, it becomes more practical to include in the specification that the sampling frame interval can be increased or that the optical flow can be calculated only within the region of interest.
[0171] According to one embodiment, the video server may be configured to perform global alignment between frames to estimate global transformation parameters and to use residual motion components, from which motion components described by the global transformation have been removed, in order to prevent global motion caused by camera shake or movement of the photographer from affecting the calculation of the first and second partial motion amounts. The video server may use background feature points as a standard for global alignment, and after estimating the global transformation from the motion vector of the background feature points, it may calculate a corrected motion amount by subtracting the component corresponding to the global transformation from the coordinate movement of the user or exercise tool. Accordingly, the problem of camera movement unrelated to actual exercise performance being overcalculated as motion amount can be reduced, and the calculation of exercise time can be configured to more accurately reflect the sum of the playback times of the actual exercise segments.
[0172] According to one embodiment, the video server (220) can identify a frame as a target frame in which the value obtained by subtracting the first average movement amount from the first partial movement amount is greater than or equal to a positive first value and the value obtained by subtracting the second average movement amount from the second partial movement amount is greater than or equal to a positive second value (S420).
[0173] The video server checks the average movement amount calculated in operation S410 as a baseline and can calculate the subtraction value indicating how much more dynamic movement occurred in each frame compared to the average.
[0174] A positive “subtraction value” means that the frame in question shows a greater movement than the overall average, and the condition of being greater than or equal to a positive first value or a positive second value can be interpreted to mean that only frames where the excess amount relative to the average exceeds a specific threshold will be selected.
[0175] For example, in a soccer scene, even if the user's movement increases only slightly above the overall average, it can become a meaningful situation (especially the movement of the foot, which is the target body corresponding to soccer among the user's body parts), but since the moment when the movement of the ball increases suddenly and significantly can be a decisive scene, if the frame satisfying both the user movement excess and the ball movement excess is targeted, the likelihood of selecting the frame where the scene transition or decisive play actually occurred increases.
[0176] For example, in a Pilates setting, exercise equipment may be unavailable or movement may be restricted, so conditions for excess movement of exercise equipment can be defined to match the movement of bands or equipment, or an operational policy can be established to apply them strictly only in disciplines where equipment is present.
[0177] Since the video server may have lower accuracy due to noise when only a single frame satisfies the condition, in the step of checking the target frame, it may check whether a certain number of consecutive frames satisfy the condition together, or perform post-processing to cluster the frames that satisfy the condition on the time axis into meaningful chunks.
[0178] Since this target frame verification procedure uses a relative standard compared to the average, it operates relatively stably even if there are differences in shooting environments or individual users, and offers higher universality and reproducibility compared to methods based on absolute thresholds.
[0179] According to one embodiment, the video server (220) can identify at least one highlight image based on the target frame (S430).
[0180] The video server can generate highlight candidate segments based on the time index at which the target frame occurred, and the candidate segments can be set to a time range of a certain length that includes the target frame.
[0181] For example, if target frames are concentrated at a specific time, the video server can find the start and end times of that distribution and construct a single continuous segment by including a small buffer time before and after that range.
[0182] For example, if the target frame exists divided into multiple chunks, candidate segments are created for each chunk, and among them, the segment with a high motion time ratio or the segment with high proximity to the second pressure input time can be selected as a priority to identify the final highlight segment.
[0183] If the identified highlight segment is longer than 15 seconds, the video server extracts only the representative 15 seconds, and if it is shorter than 15 seconds, it can extend the segment to include the target frame distribution to match 15 seconds.
[0184] The identification of highlight footage is not limited to simply determining time intervals; it can be configured to segment the corresponding section into an encodeable format, align frame boundaries, and maintain audio-video synchronization. During this process, the video server saves information such as generation rules, threshold values used, and the start and end times of the selected section as metadata, which can then be utilized for configuring playback screens or for additional editing functions.
[0185] Consequently, the structure for identifying highlight videos based on target frames generates highlight videos centered on sections where the user's actual movement and the movement of the exercise tool have changed significantly compared to the average, thereby increasing the likelihood of including scenes that are perceptually dynamic and meaningful.
[0186] Meanwhile, the exercise type can be entered by the user when uploading a video from a user terminal, by specifying a sport tag for the video from a video list, or by selecting it when creating a recording session, and the corresponding exercise type can be stored on a video server as metadata for the target exercise video.
[0187] The video server can identify the type of exercise from metadata obtained upon a highlight request signal or a target video query, and refer to a policy table to determine the first and second values corresponding to the exercise type. For example, the policy table may reflect empirical settings regarding which signal—user movement or exercise tool movement—better represents the highlight for each exercise type. As an example, different threshold values can be pre-stored for each sport, such as soccer, tennis, golf, and Pilates. Furthermore, the policy table may be configured to be adjusted based on the user's proficiency or the shooting environment, rather than containing only fixed values. Since the distribution of average movement may differ between children's practice videos and adult game videos even for the same type of exercise, the video server can fine-tune the first and second values by reflecting the average movement statistics of the most recent videos.
[0188] In contrast to the conventional method of applying a single threshold uniformly to all events, this exercise type-based threshold determination improves the accuracy of highlight selection reflecting event-specific characteristics and provides the effect of reducing unnecessary false positives or misses.
[0189] For example, according to the policy table, it can be configured so that if the exercise type is soccer, the first value is set smaller than the second value, and if the exercise type is Pilates, the first value is set larger than the second value. This relative size relationship reflects the fact that the scene characteristics of soccer and Pilates are different.
[0190] For example, in soccer, the rapid movement of the ball, which is a sporting tool, often strongly indicates a decisive moment, so the second value may be set larger in the direction of applying the condition of excess amount relative to the average of the ball relatively strictly; conversely, since the condition of excess amount of user movement can be a meaningful scene even if it does not occur as significantly as the movement of the ball, the first value may be set smaller than the second value. In this structure, since the target frame must simultaneously satisfy the excess amount conditions of both the user and the ball, there is a higher probability that a scene in which the user's movement is accompanied by a moment in which the ball moves noticeably dynamically will be selected as the target frame.
[0191] For example, in Pilates, changes in the user's posture or shifts in the center of gravity often determine the quality of the exercise and the meaning of the scene. Since the movement of equipment may be relatively limited or subtle even when it is present, the first value may be set to be greater than the second value in order to apply the condition for excess user movement relatively strictly. In this case, the definition of exercise equipment can be extended to include Pilates rings, bands, reformer components, etc. Furthermore, since the second average movement amount itself is calculated to be low in shootings where there is almost no movement of the equipment, which may alter the distribution of the subtraction value, the video server may implement an operational policy where the second value plays a supplementary role depending on the exercise type, or actively applies it only in Pilates types where equipment is present.
[0192] This relative size setting is not a simple comparison of values, but a technical design that reflects which signals represent highlights for each stock, serving as a concrete means to implement stock optimization within the same excess-to-average method.
[0194] The user terminal is configured to calculate the image quality score based on the following mathematical formula 1, and
[0195]
[0196] SQ is the image quality score, K is a positive constant for determining the scale of the score, a, b, and c are positive exponential coefficients for determining the sensitivity of each term, p is a positive coefficient for determining the intensity of the blur penalty, and m is a small positive constant for ensuring numerical stability of the logarithmic calculation when B is close to 0,
[0197] R is a value corresponding to resolution, F is a value corresponding to frame rate, S is a value corresponding to image sharpness index, R_0, F_0, and S_0 are normalized reference values corresponding to reference resolution, reference frame rate, and reference sharpness, respectively, B is a blur index representing the degree of motion blur or blurring, B0 is a reference blur index, and J is a jitter index representing shaking or stabilization quality, corresponding to a value such as the variability of global movement between frames or stabilization residual.
[0198] R is a value corresponding to resolution and can typically be defined as the total number of pixels in a frame. For example, it is defined as the product of the horizontal and vertical pixels and can be easily extracted from the encoding header or container metadata. R0 is a normalization reference value corresponding to the reference resolution and can be set, for example, to the number of pixels corresponding to the resolution recognized by the service as minimum quality. Since R / R0 represents the relative resolution relative to the reference, it enables the comparison of images from different shooting devices or encoding conditions at the same scale.
[0199] According to Equation 1, the image quality score increases as the resolution increases, but the perceived quality can be reflected at a resolution above a certain level. Excessive preferential treatment of super-high resolution images, which is prone to occur in simple linear weighted sums, is suppressed, and the separation of boundary cases in threshold score-based judgment is stabilized.
[0200] R0 may not be fixed as a single value and may vary, for example, depending on the user class.
[0201] F corresponds to the frame rate and can be defined as the number of frames per second. If it is a fixed frame rate, it can be checked directly in the metadata; if it is a variable frame rate, the effective frame rate can be calculated using the frame timestamp interval of the sample period. F0 is the reference frame rate and can be set, for example, as a minimum frame rate recognized by the service or as a reference value linked to the target frame rate of the highlight results.
[0202] According to mathematical formula 1, as the frame rate increases, motion rendering improves and the score increases; however, beyond a certain level, the perceived improvement becomes gradual. Particularly in sports videos, when the frame rate is low, highlight satisfaction drops sharply due to ghosting and stuttering; this term helps ensure that low-frame-rate videos naturally receive low scores and are excluded from server processing.
[0203] For example, even if F is high in the metadata, the perceived quality may be lower if the actual shutter speed or blur is severe, so overestimation can be prevented if it is implemented in a form combined with the blur penalty term of Equation 1.
[0204] S is an image clarity metric that is not merely a subjective concept but can be defined as a numerical value representing the spatial high-frequency components or edge components of a frame. It can be implemented by calculating an edge intensity-based metric for representative selected sample frames or by calculating a scalar value similar in form to the variance of a Laplacian response and averaging it. The video server can calculate S immediately after upload or during the pre-processing stage prior to a request and store it as metadata, and the user terminal can be configured to query this value to use for quality judgment. S0 is a reference clarity metric and can be set as a reference value signifying minimum readability or minimum detail reproduction.
[0205] According to mathematical formula 1, the score increases as sharpness increases, but saturation characteristics are applied so that the score does not increase indefinitely even in situations where the sharpness index increases due to excessive sharpening or increased noise. As a result, the relative evaluation of high-resolution but blurry images and medium-resolution but sharp images is aligned closer to perceived quality.
[0206] When calculating S, the user terminal and / or video server may calculate it only from a specific set of sample frames instead of processing all frames.
[0207] B can be defined as a blur index representing the degree of motion blur or blurring. B can be designed to have a complementary relationship with the sharpness index S. For example, it can be calculated using the characteristic that edge width widens or high-frequency components decrease as the blur increases. B0 is a reference blur index and can be set as a reference value representing an acceptable level of blur. M is a small positive constant to prevent numerical instability in logarithmic calculations when B is very small or close to zero, and is an implementation parameter for the stable operation of the system.
[0208] According to Equation 1, the score can be attenuated non-linearly as the blur moves further away from the reference value. In particular, by using the square of the ln ratio, a shape is created that is gradual near the reference and decreases sharply when deviating significantly from it. This is technically significant as it reflects the tendency for the impact of blur on the perceived quality of the highlight result to deteriorate sharply after a critical point, rather than being linear.
[0209] If B is treated as an independent metric, the problem of heavily blurred images receiving high scores even when F and R are high can be reduced, allowing the setting of the second condition threshold score to be implemented intuitively.
[0210] J can be defined as a jitter metric representing shaking or stabilization quality. A video server can calculate J using values such as the variability of the global motion vector between frames, frame alignment error due to camera shake, or residual motion after stabilization. For example, if the variance of the change in the global motion of the background is defined as jitter after estimating it for each frame, it can be implemented in a form where J increases as camera shake becomes more severe. The 1 / (1+J) structure creates the effect of the score decreasing directly in the denominator as J increases, helping to exclude videos with significant shaking from being highlighted.
[0211] Even if the resolution and frame rate are good, videos with significant shaking can have a significantly lower perceived quality, and in some cases, the highlight results may become unwatchable. Therefore, Equation 1 ensures that such videos are automatically disadvantaged in the quality score, thereby preventing the waste of server resources.
[0212] K is a positive constant used to determine the scale of the score and can be used to facilitate comparison with the threshold score. While K itself is not essential for ranking determination, it has the advantage of helping to intuitively design threshold score policies in service operations.
[0213] a, b, and c are exponential coefficients that adjust the sensitivity of the resolution, frame rate, and sharpness terms, respectively, controlling the extent to which the same input change affects the score. For example, if you want to increase the importance of the frame rate, you can set b to a relatively large value. p is a coefficient that controls the intensity of the blur penalty, controlling the abruptness of the score drop when the blur deviates from the threshold value.
[0214] These coefficients can be stored on the server as service policy values or managed as a set of parameters selected based on the type of exercise or the shooting environment. For example, for sports with many fast movements like soccer, the impact of frame rate and blur is significant, so b and p can be set relatively high; conversely, for sports with many static poses like Pilates, the impact of sharpness is significant, so c can be set relatively high. This approach naturally combines with exercise type-based threshold settings, allowing quality judgments to adapt to the characteristics of the sport.
[0215] Mathematical Equation 1 is configured to simultaneously reflect the saturation characteristics of quality factors and the non-linear penalty of quality degradation factors. The ln term prevents the score from skyrocketing by reflecting the property that perceived improvement slows down once enhancement factors such as resolution, frame rate, and sharpness exceed a certain level. The exp penalty term drastically lowers the score as blur deviates from the standard, effectively excluding severely blurred images that are difficult to capture with simple linear penalties. The 1 / (1+J) term directly suppresses the problem where images with significant shaking consume system resources and degrade highlight quality.
[0216] Furthermore, by using a multiplication structure, the overall score is limited even if only one element is excellent, provided that other elements are lacking. This reflects the technical requirement that multiple quality standards must simultaneously satisfy a certain level for highlight videos to be provided at a quality that is actually watchable. Consequently, the second condition filters out defective videos much more reliably than a simple weighted sum and reduces processing latency and traffic by allowing the server to concentrate resources on videos worthy of generating highlights.
[0218] The above-described highlight video providing system is configured such that the user terminal monitors the playback status of the highlight video to check the viewing duration, and if it is confirmed that the viewing duration has exceeded a threshold time, virtual points proportional to the total playback length of the target exercise video are awarded to the user's account information based on a set winning probability.
[0219] This configuration contributes to implementing sophisticated resource management logic that goes beyond simple reward payments to precisely verify users' actual viewing behavior and reflect the value of original data in the reward system.
[0220] The user terminal monitors the playback status of highlight videos in real time through an embedded media engine to calculate the valid viewing duration. Rather than simply granting a reward based on the act of pressing the video play button, the system comprehensively tracks playback, pause, skipping, and background switching of the app to verify whether the user has actually watched the video. The system records these events to calculate the cumulative time the video is displayed on the screen and determines whether the viewing duration exceeds a pre-set threshold (e.g., 10 seconds or more for a 15-second video). This functions as a verification procedure that prevents abuse of the reward system and technically guarantees the validity of system traffic.
[0221] When valid viewing is confirmed, the system determines whether to award a reward based on the set winning probability, but dynamically determines the amount of points awarded in conjunction with the total playback length of the target exercise video. For example, even for the same 15-second highlight, the system can be designed to award higher virtual points when watching a highlight extracted from a 1-hour full match video compared to one extracted from a 10-minute original video. This significantly contributes to reflecting in the reward the technical value that the larger the source data for the highlight, the higher the user's exercise activity and data generation contribution. Consequently, this design encourages users to generate exercise data for longer periods and provides the effect of optimizing system activity by organically combining metadata from the generated original video (e.g., total playback length) with consumption data from the highlight (e.g., completion status).
[0223] The above-described highlight video providing system is configured such that the video server calculates the total sum of target virtual assets consumed as a third user accessing the highlight video via a shared link uses a paid function, and distributes reward virtual assets corresponding to a predetermined ratio of the total sum to the user's account information.
[0224] In other words, the video server assigns a unique tracking ID to the shared URL generated for each highlight video and tracks revenue data in real time when other viewers who click the link utilize paid services such as high-quality downloads, AI photo extraction, or membership subscriptions. Subsequently, the video server distributes rewards equivalent to the corresponding amount to the video sharers (i.e., users) according to a predefined distribution logic (e.g., 5% of the viewer's spending) and provides the distribution results through a dashboard.
[0225] This possesses technical originality as a network-based asset settlement logic that automatically distributes revenue based on contribution to content distribution, rather than simply awarding points. In particular, since it serves as a powerful incentive that encourages users to voluntarily promote the highlights they create, its effectiveness in terms of system operational efficiency is greatly maximized.
[0227] The above highlight video providing system may be configured such that the video server sets a video time point corresponding to the occurrence time of the second pressure input as a reference time point, searches for a peak section in which the time change pattern of the user's first partial movement amount and the exercise tool's second partial movement amount has maximum correlation within a predetermined time range before and after the reference time point, determines the start time point and end time point of the highlight section to include the peak section, and generates a highlight video by normalizing the determined highlight section to a length of 15 seconds.
[0228] In other words, the video server identifies the second pressure input not merely as a signal triggering a highlight, but as a reference point that strongly restricts the search range along the video time axis. Since the video server analyzes only short candidate regions around the reference point, the computational load is significantly reduced compared to methods that analyze the entire video in batches. Nevertheless, rather than simply cropping the area around the reference point, this approach identifies the peak where the temporal variation patterns of user movement and tool movement interlock most strongly; consequently, it can generate highlights that more accurately capture the moment when the actual impact occurs.
[0229] For example, the video server can divide candidate segments into sliding windows, calculate the correlation value or degree of simultaneous rise between the user movement sequence and the tool movement sequence in each window, and determine the highlight segment centered on the window with the largest value. Subsequently, in the 15-second normalization step, frame boundary alignment can be performed while maintaining the peak center to adjust the start and end times to be exactly 15 seconds. Since this structure combines the reflection of user intent based on a second pressure input with objective event detection based on movement volume, it can strongly ensure consistency in highlight quality.
[0231] The highlight video providing system described above may be configured such that the video server identifies a cluster of motion videos including two or more videos determined to be the same motion session among at least one motion video, selects the highest quality section by comparing the video quality scores of the two or more motion videos with respect to a reference section based on the time of occurrence of the second pressure input, and synthesizes the highest quality section into a single 15-second highlight video by time-axis alignment.
[0232] In other words, the video server performs the aforementioned operation in situations where multiple videos exist for the same session, such as multi-angle shooting, recording of amateur club matches, or simultaneous recording of a coach and a player. Conventionally, highlights are generated from only one video, which can result in unstable outcomes if the quality is low or the subject is obscured; however, this embodiment constructs highlights by selecting only the highest-quality segments from among multiple videos containing the same event. Consequently, the quality of the highlights is not dependent on shooting luck and can be provided at consistently high quality.
[0233] For example, the video server can reduce false positives by using the proximity of the shooting time, the same location tag, the same user or same team tag, and the proximity of the time of the second pressurization input together for session determination. The simplest method for time-axis alignment is based on the input time, and if necessary, fine alignment can be performed using audio features or screen-wide motion patterns. Final synthesis can be implemented not only by taking the entire 15 seconds from a single camera, but also by replacing only the relevant sections with different video if occlusion or shaking occurs in the middle, thereby significantly improving the actual user-perceived quality.
[0235] The highlight video providing system is configured such that the video server identifies a highlight section and generates explanatory metadata indicating the basis for the identification of the highlight section, and provides this metadata to the user terminal, wherein the explanatory metadata is configured to include at least one of the time of occurrence of the second pressure input, the time distribution of the target frame, the magnitude of the excess amount of movement, and the result of applying the motion type-based threshold.
[0236] This embodiment is a structure that provides users with an explanatory form of why a highlight was selected. Instead of simply discarding the result when a user does not like the highlight, they can verify which criteria were applied, making it easier for the system to gain user trust. Furthermore, from the server's perspective, explanatory metadata directly aids in quality improvement, debugging, and the collection of user feedback.
[0237] The explanatory metadata may be provided as a numerical summary or as segment-based summary values to enable visualization on the user terminal. For example, it may provide information such as how much the selected segment has shifted forward or backward based on the second pressurization input time, at what moment the excess amount of movement was greatest, and how threshold relationships were applied in soccer or Pilates types. This approach contributes to highlight generation technology functioning as a verifiable procedure rather than a simple output.
[0239] The above description is merely an illustrative explanation of the technical concept of the present invention, and those skilled in the art to which the present invention pertains will be able to make various modifications and variations within the scope of the essential characteristics of the present invention.
[0240] Accordingly, the embodiments disclosed in this invention are intended to illustrate, not limit, the technical concept of the invention, and the scope of the technical concept of the invention is not limited by these embodiments. The scope of protection of this invention shall be interpreted by the claims below, and all technical concepts within an equivalent scope shall be interpreted as being included within the scope of rights of this invention.
Claims
Claim 1 A highlight video providing system comprising a wearable device, a user terminal, and a video server, wherein the system generates at least one highlight video through an exercise video based on interaction between the devices and provides it to a user, wherein the wearable device transmits a highlight video generation signal to the user terminal in response to receiving a predetermined first pressure input from the user, and the user terminal provides the user with a video list including a thumbnail of at least one exercise video received from the video server in response to receiving the highlight video generation signal, and when the user terminal receives a selection input for a target exercise video from the user, determines whether target metadata corresponding to the target exercise video satisfies a predetermined condition including a video length condition and a video quality condition, and if the target exercise video satisfies the predetermined condition, the user terminal transmits a highlight video request signal including the target exercise video to the video server, and the video server generates at least one highlight video corresponding to 15 seconds based on the target exercise video and provides it to the user terminal, and wherein the highlight video providing system is configured such that the video server, among a plurality of exercise videos registered corresponding to the user, receives the first pressure from the wearable device during shooting A highlight video providing system configured to identify an image in which a second pressure input distinct from the input is acquired as at least one exercise video, and for the at least one exercise video, the image server identifies the sum of the playback times of the intervals in which the user's first movement and the second movement of the exercise tool are each greater than or equal to the first movement amount and the second movement amount, respectively, as the exercise time, and then match and store metadata including the identification result to the at least one exercise video. Claim 2 delete Claim 3 In claim 1, the highlight video providing system is configured such that the user terminal checks the user grade corresponding to the wearable device, and if the user terminal's user grade corresponds to the premium grade, transmits the highlight video request signal to the video server even if the target metadata does not satisfy the predetermined conditions, and if the user terminal's user grade corresponds to the general grade, checks the total playback time and exercise time of the target exercise video based on the metadata, and determines that the target metadata satisfies the predetermined conditions only if the total playback time of the target exercise video is greater than or equal to the first time and the ratio of the exercise time to the total playback time is greater than or equal to the threshold ratio. Claim 4 In paragraph 3, the highlight video providing system is configured such that the user terminal checks quality parameters including resolution, frame rate, and image clarity of the target motion video based on the metadata, and then checks whether a first condition regarding whether each of the quality parameters is greater than or equal to a set reference value is satisfied; if the quality parameters satisfy the first condition, the user terminal checks whether a second condition regarding whether an image quality score calculated based on the resolution, frame rate, and image clarity is greater than or equal to a threshold score is satisfied; and the user terminal determines that the target metadata satisfies the predetermined condition only when the image quality score is greater than or equal to the threshold score. Claim 5 In claim 4, the highlight video providing system is configured such that the video server divides the target exercise video into a plurality of frames, checks the first average movement amount of the user and the second average movement amount of the exercise tool during the total playback time of the target exercise video, and the first partial movement amount of the user and the second partial movement amount of the exercise tool for each of the plurality of frames, the video server identifies a frame among the plurality of frames in which the value obtained by subtracting the first average movement amount from the first partial movement amount is greater than or equal to a positive first value, and the value obtained by subtracting the second average movement amount from the second partial movement amount is greater than or equal to a positive second value as a target frame, and the video server is configured to identify at least one highlight video based on the target frame, the video server determines the first value and the second value based on the exercise type set by the user corresponding to the target exercise video, and the video server is configured to set the first value smaller than the second value if the exercise type corresponds to soccer, and to set the first value larger than the second value if the exercise type corresponds to Pilates.
Citation Information
Patent Citations
Exercise system through an image recognition-based experiential simulator robot
KR102626383B1
System for LED screens that includes function to return super resolution image data using a learning model by chracteristics of image data
KR102820163B1
Device and method for generating sports highlights by linking a wearable device
KR102825358B1