System

The system efficiently identifies and prioritizes interesting or important video segments by clustering, scoring, and generating playlists, allowing users to watch videos flexibly and within time constraints.

JP2026035239APending Publication Date: 2026-03-04SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024138082
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-08-19
Publication Date
2026-03-04

AI Technical Summary

Technical Problem

Conventional video watching methods are inefficient, requiring manual skipping or fast-forwarding to find interesting or important parts, and often result in missing critical information due to the length of videos.

Method used

A system that uploads video data to a server, divides it into frames, clusters scenes, evaluates and scores their interest and importance, generates a playlist based on user settings, and allows real-time updates to prioritize and play only the important or interesting parts.

Benefits of technology

Enables users to efficiently watch only the important or interesting parts of videos within a limited time, saving time and ensuring no critical information is missed.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026035239000001_ABST
    Figure 2026035239000001_ABST
Patent Text Reader

Abstract

A system is provided.SOLUTION: A system including means for uploading moving image data to a server, means for dividing the moving image data into frames and clustering the moving image data for each scene, means for evaluating interest and importance of each scene and scoring the moving image data, means for prioritizing the scenes based on the evaluated scores, means for generating a reproduction list based on time-shortening reproduction setting of a user and reproducing the moving image, and means for updating the reproduction list according to an interaction of the user.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The technology of the present disclosure relates to a system. [Background technology]

[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]

[0004] When watching videos, users often want to efficiently watch only the interesting or important parts. However, with conventional methods, it is not clear which parts to skip, and fast-forwarding and using the seek bar take time, making it difficult to watch efficiently. In addition, since many videos are long, it is not possible to watch the entire video, and there is a risk of missing important information. [Means for solving the problem]

[0005] The present invention provides a means for uploading video data to a server, dividing the video data into frames, and clustering them by scene. Furthermore, by using a means for evaluating and scoring the interest and importance of each scene, it is possible to prioritize each scene. The present invention also includes a means for generating a playlist based on the user's time-saving playback settings and playing the video, and a means for the user to select specific scenes and update the playlist while watching. This allows the user to efficiently watch only the important or interesting parts, thereby saving time.

[0006] "Video Data" means visual and audio information stored in the form of a video file.

[0007] A "server" is a computer system for storing, processing, and distributing video data over a network.

[0008] A "frame" is an individual still image that makes up video data, and when combined together they form a video.

[0009] "Clustering" is the process of grouping data with similar characteristics, in this case, similar scenes.

[0010] A "scene" refers to a specific continuous time span within a video that forms a story or section.

[0011] "Fun" refers to the elements that attract and entertain users, and is a criterion for evaluating entertainment value.

[0012] "Importance" refers to elements that are highly valuable to users and are important as knowledge or information.

[0013] "Scoring" is the process of assigning evaluation points based on specific attributes of an object.

[0014] "Prioritization" is the process of comparing the importance and interest of each item based on the evaluation results and determining the order.

[0015] A "playlist" is a list that sets the order in which videos are played, and is generated based on the user's viewing preferences.

[0016] The "time-saving playback setting" is a setting that allows users to select and play only the important or interesting parts of a video in order to watch it efficiently.

[0017] "User interaction" refers to the operations or instructions that a user gives to the system, including updating a playlist and selecting a scene. [Brief explanation of the drawings]

[0018] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10]1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION

[0019] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.

[0020] First, the terms used in the following description will be explained.

[0021] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).

[0022] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.

[0023] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.

[0024] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.

[0025] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."

[0026] [First embodiment]

[0027] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.

[0028] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0029] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0030] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.

[0031] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0032] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0033] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.

[0034] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0035] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0036] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0037] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0038] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0039] This invention is a system that allows users to effectively shorten the time spent watching videos and efficiently watch only the interesting or important parts. This system consists of three main components: a server, a terminal, and a user.

[0040] Server Operation

[0041] The server receives video data uploaded by users. After receiving the video data, the server divides the video into frames and clusters each frame into scenes. The clustered scenes are evaluated through audio data analysis, text data analysis, and visual data analysis, and each scene is scored for its interest and importance. The scored scenes are prioritized based on the evaluation results and stored in a database on the server.

[0042] Device behavior

[0043] When a user wants to watch a video, they first launch the application on their device and select a video. Users can set viewing preferences for time-saving playback, such as "play only the important parts for 50%" or "play only the interesting parts for 30%." These preferences are sent from the device to the server, which then obtains scoring information for prioritized scenes. The device generates a playlist based on this information and plays the video according to the preferences. If the user selects a specific scene or wants to view a skipped scene while watching, the playlist is updated in real time. The server reflects the user's actions, and a regenerated playlist is sent to the device.

[0044] User behavior

[0045] When a user wants to watch a video, they first upload the video file to the server. Next, they launch the application on their device and select the video they want to watch. The user then configures time-saving playback settings to suit their convenience and interests. Once playback settings are configured, the device automatically obtains scene scoring information from the server and generates a playlist. If the user wants to check a scene that was skipped during viewing or cancel time-saving playback, they can press a button to update the playlist in real time, allowing them to view content according to their needs.

[0046] Specific examples

[0047] For example, consider a case where a user wants to watch a 30-minute talk show but only has 50% of the time. The user sets the device's application to "play only 50% of the important parts." When the video is uploaded to the server, the server divides the video into frames and uses an AI model to score the importance of each scene. Based on the scoring results, the server prioritizes the scenes and sends them back to the device. The device then lists only the important scenes based on the settings and plays them back so that the user can watch efficiently. If a scene of interest is skipped while watching, the user can press a dedicated button to jump to that scene. This system allows users to watch within a limited time without missing any important information.

[0048] The processing flow will be explained below.

[0049] Server Processing

[0050] Step 1:

[0051] The server receives the video data from the user and stores it in temporary storage.

[0052] Step 2:

[0053] The server divides the video data into frames, i.e., converts the video into multiple still images.

[0054] Step 3:

[0055] The server clusters frames by scene, which groups frames with similar characteristics together.

[0056] Step 4:

[0057] To evaluate each scene, the server analyzes audio data, text data, and visual data. For example, it analyzes the speaker's content from audio data and subtitles and descriptions from text data.

[0058] Step 5:

[0059] The server scores each scene based on its interest and importance, and assigns a score to each scene based on the evaluation results.

[0060] Step 6:

[0061] The server prioritizes scenes based on the evaluated scores, and the priority information and scoring results are stored in a database.

[0062] Terminal handling

[0063] Step 1:

[0064] The device receives the user's time-saving playback settings, such as "play only the important parts at 50%."

[0065] Step 2:

[0066] The device retrieves scene scoring information from the server, including scene priority information.

[0067] Step 3:

[0068] The device will generate a playlist based on the user's settings, prioritizing important or interesting scenes.

[0069] Step 4:

[0070] The device plays videos based on the playlist, and unnecessary scenes are skipped as the user watches.

[0071] Step 5:

[0072] If a user wants to view a scene that was skipped during viewing, the device will accept a dedicated button and will readjust the playlist based on that information.

[0073] Step 6:

[0074] The device sends the updated playlist to the server and continues playing the video based on the regenerated playlist.

[0075] User Action

[0076] Step 1:

[0077] The user selects a video and uploads it to the server.

[0078] Step 2:

[0079] The user launches the application on the device, selects the video they want to watch, and enters the time-saving playback settings.

[0080] Step 3:

[0081] The user watches a video that has been efficiently edited based on the settings.

[0082] Step 4:

[0083] If a scene of interest to the user is skipped while watching, the user can press a dedicated button to jump to that scene.

[0084] By following the above steps, the user can efficiently view only the important or interesting parts of the content, thereby saving time.

[0085] Example 1

[0086] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0087] Previously, it was difficult to efficiently watch only the important or interesting parts of a long video. In particular, to obtain the information you wanted to watch within a limited time, you had to manually skip or fast-forward, which was inconvenient. Furthermore, there was no function to update the playlist in real time in response to user interactions, which meant that a flexible viewing experience was not provided.

[0088] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0089] In this invention, the server includes means for uploading video data to the server, means for dividing the video data into frames and clustering them by scene, means for evaluating and scoring the interest and importance of each scene, means for evaluating the scenes using audio data analysis, text data analysis, and visual data analysis, means for prioritizing the scenes based on the evaluated scores, means for generating a playlist based on a user's time-saving playback setting and playing the video, and means for updating the playlist in response to user interactions. This allows a user to efficiently view important or interesting parts within a limited time and update the playlist in real time.

[0090] "Video data" is data in the form of dynamic media that includes visual and audio information.

[0091] A "server" is a system that provides services to other computers and devices on a network and processes and stores data.

[0092] A "frame" refers to each still image that makes up a video, and when played back in succession, they form a moving image.

[0093] "Clustering" is a process of classifying data into groups with common characteristics, and is a technique used to divide scenes in videos.

[0094] A "scene" is a group of consecutive frames that are semantically consistent within a video, and is a part that expresses a story or event.

[0095] "Fun" refers to elements that interest and entertain viewers, and is one of the criteria for evaluating a scene or the entire video.

[0096] "Importance" is an index that indicates how valuable the content of a video is to the viewer, and is a criterion for determining the priority of scenes.

[0097] "Scoring" is the process of assigning a numerical score to each scene based on an evaluation standard.

[0098] "Audio data analysis" is a technology that converts the audio portion of a video into text and analyzes its content.

[0099] "Text data analysis" is a technique for analyzing the content of text and evaluating its meaning and importance.

[0100] "Visual data analysis" is a technology that analyzes the content of images and videos and identifies objects and people.

[0101] A "playlist" is a list created to determine the playback order and content of videos.

[0102] "Time-saving playback settings" are options that users can set to shorten the playback time of videos under certain conditions.

[0103] "User interaction" refers to the operations and inputs that a user makes to a system.

[0104] The present invention provides a system that enables users to efficiently watch videos. The system mainly comprises a server, a terminal, and a user.

[0105] The server receives video data uploaded by users. After receiving the video data, the server uses common video processing tools such as FFmpeg to divide the video into frames. For example, a 30-minute video at 30 FPS would be divided into 54,000 frames. Next, each frame is clustered into scenes. This process involves using OpenCV to detect changes in color and motion, and then using a clustering algorithm (e.g., K-means) to divide the scenes. Each scene is evaluated using audio data analysis (e.g., Google® Cloud Speech-to-Text, which converts speech to text), text data analysis (e.g., OpenAI® GPT-4®, which analyzes the meaning of text), and visual data analysis (e.g., OpenCV, which identifies objects and people). Based on the evaluation results, each scene is scored and stored in a database. This scoring information is used to generate a playlist when users watch the video.

[0106] The device provides the user with an interface to select the video they want to watch. The user can set time-saving playback settings, such as "play only the important parts for 50%" or "play only the interesting parts for 30%." These settings are sent to the server, which provides the device with data to generate a playlist based on the stored scoring information. The device generates a playlist based on this data and plays the video according to the settings selected by the user. If the user wants to select a specific scene or view a skipped scene while watching, the playlist is updated in real time and the information is synchronized with the server.

[0107] To watch a video, a user first uploads the video file to the server. Next, they launch a dedicated application on their device and select the video they want to watch. The user configures time-saving playback settings based on their interests and convenience. Once this configuration is complete, the server obtains the scene scoring information and sends it to the device. The device then generates a playlist based on this information and plays the video according to the settings. As the user interacts with the video while watching, the playlist is updated in real time, allowing for flexible viewing.

[0108] For example, suppose a user wants to watch a 30-minute talk show, but only has 15 minutes available. The user sets the device app to "play only 50% of the important parts." Use this prompt:

[0109] "I would like to upload the following 30-minute talk show video and play it based on the setting of 'play only 50% of the important parts'. Please rate the importance of each scene using a system that uses Google Cloud Speech-to-Text for audio analysis, OpenAI GPT-4 for text analysis, and OpenCV for visual analysis. Based on the evaluation results, please make a list of only the important scenes, so that it can be viewed in 15 minutes."

[0110] As described above, the present invention provides a system that allows users to efficiently view important or interesting parts within a limited time. The advanced analysis technology of the server and the user-friendly interface of the terminal realize a flexible and effective viewing experience.

[0111] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0112] Step 1:

[0113] The server receives the video data. The user selects a video file from the device and uploads it to the server. The input is the video file provided by the user, and the output is a video file saved in a temporary directory on the server. At this stage, data is sent and received using the HTTP protocol.

[0114] Step 2:

[0115] The server splits the video data into frames. The input is the video file saved in step 1, and the output is a set of still images for each frame. This process uses FFmpeg to split the video based on the number of seconds and frame rate. Specifically, a 30-minute video (30 FPS) will generate 54,000 still frames.

[0116] Step 3:

[0117] The server clusters each frame by scene. The input is the set of still images generated in step 2, and the output is a list of frames clustered by scene. OpenCV is used to analyze the color and motion of the images, and a clustering algorithm (e.g., K-means) is used to divide the scenes. Specifically, identical scenes are grouped together based on the similarity of the frames.

[0118] Step 4:

[0119] The server evaluates the scenes. The input is the list of scenes clustered in Step 3, and the output is an evaluation score for each scene. Google Cloud Speech-to-Text is used for audio data analysis, OpenAI GPT-4 for text data analysis, and OpenCV is used again for visual data analysis. Specifically, the server converts audio into text, analyzes its content, and then analyzes visual information to evaluate the importance and interest of the scene.

[0120] Step 5:

[0121] The server scores and prioritizes scenes based on the evaluation results. The input is the evaluation data obtained in step 4, and the output is a score list for each scene. Scoring is performed on a scale of 0 to 100, and scenes with higher scores have a higher viewing priority. Specifically, the evaluation score for each scene is quantified, and the scenes are sorted based on that.

[0122] Step 6:

[0123] The server saves the scored scenes in a database. The input is a list of scenes scored in step 5, and the output is the scoring information saved in the database. Specifically, data such as the scene frame number, start and end time, and score are stored in the database.

[0124] Step 7:

[0125] The user sets viewing preferences on the device. The user launches a dedicated application and selects the video they want to watch. The input is the user's viewing preferences (e.g., "play only the important parts at 50%), and the output is the viewing preferences data sent to the server. This setting is made via the application's UI.

[0126] Step 8:

[0127] The device retrieves the scoring information from the server. The input is the viewing preference data sent in step 7, and the output is the scoring information received from the server. The server filters the scoring information based on the user-specified preferences and sends it back to the device.

[0128] Step 9:

[0129] The device generates a playlist. The input is the scoring information obtained in step 8, and the output is the generated playlist. The playlist is ordered by scenes with the highest scores based on the user's viewing preferences, and is adjusted to match the configured playback percentage.

[0130] Step 10:

[0131] As the user watches the video, the playlist is updated in real time. The input is the user's viewing actions and interactions, and the output is the updated playlist and the video being played. If the user wants to select a specific scene or rewatch a skipped scene, they can do so using the device interface, and the content is synchronized with the server in real time. This allows users to flexibly customize their viewing experience.

[0132] (Application example 1)

[0133] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0134] Currently, many factories are taking various measures to improve work efficiency, but optimal suggestions and real-time improvements are still often done manually. Therefore, there is a need for a system that can automatically generate suggestions for efficiency and quickly identify and play back important scenes in videos. Currently, there is no effective method to solve this problem.

[0135] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0136] In this invention, the server includes a means for uploading video data to the server, a means for dividing the video data into frames and clustering them by scene, a means for evaluating and scoring the interest and importance of each scene, a means for prioritizing the scenes based on the evaluated scores, and a means for analyzing work videos in an industrial environment and generating suggestions for improving efficiency, thereby enabling automatic and effective improvement of work efficiency in a factory.

[0137] "Video data" means digital files containing moving images and audio.

[0138] "Uploading" is the act of a user sending data from their device to a server.

[0139] A "scene" is a part of a video that consists of a certain series of frames and represents a series of movements or situations.

[0140] "Clustering" is a data processing technique that classifies many frames into groups with similar attributes.

[0141] "Evaluation" is the act of quantitatively or qualitatively judging the importance and interest of each scene.

[0142] "Scoring" is the process of assigning a numerical value to each scene based on the evaluation results.

[0143] "Prioritization" is the act of ranking scored scenes according to their importance and interest.

[0144] "Industrial environment" means the place or conditions in which manufacturing or production takes place.

[0145] "Efficiency" refers to reducing waste in work and processes and improving productivity.

[0146] A "proposal" refers to a solution or improvement to a specific problem.

[0147] System Overview

[0148] This invention is a video analysis system for improving work efficiency in an industrial environment. The system mainly includes a means for uploading video data to a server, a means for dividing the video data into frames and clustering them by scene, a means for evaluating and scoring the interest and importance of each scene, a means for prioritizing scenes based on the evaluated scores, and a means for analyzing work videos in an industrial environment and generating suggestions for improving efficiency.

[0149] Hardware and software used

[0150] Hardware: Robots used in factories (e.g. drones for monitoring work)

[0151] Software: Python, OpenCV, scikit-learn, cloud AI services (e.g., Google Cloud AI)

[0152] Data processing and calculation process

[0153] Server behavior:

[0154] 1. Uploading video data: Users upload videos of work they have taken in the factory to the server. The uploaded video data is stored on the server.

[0155] 2. Video Frame Segmentation: The server splits the received video data into frames. OpenCV is used to split the video data into individual frames.

[0156] 3. Frame Clustering: Each frame is then classified into scenes using K-means clustering, which groups similar scenes together.

[0157] 4. Scene Scoring: The clustered scenes are evaluated through audio, text, and visual data analysis. Based on the evaluation results, the importance of each scene is scored. Specifically, a machine learning algorithm is applied using scikit-learn.

[0158] 5. Prioritization: Based on the scoring results, scenes are prioritized in order of importance.

[0159] Terminal behavior:

[0160] 1. Playlist generation: The user uses the device application to select videos and set time-saving playback settings. Based on this, the device retrieves prioritized scene information from the server and generates a playlist.

[0161] 2. Work efficiency suggestions: Based on the generated playlist, suggestions for work efficiency are displayed, allowing users to quickly review only the important scenes.

[0162] 3. Updates based on user interaction: The playlist is updated in real time based on user actions.

[0163] Specific examples

[0164] For example, consider a video recording of assembly work in a factory. When a user uploads the video to the server, the server divides the video into frames and classifies them into scenes using K-means clustering. Each scene is evaluated through audio data analysis, text data analysis, and visual data analysis, and the importance of the scene is scored based on the results. Scenes with higher importance are given higher priority and displayed on the user's device based on the generated playlist. This allows users to quickly check the efficiency of work in the factory and receive suggestions for improvement.

[0165] Prompt Sentence Examples

[0166] "Analyze videos recorded in a factory to determine which scenes are important and which parts can be made more efficient. Then, generate suggestions for improving efficiency for each scene."

[0167] In this way, a system that dramatically improves work efficiency within a factory can be realized.

[0168] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0169] Step 1:

[0170] The server receives the work videos taken by the user in the factory and stores the video data uploaded to the server.

[0171] Input: Video file uploaded from the user's device

[0172] Data processing: Save the video data on the server.

[0173] Output: Video data on the server

[0174] Step 2:

[0175] The server divides the video data into frames.

[0176] Input: Saved video data

[0177] Data processing: Video data is divided into individual frames using OpenCV.

[0178] Output: List of frame images

[0179] Step 3:

[0180] The server clusters each of the divided frames.

[0181] Input: List of frame images

[0182] Data operations: Use K-means clustering to group frames that belong to the same scene. Vectorize the frame images and cluster them in feature space.

[0183] Output: Clustered groups of frames

[0184] Step 4:

[0185] The server analyzes the clustered scenes and performs scoring.

[0186] Input: A group of clustered frames

[0187] Data calculation: Evaluate the importance of each scene through audio data analysis, text data analysis, and visual data analysis. Specifically, audio analysis uses the Google Cloud Speech-to-Text API, text analysis uses NLP (natural language processing) models, and visual data analysis applies machine learning algorithms.

[0188] Output: Importance score for each scene

[0189] Step 5:

[0190] The server prioritizes the scenes based on the scoring results.

[0191] Input: Importance score for each scene

[0192] Data calculation: sorting scenes in order of importance based on their scores.

[0193] Output: A prioritized list of scenes

[0194] Step 6:

[0195] The terminal generates a playlist based on the user's time-saving playback settings.

[0196] Input: A prioritized list of scenes and the user's time-saving playback settings

[0197] Data processing: Generate a playlist containing only important scenes, taking into account user preferences.

[0198] Output: Playlist

[0199] Step 7:

[0200] The terminal plays the videos based on the playlist.

[0201] Input: Playlist

[0202] Data processing: Play only the scenes included in the playlist in sequence.

[0203] Output: Video playback

[0204] Step 8:

[0205] The terminal updates the playlist in real time in response to user interactions.

[0206] Input: User actions (selecting a specific scene, skipping, etc.)

[0207] Data processing: Regenerate the playlist based on user actions to reflect updated information.

[0208] Output: Updated playlist and video playback changes

[0209] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.

[0210] This invention is a system that allows users to effectively shorten the time spent watching videos and efficiently watch only the interesting or important parts. This system consists of four main components: a server, a terminal, a user, and an emotion engine.

[0211] Server Operation

[0212] The server receives video data uploaded by users. After receiving the video data, the server divides the video into frames and clusters each frame into scenes. The clustered scenes are evaluated through audio data analysis, text data analysis, and visual data analysis, and each scene is scored for its interest and importance. The scored scenes are prioritized based on the evaluation results and stored in a database on the server.

[0213] Device behavior

[0214] When a user wants to watch a video, they first launch the application on their device and select a video. Users can set viewing preferences for time-saving playback, such as "play only the important parts for 50%" or "play only the interesting parts for 30%." These preferences are sent from the device to the server, which then obtains scoring information for prioritized scenes. The device generates a playlist based on this information and plays the video according to the preferences. If the user selects a specific scene or wants to view a skipped scene while watching, the playlist is updated in real time. The server reflects the user's actions, and a regenerated playlist is sent to the device.

[0215] User behavior

[0216] When a user wants to watch a video, they first upload the video file to the server. Next, they launch the application on their device and select the video they want to watch. The user then configures time-saving playback settings to suit their convenience and interests. Once playback settings are configured, the device automatically obtains scene scoring information from the server and generates a playlist. If the user wants to check a scene that was skipped during viewing or cancel time-saving playback, they can press a button to update the playlist in real time, allowing them to view content according to their needs.

[0217] Emotion Engine Operation

[0218] The device is equipped with an emotion engine that monitors the user's emotions while watching videos. The engine analyzes the user's facial expressions, tone of voice, and eye movements to recognize the user's emotional state in real time. The emotion data recognized by the emotion engine is fed back to the device and used to adjust the playlist.

[0219] For example, if a user is enjoying a scene, the score for that scene will increase, and similar scenes will be added to the list with priority. Emotional data will also be accumulated and used as learning data to reflect in future scoring and prioritization. This function allows video viewing to be more tailored to the user's preferences.

[0220] Specific examples

[0221] For example, consider a case where a user wants to watch a 30-minute talk show but only has 50% of the time. The user sets the device application to "play only 50% of the important parts." When the video is uploaded to the server, the server divides the video into frames and uses an AI model to score the importance of each scene. Based on the scoring results, the server prioritizes the scenes and sends them back to the device. Based on the settings, the device lists only the important scenes and plays them so that the user can watch efficiently.

[0222] While watching, the emotion engine monitors the user's reactions and, if they are enjoying the content, adjusts the priority of the corresponding scene in real time. If a scene that interests the user is skipped while watching, they can press a dedicated button to jump to that scene. This system not only allows users to efficiently view important information and interesting parts within a limited time, but also provides a viewing experience customized to their emotions.

[0223] The processing flow will be explained below.

[0224] Server Processing

[0225] Step 1:

[0226] The server receives video data from the user and stores it in temporary storage.

[0227] Step 2:

[0228] The server divides the video data into frames and converts the video into individual still images.

[0229] Step 3:

[0230] The server clusters frames by scene, which groups frames with similar characteristics together.

[0231] Step 4:

[0232] To evaluate each scene, the server performs audio data analysis, text data analysis, and visual data analysis, such as analyzing the speaker's content from the audio data and analyzing subtitles and descriptions from the text data.

[0233] Step 5:

[0234] The server scores each scene based on its interest and importance, and assigns a score to each scene based on the evaluation results.

[0235] Step 6:

[0236] The server prioritizes scenes based on the evaluated scores and stores the priority information and scoring results in a database.

[0237] Terminal handling

[0238] Step 1:

[0239] The device receives the user's time-saving playback settings, such as "play only the important parts at 50%."

[0240] Step 2:

[0241] The device retrieves scene scoring information, including prioritization information, from the server.

[0242] Step 3:

[0243] The device generates a playlist based on the user's settings, prioritizing important or interesting scenes.

[0244] Step 4:

[0245] The device plays videos based on the playlist, and unnecessary scenes are skipped as the user watches.

[0246] Step 5:

[0247] If the user wants to view a scene that was skipped during viewing, the device will accept a dedicated button and readjust the playlist based on that information.

[0248] Step 6:

[0249] The device sends the updated playlist to the server and continues playing the video based on the regenerated playlist.

[0250] User Action

[0251] Step 1:

[0252] The user selects a video and uploads it to the server.

[0253] Step 2:

[0254] The user launches the application on the device, selects the video they want to watch, and enters the time-saving playback settings.

[0255] Step 3:

[0256] The user watches a video that has been efficiently edited based on the settings.

[0257] Step 4:

[0258] If a scene of interest to the user is skipped while watching, the user can press a dedicated button to jump to that scene.

[0259] Emotion engine processing

[0260] Step 1:

[0261] The device's built-in emotion engine monitors the user's facial expressions, tone of voice, gaze, and other factors in real time.

[0262] Step 2:

[0263] The emotion engine analyzes the user's emotional state and generates emotion data based on that.

[0264] Step 3:

[0265] The generated emotion data is received by the device and used to adjust the playlist, for example by increasing the score of scenes that are enjoyed.

[0266] Step 4:

[0267] Emotional data is accumulated and sent to the server to be used as learning data for future scoring and prioritization.

[0268] Specific examples

[0269] For example, consider a case where a user wants to watch a 30-minute talk show but only has 50% of the time. The user sets the device application to "play only 50% of the important parts." When the video is uploaded to the server, the server divides the video into frames and uses an AI model to score the importance of each scene. Based on the scoring results, the server prioritizes the scenes and sends them back to the device. The device then lists only the important scenes based on the settings and plays them back so that the user can watch efficiently.

[0270] While watching, the emotion engine monitors the user's reactions and, if they are enjoying the content, adjusts the priority of the corresponding scene in real time. If a scene that interests the user is skipped while watching, they can press a dedicated button to jump to that scene. This system not only allows users to efficiently view important information and interesting parts within a limited time, but also provides a viewing experience that is customized according to their emotions.

[0271] Example 2

[0272] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0273] Conventional video viewing systems make it difficult for users to efficiently grasp important or interesting parts of a video. They also lack a means to meet the need to view only the most important information within a limited time. Furthermore, content optimization based on user emotions is not performed, preventing consistent satisfaction.

[0274] The identification process by the identification processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means. In this invention, the server includes a means for uploading video data to the information processing device, a means for dividing the video data into frames and classifying them by scene, a means for evaluating and scoring the interest and importance of each scene, a means for prioritizing the scenes based on the evaluated scores, a means for generating a playlist based on the user's playback settings and playing the video, a means for updating the playlist in response to the user's operations, and a means for recognizing the user's emotions during viewing and adjusting the playlist. This makes it possible to streamline the video viewing experience and provide customized playback according to the user's interests and emotions.

[0275] "Video data" means a digitally represented video file containing visual and audio information.

[0276] An "information processing device" is a computer system for processing, storing, and analyzing data.

[0277] A "user" is a person who operates the system and watches videos.

[0278] A "frame" is an individual still image that makes up a video.

[0279] A "scene" is a set of frames that represent a certain continuous time and space within a video.

[0280] "Classification" is the process of grouping frames that have the same characteristics or features.

[0281] "Scoring" is the process of quantifying the interest and importance of a scene based on certain criteria.

[0282] "Prioritization" means determining the order in which scenes are played based on the evaluated scores.

[0283] A "playlist" is a list that defines the order in which scenes are to be played.

[0284] "Playback" is the process of visually and aurally representing a moving image.

[0285] "Playback settings" are settings that specify the playback method and range of the video that the user wants to watch.

[0286] "Emotions" are the psychological and physiological reactions of users while watching.

[0287] "Recognition" refers to grasping a user's emotional state by analyzing their facial expressions, tone of voice, eye movements, etc.

[0288] "Adjusting" means changing scene priorities and playlists in real time.

[0289] "Operation" refers to the instructions or input that a user gives to the system.

[0290] This invention is a system that allows users to effectively shorten the time it takes to watch a video and efficiently watch only the interesting or important parts. This system consists of four main components: a server, a terminal, a user, and an emotion engine.

[0291] Server Operation

[0292] The server receives video data uploaded by users. Once the video data is received, the server divides the video into frames using FFmpeg. The divided frames are then clustered by scene using the FrameGenius API. The clustered scenes are analyzed for audio, text, and visual data using Google Cloud's Speech-to-Text API and NLP Cloud. Based on the analyzed data, the interestingness and importance of each scene is evaluated and scored using a proprietary scoring algorithm. The scenes are prioritized based on the scoring results, and the results are stored in a MongoDB database.

[0293] Device behavior

[0294] When a user wants to watch a video, they launch a dedicated application on their device and select the video they want to watch. Users can set time-saving playback settings such as "play only the important parts at 50%" or "play only the interesting parts at 30%." This setting information is sent from the device to the server, which then obtains scoring information for prioritized scenes. The device generates a playlist based on this information and plays the video using HLS (HTTP Live Streaming) technology. If the user selects a specific scene while watching or wants to watch a skipped scene, the playlist is updated in real time and the information is reflected on the server.

[0295] User actions

[0296] Users select the video they want to watch from their device and upload it to the server. Next, they launch a dedicated application on their device and select the video they want to watch. After selecting the video they want to watch, users can set up time-saving playback according to their interests and convenience. Once playback settings are set, the device automatically obtains scene scoring information from the server and generates a playlist. If users want to check scenes that were skipped during viewing or cancel time-saving playback, they can operate a dedicated button to update the playlist in real time, allowing them to watch according to their needs.

[0297] Emotion Engine Operation

[0298] The device is equipped with an emotion engine that monitors the user's emotions while watching videos using the Momento AI engine. This engine analyzes the user's facial expressions, vocal tone, eye movements, and other factors to recognize the user's emotional state in real time. The recognized emotion data is fed back to the device and used to adjust the playlist. For example, if the user is enjoying a scene very much, the score for that scene will increase, and similar scenes will be added to the playlist with priority. Emotion data is also accumulated and used as learning data to reflect in future scoring and prioritization. This function enables video viewing that is more tailored to the user's preferences.

[0299] Specific examples

[0300] For example, consider the case where a user wants to watch a 30-minute talk show but only has 15 minutes to watch. The user sets the device application to "play only the important parts at 50%." When the video is uploaded to the server, the server uses FFmpeg to divide the video into frames and clusters each scene using the FrameGenius API. The server then analyzes the video using Google Cloud's Speech-to-Text API and NLP Cloud to score its importance. Prioritization is performed based on the scoring results, and the results are sent to the device. The device then lists only the important scenes in a playlist and plays the video using HLS technology.

[0301] While watching, the emotional engine monitors the user's reactions using the Momento AI engine, and if the user is enjoying the video, the priority of the corresponding scene is adjusted in real time. Users can jump to the relevant scene by pressing the "Watch Skipped Scene" button. This system not only allows users to efficiently view important information and interesting parts within a limited time, but also provides a viewing experience customized to their emotions.

[0302] Prompt Sentence Examples

[0303] "I want to set up time-saving playback so that I can watch only the important parts of a 30-minute talk show video in 50% of the viewing time."

[0304] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0305] Step 1: Upload your video

[0306] The server receives video data from the user. The user selects the video they want to watch on their device and uploads it to the server using a dedicated application. Input: Video file sent from the user's device. Output: Video data saved on the server.

[0307] Specific behavior:

[0308] A user opens the application and clicks the "Upload Video" button.

[0309] The user selects the video they want to watch from the file selection dialog and clicks "Open."

[0310] The video file is sent to the server and the upload is complete.

[0311] Step 2: Video analysis and frame segmentation

[0312] The server splits the received video into frames using FFmpeg. Input: Video data stored on the server. Output: Still images for each frame.

[0313] Specific behavior:

[0314] The server splits the received video data into frames using the FFmpeg tool.

[0315] The split frames are saved in a temporary folder.

[0316] Step 3: Scene Clustering

[0317] The server uses the FrameGenius API to classify frames into scenes. Input: Segmented frame data. Output: Clustered scene data.

[0318] Specific behavior:

[0319] The FrameGenius API is used to analyze the visual characteristics of each frame.

[0320] Frames are grouped into clusters based on similarity and organized into scenes.

[0321] Step 4: Analyze and score the scene

[0322] The server uses Google Cloud's Speech-to-Text API and NLP Cloud to analyze the audio, text, and visual data for each scene. Input: Clustered scene data. Output: Analysis results and scoring data for each scene.

[0323] Specific behavior:

[0324] The audio data contained in each clustered scene is converted into text data using Google Cloud's Speech-to-Text API.

[0325] Text data is analyzed using NLP Cloud to extract keywords and emotions.

[0326] The visual data of each scene is also analyzed to evaluate its importance and interest.

[0327] Step 5: Prioritize and save your scenes

[0328] The server uses a proprietary algorithm to determine the priority of scenes based on the analysis results and stores the results in a MongoDB database. Input: Scoring data for each scene. Output: A prioritized scene list.

[0329] Specific behavior:

[0330] Scenes are sorted by priority based on the scoring data for each scene.

[0331] Store the prioritized scene list in a MongoDB database.

[0332] Step 6: Submit viewing preferences

[0333] The user configures the time-saving playback settings using the application on the device and sends this setting information to the server. Input: User's viewing settings. Output: Viewing setting data sent to the server.

[0334] Specific behavior:

[0335] The user opens the application and selects the video they want to watch.

[0336] Users can select viewing settings such as "play only the important parts at 50%" or "play only the interesting parts at 30%."

[0337] The configuration information is sent to the server.

[0338] Step 7: Generate and play a playlist

[0339] The server sends prioritized scene scoring information based on the received viewing preferences to the device, which then generates a playlist based on that information and plays the video using HLS technology. Input: User's viewing preferences, prioritized scene list. Output: Playlist, video to be played.

[0340] Specific behavior:

[0341] The server filters out priority scenes according to viewing settings and transmits them to the terminal.

[0342] The device uses HLS technology to generate a playlist and plays the video in scene order.

[0343] Step 8: Emotion Engine Feedback

[0344] The emotion engine built into the device monitors the user's emotional state in real time and adjusts the playlist. Input: Real-time emotion data while the user is watching. Output: Adjusted playlist.

[0345] Specific behavior:

[0346] The emotion engine analyzes the user's facial expressions, tone of voice, and eye movements.

[0347] Using the Momento AI engine, it recognizes the user's emotional state and provides feedback in real time.

[0348] Adjust the playlist to include scenes that users enjoy or find interesting.

[0349] (Application example 2)

[0350] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0351] While modern consumers have access to a wealth of video content, it is difficult for them to efficiently obtain the most interesting content and important information within a limited viewing time. Furthermore, the ability to adjust playback content in real time based on the viewer's emotions and reactions is also lacking. In response to this problem, the present invention aims to provide a system that provides a video viewing experience that maximizes user engagement within a limited time.

[0352] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[0353] In this invention, the server includes means for uploading video data to the server, means for dividing the video data into frames and clustering them by scene, means for evaluating and scoring the interest and importance of each scene, means for prioritizing the scenes based on the evaluated scores, means for generating a playlist based on a user's time-saving playback settings and playing the videos, means for updating the playlist in response to user interactions, and means for analyzing the user's facial expressions, tone of voice, and eye movements and adjusting the playlist based on emotion data. This allows users to efficiently watch only interesting or important parts even within a limited time, and further enables a personalized viewing experience based on emotions.

[0354] "Means for uploading video data to a server" refers to a function that allows a user to send video files from their own device to a server.

[0355] The "means of dividing video data into frames and clustering by scene" is a process of dividing a video uploaded to a server into frames and grouping similar frames.

[0356] The "means for evaluating and scoring the interest and importance of each scene" is a process of analyzing the clustered scenes and assigning a numerical evaluation of their interest and importance.

[0357] The "means for prioritizing scenes based on the evaluated scores" is a mechanism for sorting scored scenes in order of importance based on their scores.

[0358] "Means for generating a playlist based on the user's time-saving playback settings and playing videos" refers to a function that creates an optimized playlist based on the user's specified playback time and priority scene conditions, and plays videos according to that list.

[0359] The "means for updating the playlist in response to user interaction" refers to a system that dynamically changes the playlist based on the operations performed by the user while viewing (e.g., selecting or skipping a specific scene).

[0360] "Means of analyzing the user's facial expressions, tone of voice, and eye movements, and adjusting the playlist based on emotional data" refers to a function that analyzes the user's emotions and reactions while watching, and adjusts the playlist on the spot based on the results.

[0361] This invention provides a system that enables users to effectively shorten the time spent watching videos and efficiently watch only the interesting or important parts. This system consists of four main components: a server, a terminal, a user, and an emotion engine.

[0362] Server Operation

[0363] The server receives video data uploaded by users. After receiving the video data, the server divides the video into frames and clusters each frame into scenes. The clustered scenes are evaluated through audio data analysis, text data analysis, and visual data analysis to score each scene's interest and importance. The scenes are prioritized based on the evaluation results and stored in a database.

[0364] Device behavior

[0365] When a user wants to watch a video, they first launch the application on their device and select a video. Users can set viewing preferences for time-saving playback, such as "play only the important parts for 50%" or "play only the interesting parts for 30%." These preferences are sent from the device to the server, which then obtains scoring information for prioritized scenes. The device generates a playlist based on this information and plays the video according to the preferences. If the user selects a specific scene or wants to view a skipped scene while watching, the playlist is updated in real time. The user's actions are reflected on the server, and a regenerated playlist is sent to the device.

[0366] User behavior

[0367] When a user wants to watch a video, they first upload the video file to the server. Next, they launch the application on their device and select the video they want to watch. The user then configures time-saving playback settings to suit their convenience and interests. Once playback settings are configured, the device automatically obtains scene scoring information from the server and generates a playlist. If the user wants to check a scene that was skipped during viewing or cancel time-saving playback, they can press a button to update the playlist in real time, allowing them to watch according to their needs.

[0368] Emotion Engine Operation

[0369] The device is equipped with an emotion engine that monitors the user's emotions while watching videos. This engine analyzes the user's facial expressions, vocal tone, eye movements, and other factors to recognize the user's emotional state in real time. The emotion data recognized by the emotion engine is fed back to the device and used to adjust the playlist. For example, if the user is enjoying a video very much, the score for that scene will increase and similar scenes will be added to the playlist with priority. Emotion data is also accumulated and used as learning data to reflect in future scoring and prioritization. This function enables video viewing that is more tailored to the user's preferences.

[0370] Specific examples

[0371] Consider a scenario where a user wants to watch a 30-minute talk show but only has 50% of the time available. The user selects "Play only 50% of the important parts" in the device application. When the video is uploaded to the server, the server divides it into frames and uses a generative AI model to score the importance of each scene. Based on the scoring results, the server prioritizes the scenes and sends them back to the device. The device then lists only the important scenes based on the user's settings and plays them for efficient viewing. While watching, the emotion engine monitors the user's reactions. If the user is enjoying the video, the priority of the corresponding scene is adjusted in real time. If a scene that interests the user is skipped while watching, the user can press a dedicated button to jump back to that scene. This system not only allows users to efficiently view important information and interesting parts within a limited time, but also provides a customized viewing experience tailored to their emotions.

[0372] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0373] Step 1:

[0374] The user uploads a video file to the server. The user selects a video file using an application on their device and presses the upload button, sending the video data to the server. The input is the video file selected by the user, and the output is the video data saved on the server.

[0375] Step 2:

[0376] The server divides the uploaded video data into frames. When the server receives the video data, it divides the video into frames and stores each frame in memory. The input is video data, and the output is image data divided into frames.

[0377] Step 3:

[0378] The server clusters the image data for each frame by scene. The server uses a clustering algorithm to classify similar frames into the same scene cluster. The input is the divided frame image data, and the output is the clustered scene groups.

[0379] Step 4:

[0380] The server evaluates and scores the interest and importance of each scene group. The server performs audio, text, and visual analysis, and uses a generative AI model to evaluate the importance and interest of each scene as a numerical value. The input is the clustered scene data, and the output is the evaluation score.

[0381] Step 5:

[0382] The server prioritizes scenes based on the evaluated scores. The server prioritizes the playback order of each scene based on the scoring results. The input is the evaluation scores, and the output is a list of prioritized scene clusters.

[0383] Step 6:

[0384] The device sends the user's time-saving playback settings to the server and generates a playlist. The user configures playback settings within the application, and the device sends them to the server. The server creates a playlist based on the prioritized scene list and matches the user's settings. The input is the user's playback settings, and the output is an optimized playlist.

[0385] Step 7:

[0386] The device plays videos based on the generated playlist. The device receives the playlist sent from the server and plays videos according to that list. The input is the playlist and the output is the playback of the video.

[0387] Step 8:

[0388] The device updates the playlist according to the user's interactions. When the user skips or selects a specific scene while watching a video, the device sends that information to the server in real time and updates the playlist. The input is the user's interaction data, and the output is the updated playlist.

[0389] Step 9:

[0390] The device uses an emotion engine to monitor the user's emotional state and adjust the playlist accordingly. The emotion engine built into the device analyzes the user's facial expressions, tone of voice, and eye movements to obtain emotional data in real time. The obtained emotional data is sent to the server, which then adjusts the playlist based on that data. The input is the user's emotional data, and the output is a playlist adjusted based on the emotions.

[0391] Step 10:

[0392] The server sends the adjusted playlist to the terminal, and the terminal updates the playlist. The input is the adjusted playlist, and the output is the playlist update.

[0393] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0394] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (registered trademark) (Internet search engine).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0395] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.

[0396] [Second embodiment]

[0397] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.

[0398] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0399] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0400] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.

[0401] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0402] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0403] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0404] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0405] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0406] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0407] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0408] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."

[0409] This invention is a system that allows users to effectively shorten the time spent watching videos and efficiently watch only the interesting or important parts. This system consists of three main components: a server, a terminal, and a user.

[0410] Server Operation

[0411] The server receives video data uploaded by users. After receiving the video data, the server divides the video into frames and clusters each frame into scenes. The clustered scenes are evaluated through audio data analysis, text data analysis, and visual data analysis, and each scene is scored for its interest and importance. The scored scenes are prioritized based on the evaluation results and stored in a database on the server.

[0412] Device behavior

[0413] When a user wants to watch a video, they first launch the application on their device and select a video. Users can set viewing preferences for time-saving playback, such as "play only the important parts for 50%" or "play only the interesting parts for 30%." These preferences are sent from the device to the server, which then obtains scoring information for prioritized scenes. The device generates a playlist based on this information and plays the video according to the preferences. If the user selects a specific scene or wants to view a skipped scene while watching, the playlist is updated in real time. The server reflects the user's actions, and a regenerated playlist is sent to the device.

[0414] User behavior

[0415] When a user wants to watch a video, they first upload the video file to the server. Next, they launch the application on their device and select the video they want to watch. The user then configures time-saving playback settings to suit their convenience and interests. Once playback settings are configured, the device automatically obtains scene scoring information from the server and generates a playlist. If the user wants to check a scene that was skipped during viewing or cancel time-saving playback, they can press a button to update the playlist in real time, allowing them to view content according to their needs.

[0416] Specific examples

[0417] For example, consider a case where a user wants to watch a 30-minute talk show but only has 50% of the time. The user sets the device's application to "play only 50% of the important parts." When the video is uploaded to the server, the server divides the video into frames and uses an AI model to score the importance of each scene. Based on the scoring results, the server prioritizes the scenes and sends them back to the device. The device then lists only the important scenes based on the settings and plays them back so that the user can watch efficiently. If a scene of interest is skipped while watching, the user can press a dedicated button to jump to that scene. This system allows users to watch within a limited time without missing any important information.

[0418] The processing flow will be explained below.

[0419] Server Processing

[0420] Step 1:

[0421] The server receives the video data from the user and stores it in temporary storage.

[0422] Step 2:

[0423] The server divides the video data into frames, i.e., converts the video into multiple still images.

[0424] Step 3:

[0425] The server clusters frames by scene, which groups frames with similar characteristics together.

[0426] Step 4:

[0427] To evaluate each scene, the server analyzes audio data, text data, and visual data. For example, it analyzes the speaker's content from audio data and subtitles and descriptions from text data.

[0428] Step 5:

[0429] The server scores each scene based on its interest and importance, and assigns a score to each scene based on the evaluation results.

[0430] Step 6:

[0431] The server prioritizes scenes based on the evaluated scores, and the priority information and scoring results are stored in a database.

[0432] Terminal handling

[0433] Step 1:

[0434] The device receives the user's time-saving playback settings, such as "play only the important parts at 50%."

[0435] Step 2:

[0436] The device retrieves scene scoring information from the server, including scene priority information.

[0437] Step 3:

[0438] The device will generate a playlist based on the user's settings, prioritizing important or interesting scenes.

[0439] Step 4:

[0440] The device plays videos based on the playlist, and unnecessary scenes are skipped as the user watches.

[0441] Step 5:

[0442] If a user wants to view a scene that was skipped during viewing, the device will accept a dedicated button and will readjust the playlist based on that information.

[0443] Step 6:

[0444] The device sends the updated playlist to the server and continues playing the video based on the regenerated playlist.

[0445] User Action

[0446] Step 1:

[0447] The user selects a video and uploads it to the server.

[0448] Step 2:

[0449] The user launches the application on the device, selects the video they want to watch, and enters the time-saving playback settings.

[0450] Step 3:

[0451] The user watches a video that has been efficiently edited based on the settings.

[0452] Step 4:

[0453] If a scene of interest to the user is skipped while watching, the user can press a dedicated button to jump to that scene.

[0454] By following the above steps, the user can efficiently view only the important or interesting parts of the content, thereby saving time.

[0455] Example 1

[0456] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0457] Previously, it was difficult to efficiently watch only the important or interesting parts of a long video. In particular, to obtain the information you wanted to watch within a limited time, you had to manually skip or fast-forward, which was inconvenient. Furthermore, there was no function to update the playlist in real time in response to user interactions, which meant that a flexible viewing experience was not provided.

[0458] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0459] In this invention, the server includes means for uploading video data to the server, means for dividing the video data into frames and clustering them by scene, means for evaluating and scoring the interest and importance of each scene, means for evaluating the scenes using audio data analysis, text data analysis, and visual data analysis, means for prioritizing the scenes based on the evaluated scores, means for generating a playlist based on a user's time-saving playback setting and playing the video, and means for updating the playlist in response to user interactions. This allows a user to efficiently view important or interesting parts within a limited time and update the playlist in real time.

[0460] "Video data" is data in the form of dynamic media that includes visual and audio information.

[0461] A "server" is a system that provides services to other computers and devices on a network and processes and stores data.

[0462] A "frame" refers to each still image that makes up a video, and when played back in succession, they form a moving image.

[0463] "Clustering" is a process of classifying data into groups with common characteristics, and is a technique used to divide scenes in videos.

[0464] A "scene" is a group of consecutive frames that are semantically consistent within a video, and is a part that expresses a story or event.

[0465] "Fun" refers to elements that interest and entertain viewers, and is one of the criteria for evaluating a scene or the entire video.

[0466] "Importance" is an index that indicates how valuable the content of a video is to the viewer, and is a criterion for determining the priority of scenes.

[0467] "Scoring" is the process of assigning a numerical score to each scene based on an evaluation standard.

[0468] "Audio data analysis" is a technology that converts the audio portion of a video into text and analyzes its content.

[0469] "Text data analysis" is a technique for analyzing the content of text and evaluating its meaning and importance.

[0470] "Visual data analysis" is a technology that analyzes the content of images and videos and identifies objects and people.

[0471] A "playlist" is a list created to determine the playback order and content of videos.

[0472] "Time-saving playback settings" are options that users can set to shorten the playback time of videos under certain conditions.

[0473] "User interaction" refers to the operations and inputs that a user makes to a system.

[0474] The present invention provides a system that enables users to efficiently watch videos. The system mainly comprises a server, a terminal, and a user.

[0475] The server receives video data uploaded by users. After receiving the video data, the server uses common video processing tools such as FFmpeg to divide the video into frames. For example, a 30-minute video at 30 FPS would be divided into 54,000 frames. Next, each frame is clustered into scenes. This process involves using OpenCV to detect changes in color and motion, and then using a clustering algorithm (e.g., K-means) to divide the scenes. Each scene is evaluated using audio data analysis (e.g., Google Cloud Speech-to-Text, which converts speech to text), text data analysis (e.g., OpenAI GPT-4, which analyzes the meaning of text), and visual data analysis (e.g., OpenCV, which identifies objects and people). Each scene is scored based on the evaluation results and stored in a database. This scoring information is used to generate a playlist when users watch the video.

[0476] The device provides the user with an interface to select the video they want to watch. The user can set time-saving playback settings, such as "play only the important parts for 50%" or "play only the interesting parts for 30%." These settings are sent to the server, which provides the device with data to generate a playlist based on the stored scoring information. The device generates a playlist based on this data and plays the video according to the settings selected by the user. If the user wants to select a specific scene or view a skipped scene while watching, the playlist is updated in real time and the information is synchronized with the server.

[0477] To watch a video, a user first uploads the video file to the server. Next, they launch a dedicated application on their device and select the video they want to watch. The user configures time-saving playback settings based on their interests and convenience. Once this configuration is complete, the server obtains the scene scoring information and sends it to the device. The device then generates a playlist based on this information and plays the video according to the settings. As the user interacts with the video while watching, the playlist is updated in real time, allowing for flexible viewing.

[0478] For example, suppose a user wants to watch a 30-minute talk show, but only has 15 minutes available. The user sets the device app to "play only 50% of the important parts." Use this prompt:

[0479] "I would like to upload the following 30-minute talk show video and play it based on the setting of 'play only 50% of the important parts'. Please rate the importance of each scene using a system that uses Google Cloud Speech-to-Text for audio analysis, OpenAI GPT-4 for text analysis, and OpenCV for visual analysis. Based on the evaluation results, please make a list of only the important scenes, so that it can be viewed in 15 minutes."

[0480] As described above, the present invention provides a system that allows users to efficiently view important or interesting parts within a limited time. The advanced analysis technology of the server and the user-friendly interface of the terminal realize a flexible and effective viewing experience.

[0481] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0482] Step 1:

[0483] The server receives the video data. The user selects a video file from the device and uploads it to the server. The input is the video file provided by the user, and the output is a video file saved in a temporary directory on the server. At this stage, data is sent and received using the HTTP protocol.

[0484] Step 2:

[0485] The server splits the video data into frames. The input is the video file saved in step 1, and the output is a set of still images for each frame. This process uses FFmpeg to split the video based on the number of seconds and frame rate. Specifically, a 30-minute video (30 FPS) will generate 54,000 still frames.

[0486] Step 3:

[0487] The server clusters each frame by scene. The input is the set of still images generated in step 2, and the output is a list of frames clustered by scene. OpenCV is used to analyze the color and motion of the images, and a clustering algorithm (e.g., K-means) is used to divide the scenes. Specifically, identical scenes are grouped together based on the similarity of the frames.

[0488] Step 4:

[0489] The server evaluates the scenes. The input is the list of scenes clustered in Step 3, and the output is an evaluation score for each scene. Google Cloud Speech-to-Text is used for audio data analysis, OpenAI GPT-4 for text data analysis, and OpenCV is used again for visual data analysis. Specifically, the server converts audio into text, analyzes its content, and then analyzes visual information to evaluate the importance and interest of the scene.

[0490] Step 5:

[0491] The server scores and prioritizes scenes based on the evaluation results. The input is the evaluation data obtained in step 4, and the output is a score list for each scene. Scoring is performed on a scale of 0 to 100, and scenes with higher scores have a higher viewing priority. Specifically, the evaluation score for each scene is quantified, and the scenes are sorted based on that.

[0492] Step 6:

[0493] The server saves the scored scenes in a database. The input is a list of scenes scored in step 5, and the output is the scoring information saved in the database. Specifically, data such as the scene frame number, start and end time, and score are stored in the database.

[0494] Step 7:

[0495] The user sets viewing preferences on the device. The user launches a dedicated application and selects the video they want to watch. The input is the user's viewing preferences (e.g., "play only the important parts at 50%), and the output is the viewing preferences data sent to the server. This setting is made via the application's UI.

[0496] Step 8:

[0497] The device retrieves the scoring information from the server. The input is the viewing preference data sent in step 7, and the output is the scoring information received from the server. The server filters the scoring information based on the user-specified preferences and sends it back to the device.

[0498] Step 9:

[0499] The device generates a playlist. The input is the scoring information obtained in step 8, and the output is the generated playlist. The playlist is ordered by scenes with the highest scores based on the user's viewing preferences, and is adjusted to match the configured playback percentage.

[0500] Step 10:

[0501] As the user watches the video, the playlist is updated in real time. The input is the user's viewing actions and interactions, and the output is the updated playlist and the video being played. If the user wants to select a specific scene or rewatch a skipped scene, they can do so using the device interface, and the content is synchronized with the server in real time. This allows users to flexibly customize their viewing experience.

[0502] (Application example 1)

[0503] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0504] Currently, many factories are taking various measures to improve work efficiency, but optimal suggestions and real-time improvements are still often done manually. Therefore, there is a need for a system that can automatically generate suggestions for efficiency and quickly identify and play back important scenes in videos. Currently, there is no effective method to solve this problem.

[0505] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0506] In this invention, the server includes a means for uploading video data to the server, a means for dividing the video data into frames and clustering them by scene, a means for evaluating and scoring the interest and importance of each scene, a means for prioritizing the scenes based on the evaluated scores, and a means for analyzing work videos in an industrial environment and generating suggestions for improving efficiency, thereby enabling automatic and effective improvement of work efficiency in a factory.

[0507] "Video data" means digital files containing moving images and audio.

[0508] "Uploading" is the act of a user sending data from their device to a server.

[0509] A "scene" is a part of a video that consists of a certain series of frames and represents a series of movements or situations.

[0510] "Clustering" is a data processing technique that classifies many frames into groups with similar attributes.

[0511] "Evaluation" is the act of quantitatively or qualitatively judging the importance and interest of each scene.

[0512] "Scoring" is the process of assigning a numerical value to each scene based on the evaluation results.

[0513] "Prioritization" is the act of ranking scored scenes according to their importance and interest.

[0514] "Industrial environment" means the place or conditions in which manufacturing or production takes place.

[0515] "Efficiency" refers to reducing waste in work and processes and improving productivity.

[0516] A "proposal" refers to a solution or improvement to a specific problem.

[0517] System Overview

[0518] This invention is a video analysis system for improving work efficiency in an industrial environment. The system mainly includes a means for uploading video data to a server, a means for dividing the video data into frames and clustering them by scene, a means for evaluating and scoring the interest and importance of each scene, a means for prioritizing scenes based on the evaluated scores, and a means for analyzing work videos in an industrial environment and generating suggestions for improving efficiency.

[0519] Hardware and software used

[0520] Hardware: Robots used in factories (e.g. drones for monitoring work)

[0521] Software: Python, OpenCV, scikit-learn, cloud AI services (e.g., Google Cloud AI)

[0522] Data processing and calculation process

[0523] Server behavior:

[0524] 1. Uploading video data: Users upload videos of work they have taken in the factory to the server. The uploaded video data is stored on the server.

[0525] 2. Video Frame Segmentation: The server splits the received video data into frames. OpenCV is used to split the video data into individual frames.

[0526] 3. Frame Clustering: Each frame is then classified into scenes using K-means clustering, which groups similar scenes together.

[0527] 4. Scene Scoring: The clustered scenes are evaluated through audio, text, and visual data analysis. Based on the evaluation results, the importance of each scene is scored. Specifically, a machine learning algorithm is applied using scikit-learn.

[0528] 5. Prioritization: Based on the scoring results, scenes are prioritized in order of importance.

[0529] Terminal behavior:

[0530] 1. Playlist generation: The user uses the device application to select videos and set time-saving playback settings. Based on this, the device retrieves prioritized scene information from the server and generates a playlist.

[0531] 2. Work efficiency suggestions: Based on the generated playlist, suggestions for work efficiency are displayed, allowing users to quickly review only the important scenes.

[0532] 3. Updates based on user interaction: The playlist is updated in real time based on user actions.

[0533] Specific examples

[0534] For example, consider a video recording of assembly work in a factory. When a user uploads the video to the server, the server divides the video into frames and classifies them into scenes using K-means clustering. Each scene is evaluated through audio data analysis, text data analysis, and visual data analysis, and the importance of the scene is scored based on the results. Scenes with higher importance are given higher priority and displayed on the user's device based on the generated playlist. This allows users to quickly check the efficiency of work in the factory and receive suggestions for improvement.

[0535] Prompt Sentence Examples

[0536] "Analyze videos recorded in a factory to determine which scenes are important and which parts can be made more efficient. Then, generate suggestions for improving efficiency for each scene."

[0537] In this way, a system that dramatically improves work efficiency within a factory can be realized.

[0538] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0539] Step 1:

[0540] The server receives the work videos taken by the user in the factory and stores the video data uploaded to the server.

[0541] Input: Video file uploaded from the user's device

[0542] Data processing: Save the video data on the server.

[0543] Output: Video data on the server

[0544] Step 2:

[0545] The server divides the video data into frames.

[0546] Input: Saved video data

[0547] Data processing: Video data is divided into individual frames using OpenCV.

[0548] Output: List of frame images

[0549] Step 3:

[0550] The server clusters each of the divided frames.

[0551] Input: List of frame images

[0552] Data operations: Use K-means clustering to group frames that belong to the same scene. Vectorize the frame images and cluster them in feature space.

[0553] Output: Clustered groups of frames

[0554] Step 4:

[0555] The server analyzes the clustered scenes and performs scoring.

[0556] Input: A group of clustered frames

[0557] Data calculation: Evaluate the importance of each scene through audio data analysis, text data analysis, and visual data analysis. Specifically, audio analysis uses the Google Cloud Speech-to-Text API, text analysis uses NLP (natural language processing) models, and visual data analysis applies machine learning algorithms.

[0558] Output: Importance score for each scene

[0559] Step 5:

[0560] The server prioritizes the scenes based on the scoring results.

[0561] Input: Importance score for each scene

[0562] Data calculation: sorting scenes in order of importance based on their scores.

[0563] Output: A prioritized list of scenes

[0564] Step 6:

[0565] The terminal generates a playlist based on the user's time-saving playback settings.

[0566] Input: A prioritized list of scenes and the user's time-saving playback settings

[0567] Data processing: Generate a playlist containing only important scenes, taking into account user preferences.

[0568] Output: Playlist

[0569] Step 7:

[0570] The terminal plays the videos based on the playlist.

[0571] Input: Playlist

[0572] Data processing: Play only the scenes included in the playlist in sequence.

[0573] Output: Video playback

[0574] Step 8:

[0575] The terminal updates the playlist in real time in response to user interactions.

[0576] Input: User actions (selecting a specific scene, skipping, etc.)

[0577] Data processing: Regenerate the playlist based on user actions to reflect updated information.

[0578] Output: Updated playlist and video playback changes

[0579] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[0580] This invention is a system that allows users to effectively shorten the time spent watching videos and efficiently watch only the interesting or important parts. This system consists of four main components: a server, a terminal, a user, and an emotion engine.

[0581] Server Operation

[0582] The server receives video data uploaded by users. After receiving the video data, the server divides the video into frames and clusters each frame into scenes. The clustered scenes are evaluated through audio data analysis, text data analysis, and visual data analysis, and each scene is scored for its interest and importance. The scored scenes are prioritized based on the evaluation results and stored in a database on the server.

[0583] Device behavior

[0584] When a user wants to watch a video, they first launch the application on their device and select a video. Users can set viewing preferences for time-saving playback, such as "play only the important parts for 50%" or "play only the interesting parts for 30%." These preferences are sent from the device to the server, which then obtains scoring information for prioritized scenes. The device generates a playlist based on this information and plays the video according to the preferences. If the user selects a specific scene or wants to view a skipped scene while watching, the playlist is updated in real time. The server reflects the user's actions, and a regenerated playlist is sent to the device.

[0585] User behavior

[0586] When a user wants to watch a video, they first upload the video file to the server. Next, they launch the application on their device and select the video they want to watch. The user then configures time-saving playback settings to suit their convenience and interests. Once playback settings are configured, the device automatically obtains scene scoring information from the server and generates a playlist. If the user wants to check a scene that was skipped during viewing or cancel time-saving playback, they can press a button to update the playlist in real time, allowing them to view content according to their needs.

[0587] Emotion Engine Operation

[0588] The device is equipped with an emotion engine that monitors the user's emotions while watching videos. The engine analyzes the user's facial expressions, tone of voice, and eye movements to recognize the user's emotional state in real time. The emotion data recognized by the emotion engine is fed back to the device and used to adjust the playlist.

[0589] For example, if a user is enjoying a scene, the score for that scene will increase, and similar scenes will be added to the list with priority. Emotional data will also be accumulated and used as learning data to reflect in future scoring and prioritization. This function allows video viewing to be more tailored to the user's preferences.

[0590] Specific examples

[0591] For example, consider a case where a user wants to watch a 30-minute talk show but only has 50% of the time. The user sets the device application to "play only 50% of the important parts." When the video is uploaded to the server, the server divides the video into frames and uses an AI model to score the importance of each scene. Based on the scoring results, the server prioritizes the scenes and sends them back to the device. Based on the settings, the device lists only the important scenes and plays them so that the user can watch efficiently.

[0592] While watching, the emotion engine monitors the user's reactions and, if they are enjoying the content, adjusts the priority of the corresponding scene in real time. If a scene that interests the user is skipped while watching, they can press a dedicated button to jump to that scene. This system not only allows users to efficiently view important information and interesting parts within a limited time, but also provides a viewing experience customized to their emotions.

[0593] The processing flow will be explained below.

[0594] Server Processing

[0595] Step 1:

[0596] The server receives video data from the user and stores it in temporary storage.

[0597] Step 2:

[0598] The server divides the video data into frames and converts the video into individual still images.

[0599] Step 3:

[0600] The server clusters frames by scene, which groups frames with similar characteristics together.

[0601] Step 4:

[0602] To evaluate each scene, the server performs audio data analysis, text data analysis, and visual data analysis, such as analyzing the speaker's content from the audio data and analyzing subtitles and descriptions from the text data.

[0603] Step 5:

[0604] The server scores each scene based on its interest and importance, and assigns a score to each scene based on the evaluation results.

[0605] Step 6:

[0606] The server prioritizes scenes based on the evaluated scores and stores the priority information and scoring results in a database.

[0607] Terminal handling

[0608] Step 1:

[0609] The device receives the user's time-saving playback settings, such as "play only the important parts at 50%."

[0610] Step 2:

[0611] The device retrieves scene scoring information, including prioritization information, from the server.

[0612] Step 3:

[0613] The device generates a playlist based on the user's settings, prioritizing important or interesting scenes.

[0614] Step 4:

[0615] The device plays videos based on the playlist, and unnecessary scenes are skipped as the user watches.

[0616] Step 5:

[0617] If the user wants to view a scene that was skipped during viewing, the device will accept a dedicated button and readjust the playlist based on that information.

[0618] Step 6:

[0619] The device sends the updated playlist to the server and continues playing the video based on the regenerated playlist.

[0620] User Action

[0621] Step 1:

[0622] The user selects a video and uploads it to the server.

[0623] Step 2:

[0624] The user launches the application on the device, selects the video they want to watch, and enters the time-saving playback settings.

[0625] Step 3:

[0626] The user watches a video that has been efficiently edited based on the settings.

[0627] Step 4:

[0628] If a scene of interest to the user is skipped while watching, the user can press a dedicated button to jump to that scene.

[0629] Emotion engine processing

[0630] Step 1:

[0631] The device's built-in emotion engine monitors the user's facial expressions, tone of voice, gaze, and other factors in real time.

[0632] Step 2:

[0633] The emotion engine analyzes the user's emotional state and generates emotion data based on that.

[0634] Step 3:

[0635] The generated emotion data is received by the device and used to adjust the playlist, for example by increasing the score of scenes that are enjoyed.

[0636] Step 4:

[0637] Emotional data is accumulated and sent to the server to be used as learning data for future scoring and prioritization.

[0638] Specific examples

[0639] For example, consider a case where a user wants to watch a 30-minute talk show but only has 50% of the time. The user sets the device application to "play only 50% of the important parts." When the video is uploaded to the server, the server divides the video into frames and uses an AI model to score the importance of each scene. Based on the scoring results, the server prioritizes the scenes and sends them back to the device. The device then lists only the important scenes based on the settings and plays them back so that the user can watch efficiently.

[0640] While watching, the emotion engine monitors the user's reactions and, if they are enjoying the content, adjusts the priority of the corresponding scene in real time. If a scene that interests the user is skipped while watching, they can press a dedicated button to jump to that scene. This system not only allows users to efficiently view important information and interesting parts within a limited time, but also provides a viewing experience that is customized according to their emotions.

[0641] Example 2

[0642] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0643] Conventional video viewing systems make it difficult for users to efficiently grasp important or interesting parts of a video. They also lack a means to meet the need to view only the most important information within a limited time. Furthermore, content optimization based on user emotions is not performed, preventing consistent satisfaction.

[0644] The identification process by the identification processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means. In this invention, the server includes a means for uploading video data to the information processing device, a means for dividing the video data into frames and classifying them by scene, a means for evaluating and scoring the interest and importance of each scene, a means for prioritizing the scenes based on the evaluated scores, a means for generating a playlist based on the user's playback settings and playing the video, a means for updating the playlist in response to the user's operations, and a means for recognizing the user's emotions during viewing and adjusting the playlist. This makes it possible to streamline the video viewing experience and provide customized playback according to the user's interests and emotions.

[0645] "Video data" means a digitally represented video file containing visual and audio information.

[0646] An "information processing device" is a computer system for processing, storing, and analyzing data.

[0647] A "user" is a person who operates the system and watches videos.

[0648] A "frame" is an individual still image that makes up a video.

[0649] A "scene" is a set of frames that represent a certain continuous time and space within a video.

[0650] "Classification" is the process of grouping frames that have the same characteristics or features.

[0651] "Scoring" is the process of quantifying the interest and importance of a scene based on certain criteria.

[0652] "Prioritization" means determining the order in which scenes are played based on the evaluated scores.

[0653] A "playlist" is a list that defines the order in which scenes are to be played.

[0654] "Playback" is the process of visually and aurally representing a moving image.

[0655] "Playback settings" are settings that specify the playback method and range of the video that the user wants to watch.

[0656] "Emotions" are the psychological and physiological reactions of users while watching.

[0657] "Recognition" refers to grasping a user's emotional state by analyzing their facial expressions, tone of voice, eye movements, etc.

[0658] "Adjusting" means changing scene priorities and playlists in real time.

[0659] "Operation" refers to the instructions or input that a user gives to the system.

[0660] This invention is a system that allows users to effectively shorten the time it takes to watch a video and efficiently watch only the interesting or important parts. This system consists of four main components: a server, a terminal, a user, and an emotion engine.

[0661] Server Operation

[0662] The server receives video data uploaded by users. Once the video data is received, the server divides the video into frames using FFmpeg. The divided frames are then clustered by scene using the FrameGenius API. The clustered scenes are analyzed for audio, text, and visual data using Google Cloud's Speech-to-Text API and NLP Cloud. Based on the analyzed data, the interestingness and importance of each scene is evaluated and scored using a proprietary scoring algorithm. The scenes are prioritized based on the scoring results, and the results are stored in a MongoDB database.

[0663] Device behavior

[0664] When a user wants to watch a video, they launch a dedicated application on their device and select the video they want to watch. Users can set time-saving playback settings such as "play only the important parts at 50%" or "play only the interesting parts at 30%." This setting information is sent from the device to the server, which then obtains scoring information for prioritized scenes. The device generates a playlist based on this information and plays the video using HLS (HTTP Live Streaming) technology. If the user selects a specific scene while watching or wants to watch a skipped scene, the playlist is updated in real time and the information is reflected on the server.

[0665] User actions

[0666] Users select the video they want to watch from their device and upload it to the server. Next, they launch a dedicated application on their device and select the video they want to watch. After selecting the video they want to watch, users can set up time-saving playback according to their interests and convenience. Once playback settings are set, the device automatically obtains scene scoring information from the server and generates a playlist. If users want to check scenes that were skipped during viewing or cancel time-saving playback, they can operate a dedicated button to update the playlist in real time, allowing them to watch according to their needs.

[0667] Emotion Engine Operation

[0668] The device is equipped with an emotion engine that monitors the user's emotions while watching videos using the Momento AI engine. This engine analyzes the user's facial expressions, vocal tone, eye movements, and other factors to recognize the user's emotional state in real time. The recognized emotion data is fed back to the device and used to adjust the playlist. For example, if the user is enjoying a scene very much, the score for that scene will increase, and similar scenes will be added to the playlist with priority. Emotion data is also accumulated and used as learning data to reflect in future scoring and prioritization. This function enables video viewing that is more tailored to the user's preferences.

[0669] Specific examples

[0670] For example, consider the case where a user wants to watch a 30-minute talk show but only has 15 minutes to watch. The user sets the device application to "play only the important parts at 50%." When the video is uploaded to the server, the server uses FFmpeg to divide the video into frames and clusters each scene using the FrameGenius API. The server then analyzes the video using Google Cloud's Speech-to-Text API and NLP Cloud to score its importance. Prioritization is performed based on the scoring results, and the results are sent to the device. The device then lists only the important scenes in a playlist and plays the video using HLS technology.

[0671] While watching, the emotional engine monitors the user's reactions using the Momento AI engine, and if the user is enjoying the video, the priority of the corresponding scene is adjusted in real time. Users can jump to the relevant scene by pressing the "Watch Skipped Scene" button. This system not only allows users to efficiently view important information and interesting parts within a limited time, but also provides a viewing experience customized to their emotions.

[0672] Prompt Sentence Examples

[0673] "I want to set up time-saving playback so that I can watch only the important parts of a 30-minute talk show video in 50% of the viewing time."

[0674] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0675] Step 1: Upload your video

[0676] The server receives video data from the user. The user selects the video they want to watch on their device and uploads it to the server using a dedicated application. Input: Video file sent from the user's device. Output: Video data saved on the server.

[0677] Specific behavior:

[0678] A user opens the application and clicks the "Upload Video" button.

[0679] The user selects the video they want to watch from the file selection dialog and clicks "Open."

[0680] The video file is sent to the server and the upload is complete.

[0681] Step 2: Video analysis and frame segmentation

[0682] The server splits the received video into frames using FFmpeg. Input: Video data stored on the server. Output: Still images for each frame.

[0683] Specific behavior:

[0684] The server splits the received video data into frames using the FFmpeg tool.

[0685] The split frames are saved in a temporary folder.

[0686] Step 3: Scene Clustering

[0687] The server uses the FrameGenius API to classify frames into scenes. Input: Segmented frame data. Output: Clustered scene data.

[0688] Specific behavior:

[0689] The FrameGenius API is used to analyze the visual characteristics of each frame.

[0690] Frames are grouped into clusters based on similarity and organized into scenes.

[0691] Step 4: Analyze and score the scene

[0692] The server uses Google Cloud's Speech-to-Text API and NLP Cloud to analyze the audio, text, and visual data for each scene. Input: Clustered scene data. Output: Analysis results and scoring data for each scene.

[0693] Specific behavior:

[0694] The audio data contained in each clustered scene is converted into text data using Google Cloud's Speech-to-Text API.

[0695] Text data is analyzed using NLP Cloud to extract keywords and emotions.

[0696] The visual data of each scene is also analyzed to evaluate its importance and interest.

[0697] Step 5: Prioritize and save your scenes

[0698] The server uses a proprietary algorithm to determine the priority of scenes based on the analysis results and stores the results in a MongoDB database. Input: Scoring data for each scene. Output: A prioritized scene list.

[0699] Specific behavior:

[0700] Scenes are sorted by priority based on the scoring data for each scene.

[0701] Store the prioritized scene list in a MongoDB database.

[0702] Step 6: Submit viewing preferences

[0703] The user configures the time-saving playback settings using the application on the device and sends this setting information to the server. Input: User's viewing settings. Output: Viewing setting data sent to the server.

[0704] Specific behavior:

[0705] The user opens the application and selects the video they want to watch.

[0706] Users can select viewing settings such as "play only the important parts at 50%" or "play only the interesting parts at 30%."

[0707] The configuration information is sent to the server.

[0708] Step 7: Generate and play a playlist

[0709] The server sends prioritized scene scoring information based on the received viewing preferences to the device, which then generates a playlist based on that information and plays the video using HLS technology. Input: User's viewing preferences, prioritized scene list. Output: Playlist, video to be played.

[0710] Specific behavior:

[0711] The server filters out priority scenes according to viewing settings and transmits them to the terminal.

[0712] The device uses HLS technology to generate a playlist and plays the video in scene order.

[0713] Step 8: Emotion Engine Feedback

[0714] The emotion engine built into the device monitors the user's emotional state in real time and adjusts the playlist. Input: Real-time emotion data while the user is watching. Output: Adjusted playlist.

[0715] Specific behavior:

[0716] The emotion engine analyzes the user's facial expressions, tone of voice, and eye movements.

[0717] Using the Momento AI engine, it recognizes the user's emotional state and provides feedback in real time.

[0718] Adjust the playlist to include scenes that users enjoy or find interesting.

[0719] (Application example 2)

[0720] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0721] While modern consumers have access to a wealth of video content, it is difficult for them to efficiently obtain the most interesting content and important information within a limited viewing time. Furthermore, the ability to adjust playback content in real time based on the viewer's emotions and reactions is also lacking. In response to this problem, the present invention aims to provide a system that provides a video viewing experience that maximizes user engagement within a limited time.

[0722] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[0723] In this invention, the server includes means for uploading video data to the server, means for dividing the video data into frames and clustering them by scene, means for evaluating and scoring the interest and importance of each scene, means for prioritizing the scenes based on the evaluated scores, means for generating a playlist based on a user's time-saving playback settings and playing the videos, means for updating the playlist in response to user interactions, and means for analyzing the user's facial expressions, tone of voice, and eye movements and adjusting the playlist based on emotion data. This allows users to efficiently watch only interesting or important parts even within a limited time, and further enables a personalized viewing experience based on emotions.

[0724] "Means for uploading video data to a server" refers to a function that allows a user to send video files from their own device to a server.

[0725] The "means of dividing video data into frames and clustering by scene" is a process of dividing a video uploaded to a server into frames and grouping similar frames.

[0726] The "means for evaluating and scoring the interest and importance of each scene" is a process of analyzing the clustered scenes and assigning a numerical evaluation of their interest and importance.

[0727] The "means for prioritizing scenes based on the evaluated scores" is a mechanism for sorting scored scenes in order of importance based on their scores.

[0728] "Means for generating a playlist based on the user's time-saving playback settings and playing videos" refers to a function that creates an optimized playlist based on the user's specified playback time and priority scene conditions, and plays videos according to that list.

[0729] The "means for updating the playlist in response to user interaction" refers to a system that dynamically changes the playlist based on the operations performed by the user while viewing (e.g., selecting or skipping a specific scene).

[0730] "Means of analyzing the user's facial expressions, tone of voice, and eye movements, and adjusting the playlist based on emotional data" refers to a function that analyzes the user's emotions and reactions while watching, and adjusts the playlist on the spot based on the results.

[0731] This invention provides a system that enables users to effectively shorten the time spent watching videos and efficiently watch only the interesting or important parts. This system consists of four main components: a server, a terminal, a user, and an emotion engine.

[0732] Server Operation

[0733] The server receives video data uploaded by users. After receiving the video data, the server divides the video into frames and clusters each frame into scenes. The clustered scenes are evaluated through audio data analysis, text data analysis, and visual data analysis to score each scene's interest and importance. The scenes are prioritized based on the evaluation results and stored in a database.

[0734] Device behavior

[0735] When a user wants to watch a video, they first launch the application on their device and select a video. Users can set viewing preferences for time-saving playback, such as "play only the important parts for 50%" or "play only the interesting parts for 30%." These preferences are sent from the device to the server, which then obtains scoring information for prioritized scenes. The device generates a playlist based on this information and plays the video according to the preferences. If the user selects a specific scene or wants to view a skipped scene while watching, the playlist is updated in real time. The user's actions are reflected on the server, and a regenerated playlist is sent to the device.

[0736] User behavior

[0737] When a user wants to watch a video, they first upload the video file to the server. Next, they launch the application on their device and select the video they want to watch. The user then configures time-saving playback settings to suit their convenience and interests. Once playback settings are configured, the device automatically obtains scene scoring information from the server and generates a playlist. If the user wants to check a scene that was skipped during viewing or cancel time-saving playback, they can press a button to update the playlist in real time, allowing them to watch according to their needs.

[0738] Emotion Engine Operation

[0739] The device is equipped with an emotion engine that monitors the user's emotions while watching videos. This engine analyzes the user's facial expressions, vocal tone, eye movements, and other factors to recognize the user's emotional state in real time. The emotion data recognized by the emotion engine is fed back to the device and used to adjust the playlist. For example, if the user is enjoying a video very much, the score for that scene will increase and similar scenes will be added to the playlist with priority. Emotion data is also accumulated and used as learning data to reflect in future scoring and prioritization. This function enables video viewing that is more tailored to the user's preferences.

[0740] Specific examples

[0741] Consider a scenario where a user wants to watch a 30-minute talk show but only has 50% of the time available. The user selects "Play only 50% of the important parts" in the device application. When the video is uploaded to the server, the server divides it into frames and uses a generative AI model to score the importance of each scene. Based on the scoring results, the server prioritizes the scenes and sends them back to the device. The device then lists only the important scenes based on the user's settings and plays them for efficient viewing. While watching, the emotion engine monitors the user's reactions. If the user is enjoying the video, the priority of the corresponding scene is adjusted in real time. If a scene that interests the user is skipped while watching, the user can press a dedicated button to jump back to that scene. This system not only allows users to efficiently view important information and interesting parts within a limited time, but also provides a customized viewing experience tailored to their emotions.

[0742] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0743] Step 1:

[0744] The user uploads a video file to the server. The user selects a video file using an application on their device and presses the upload button, sending the video data to the server. The input is the video file selected by the user, and the output is the video data saved on the server.

[0745] Step 2:

[0746] The server divides the uploaded video data into frames. When the server receives the video data, it divides the video into frames and stores each frame in memory. The input is video data, and the output is image data divided into frames.

[0747] Step 3:

[0748] The server clusters the image data for each frame by scene. The server uses a clustering algorithm to classify similar frames into the same scene cluster. The input is the divided frame image data, and the output is the clustered scene groups.

[0749] Step 4:

[0750] The server evaluates and scores the interest and importance of each scene group. The server performs audio, text, and visual analysis, and uses a generative AI model to evaluate the importance and interest of each scene as a numerical value. The input is the clustered scene data, and the output is the evaluation score.

[0751] Step 5:

[0752] The server prioritizes scenes based on the evaluated scores. The server prioritizes the playback order of each scene based on the scoring results. The input is the evaluation scores, and the output is a list of prioritized scene clusters.

[0753] Step 6:

[0754] The device sends the user's time-saving playback settings to the server and generates a playlist. The user configures playback settings within the application, and the device sends them to the server. The server creates a playlist based on the prioritized scene list and matches the user's settings. The input is the user's playback settings, and the output is an optimized playlist.

[0755] Step 7:

[0756] The device plays videos based on the generated playlist. The device receives the playlist sent from the server and plays videos according to that list. The input is the playlist and the output is the playback of the video.

[0757] Step 8:

[0758] The device updates the playlist according to the user's interactions. When the user skips or selects a specific scene while watching a video, the device sends that information to the server in real time and updates the playlist. The input is the user's interaction data, and the output is the updated playlist.

[0759] Step 9:

[0760] The device uses an emotion engine to monitor the user's emotional state and adjust the playlist accordingly. The emotion engine built into the device analyzes the user's facial expressions, tone of voice, and eye movements to obtain emotional data in real time. The obtained emotional data is sent to the server, which then adjusts the playlist based on that data. The input is the user's emotional data, and the output is a playlist adjusted based on the emotions.

[0761] Step 10:

[0762] The server sends the adjusted playlist to the terminal, and the terminal updates the playlist. The input is the adjusted playlist, and the output is the playlist update.

[0763] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0764] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0765] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.

[0766] [Third embodiment]

[0767] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.

[0768] 5, the data processing system 310 includes the data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.

[0769] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0770] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.

[0771] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0772] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0773] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0774] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0775] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0776] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0777] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0778] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."

[0779] This invention is a system that allows users to effectively shorten the time spent watching videos and efficiently watch only the interesting or important parts. This system consists of three main components: a server, a terminal, and a user.

[0780] Server Operation

[0781] The server receives video data uploaded by users. After receiving the video data, the server divides the video into frames and clusters each frame into scenes. The clustered scenes are evaluated through audio data analysis, text data analysis, and visual data analysis, and each scene is scored for its interest and importance. The scored scenes are prioritized based on the evaluation results and stored in a database on the server.

[0782] Device behavior

[0783] When a user wants to watch a video, they first launch the application on their device and select a video. Users can set viewing preferences for time-saving playback, such as "play only the important parts for 50%" or "play only the interesting parts for 30%." These preferences are sent from the device to the server, which then obtains scoring information for prioritized scenes. The device generates a playlist based on this information and plays the video according to the preferences. If the user selects a specific scene or wants to view a skipped scene while watching, the playlist is updated in real time. The server reflects the user's actions, and a regenerated playlist is sent to the device.

[0784] User behavior

[0785] When a user wants to watch a video, they first upload the video file to the server. Next, they launch the application on their device and select the video they want to watch. The user then configures time-saving playback settings to suit their convenience and interests. Once playback settings are configured, the device automatically obtains scene scoring information from the server and generates a playlist. If the user wants to check a scene that was skipped during viewing or cancel time-saving playback, they can press a button to update the playlist in real time, allowing them to view content according to their needs.

[0786] Specific examples

[0787] For example, consider a case where a user wants to watch a 30-minute talk show but only has 50% of the time. The user sets the device's application to "play only 50% of the important parts." When the video is uploaded to the server, the server divides the video into frames and uses an AI model to score the importance of each scene. Based on the scoring results, the server prioritizes the scenes and sends them back to the device. The device then lists only the important scenes based on the settings and plays them back so that the user can watch efficiently. If a scene of interest is skipped while watching, the user can press a dedicated button to jump to that scene. This system allows users to watch within a limited time without missing any important information.

[0788] The processing flow will be explained below.

[0789] Server Processing

[0790] Step 1:

[0791] The server receives the video data from the user and stores it in temporary storage.

[0792] Step 2:

[0793] The server divides the video data into frames, i.e., converts the video into multiple still images.

[0794] Step 3:

[0795] The server clusters frames by scene, which groups frames with similar characteristics together.

[0796] Step 4:

[0797] To evaluate each scene, the server analyzes audio data, text data, and visual data. For example, it analyzes the speaker's content from audio data and subtitles and descriptions from text data.

[0798] Step 5:

[0799] The server scores each scene based on its interest and importance, and assigns a score to each scene based on the evaluation results.

[0800] Step 6:

[0801] The server prioritizes scenes based on the evaluated scores, and the priority information and scoring results are stored in a database.

[0802] Terminal handling

[0803] Step 1:

[0804] The device receives the user's time-saving playback settings, such as "play only the important parts at 50%."

[0805] Step 2:

[0806] The device retrieves scene scoring information from the server, including scene priority information.

[0807] Step 3:

[0808] The device will generate a playlist based on the user's settings, prioritizing important or interesting scenes.

[0809] Step 4:

[0810] The device plays videos based on the playlist, and unnecessary scenes are skipped as the user watches.

[0811] Step 5:

[0812] If a user wants to view a scene that was skipped during viewing, the device will accept a dedicated button and will readjust the playlist based on that information.

[0813] Step 6:

[0814] The device sends the updated playlist to the server and continues playing the video based on the regenerated playlist.

[0815] User Action

[0816] Step 1:

[0817] The user selects a video and uploads it to the server.

[0818] Step 2:

[0819] The user launches the application on the device, selects the video they want to watch, and enters the time-saving playback settings.

[0820] Step 3:

[0821] The user watches a video that has been efficiently edited based on the settings.

[0822] Step 4:

[0823] If a scene of interest to the user is skipped while watching, the user can press a dedicated button to jump to that scene.

[0824] By following the above steps, the user can efficiently view only the important or interesting parts of the content, thereby saving time.

[0825] Example 1

[0826] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[0827] Previously, it was difficult to efficiently watch only the important or interesting parts of a long video. In particular, to obtain the information you wanted to watch within a limited time, you had to manually skip or fast-forward, which was inconvenient. Furthermore, there was no function to update the playlist in real time in response to user interactions, which meant that a flexible viewing experience was not provided.

[0828] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0829] In this invention, the server includes means for uploading video data to the server, means for dividing the video data into frames and clustering them by scene, means for evaluating and scoring the interest and importance of each scene, means for evaluating the scenes using audio data analysis, text data analysis, and visual data analysis, means for prioritizing the scenes based on the evaluated scores, means for generating a playlist based on a user's time-saving playback setting and playing the video, and means for updating the playlist in response to user interactions. This allows a user to efficiently view important or interesting parts within a limited time and update the playlist in real time.

[0830] "Video data" is data in the form of dynamic media that includes visual and audio information.

[0831] A "server" is a system that provides services to other computers and devices on a network and processes and stores data.

[0832] A "frame" refers to each still image that makes up a video, and when played back in succession, they form a moving image.

[0833] "Clustering" is a process of classifying data into groups with common characteristics, and is a technique used to divide scenes in videos.

[0834] A "scene" is a group of consecutive frames that are semantically consistent within a video, and is a part that expresses a story or event.

[0835] "Fun" refers to elements that interest and entertain viewers, and is one of the criteria for evaluating a scene or the entire video.

[0836] "Importance" is an index that indicates how valuable the content of a video is to the viewer, and is a criterion for determining the priority of scenes.

[0837] "Scoring" is the process of assigning a numerical score to each scene based on an evaluation standard.

[0838] "Audio data analysis" is a technology that converts the audio portion of a video into text and analyzes its content.

[0839] "Text data analysis" is a technique for analyzing the content of text and evaluating its meaning and importance.

[0840] "Visual data analysis" is a technology that analyzes the content of images and videos and identifies objects and people.

[0841] A "playlist" is a list created to determine the playback order and content of videos.

[0842] "Time-saving playback settings" are options that users can set to shorten the playback time of videos under certain conditions.

[0843] "User interaction" refers to the operations and inputs that a user makes to a system.

[0844] The present invention provides a system that enables users to efficiently watch videos. The system mainly comprises a server, a terminal, and a user.

[0845] The server receives video data uploaded by users. After receiving the video data, the server uses common video processing tools such as FFmpeg to divide the video into frames. For example, a 30-minute video at 30 FPS would be divided into 54,000 frames. Next, each frame is clustered into scenes. This process involves using OpenCV to detect changes in color and motion, and then using a clustering algorithm (e.g., K-means) to divide the scenes. Each scene is evaluated using audio data analysis (e.g., Google Cloud Speech-to-Text, which converts speech to text), text data analysis (e.g., OpenAI GPT-4, which analyzes the meaning of text), and visual data analysis (e.g., OpenCV, which identifies objects and people). Each scene is scored based on the evaluation results and stored in a database. This scoring information is used to generate a playlist when users watch the video.

[0846] The device provides the user with an interface to select the video they want to watch. The user can set time-saving playback settings, such as "play only the important parts for 50%" or "play only the interesting parts for 30%." These settings are sent to the server, which provides the device with data to generate a playlist based on the stored scoring information. The device generates a playlist based on this data and plays the video according to the settings selected by the user. If the user wants to select a specific scene or view a skipped scene while watching, the playlist is updated in real time and the information is synchronized with the server.

[0847] To watch a video, a user first uploads the video file to the server. Next, they launch a dedicated application on their device and select the video they want to watch. The user configures time-saving playback settings based on their interests and convenience. Once this configuration is complete, the server obtains the scene scoring information and sends it to the device. The device then generates a playlist based on this information and plays the video according to the settings. As the user interacts with the video while watching, the playlist is updated in real time, allowing for flexible viewing.

[0848] For example, suppose a user wants to watch a 30-minute talk show, but only has 15 minutes available. The user sets the device app to "play only 50% of the important parts." Use this prompt:

[0849] "I would like to upload the following 30-minute talk show video and play it based on the setting of 'play only 50% of the important parts'. Please rate the importance of each scene using a system that uses Google Cloud Speech-to-Text for audio analysis, OpenAI GPT-4 for text analysis, and OpenCV for visual analysis. Based on the evaluation results, please make a list of only the important scenes, so that it can be viewed in 15 minutes."

[0850] As described above, the present invention provides a system that allows users to efficiently view important or interesting parts within a limited time. The advanced analysis technology of the server and the user-friendly interface of the terminal realize a flexible and effective viewing experience.

[0851] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0852] Step 1:

[0853] The server receives the video data. The user selects a video file from the device and uploads it to the server. The input is the video file provided by the user, and the output is a video file saved in a temporary directory on the server. At this stage, data is sent and received using the HTTP protocol.

[0854] Step 2:

[0855] The server splits the video data into frames. The input is the video file saved in step 1, and the output is a set of still images for each frame. This process uses FFmpeg to split the video based on the number of seconds and frame rate. Specifically, a 30-minute video (30 FPS) will generate 54,000 still frames.

[0856] Step 3:

[0857] The server clusters each frame by scene. The input is the set of still images generated in step 2, and the output is a list of frames clustered by scene. OpenCV is used to analyze the color and motion of the images, and a clustering algorithm (e.g., K-means) is used to divide the scenes. Specifically, identical scenes are grouped together based on the similarity of the frames.

[0858] Step 4:

[0859] The server evaluates the scenes. The input is the list of scenes clustered in Step 3, and the output is an evaluation score for each scene. Google Cloud Speech-to-Text is used for audio data analysis, OpenAI GPT-4 for text data analysis, and OpenCV is used again for visual data analysis. Specifically, the server converts audio into text, analyzes its content, and then analyzes visual information to evaluate the importance and interest of the scene.

[0860] Step 5:

[0861] The server scores and prioritizes scenes based on the evaluation results. The input is the evaluation data obtained in step 4, and the output is a score list for each scene. Scoring is performed on a scale of 0 to 100, and scenes with higher scores have a higher viewing priority. Specifically, the evaluation score for each scene is quantified, and the scenes are sorted based on that.

[0862] Step 6:

[0863] The server saves the scored scenes in a database. The input is a list of scenes scored in step 5, and the output is the scoring information saved in the database. Specifically, data such as the scene frame number, start and end time, and score are stored in the database.

[0864] Step 7:

[0865] The user sets viewing preferences on the device. The user launches a dedicated application and selects the video they want to watch. The input is the user's viewing preferences (e.g., "play only the important parts at 50%), and the output is the viewing preferences data sent to the server. This setting is made via the application's UI.

[0866] Step 8:

[0867] The device retrieves the scoring information from the server. The input is the viewing preference data sent in step 7, and the output is the scoring information received from the server. The server filters the scoring information based on the user-specified preferences and sends it back to the device.

[0868] Step 9:

[0869] The device generates a playlist. The input is the scoring information obtained in step 8, and the output is the generated playlist. The playlist is ordered by scenes with the highest scores based on the user's viewing preferences, and is adjusted to match the configured playback percentage.

[0870] Step 10:

[0871] As the user watches the video, the playlist is updated in real time. The input is the user's viewing actions and interactions, and the output is the updated playlist and the video being played. If the user wants to select a specific scene or rewatch a skipped scene, they can do so using the device interface, and the content is synchronized with the server in real time. This allows users to flexibly customize their viewing experience.

[0872] (Application example 1)

[0873] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[0874] Currently, many factories are taking various measures to improve work efficiency, but optimal suggestions and real-time improvements are still often done manually. Therefore, there is a need for a system that can automatically generate suggestions for efficiency and quickly identify and play back important scenes in videos. Currently, there is no effective method to solve this problem.

[0875] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0876] In this invention, the server includes a means for uploading video data to the server, a means for dividing the video data into frames and clustering them by scene, a means for evaluating and scoring the interest and importance of each scene, a means for prioritizing the scenes based on the evaluated scores, and a means for analyzing work videos in an industrial environment and generating suggestions for improving efficiency, thereby enabling automatic and effective improvement of work efficiency in a factory.

[0877] "Video data" means digital files containing moving images and audio.

[0878] "Uploading" is the act of a user sending data from their device to a server.

[0879] A "scene" is a part of a video that consists of a certain series of frames and represents a series of movements or situations.

[0880] "Clustering" is a data processing technique that classifies many frames into groups with similar attributes.

[0881] "Evaluation" is the act of quantitatively or qualitatively judging the importance and interest of each scene.

[0882] "Scoring" is the process of assigning a numerical value to each scene based on the evaluation results.

[0883] "Prioritization" is the act of ranking scored scenes according to their importance and interest.

[0884] "Industrial environment" means the place or conditions in which manufacturing or production takes place.

[0885] "Efficiency" refers to reducing waste in work and processes and improving productivity.

[0886] A "proposal" refers to a solution or improvement to a specific problem.

[0887] System Overview

[0888] This invention is a video analysis system for improving work efficiency in an industrial environment. The system mainly includes a means for uploading video data to a server, a means for dividing the video data into frames and clustering them by scene, a means for evaluating and scoring the interest and importance of each scene, a means for prioritizing scenes based on the evaluated scores, and a means for analyzing work videos in an industrial environment and generating suggestions for improving efficiency.

[0889] Hardware and software used

[0890] Hardware: Robots used in factories (e.g. drones for monitoring work)

[0891] Software: Python, OpenCV, scikit-learn, cloud AI services (e.g., Google Cloud AI)

[0892] Data processing and calculation process

[0893] Server behavior:

[0894] 1. Uploading video data: Users upload videos of work they have taken in the factory to the server. The uploaded video data is stored on the server.

[0895] 2. Video Frame Segmentation: The server splits the received video data into frames. OpenCV is used to split the video data into individual frames.

[0896] 3. Frame Clustering: Each frame is then classified into scenes using K-means clustering, which groups similar scenes together.

[0897] 4. Scene Scoring: The clustered scenes are evaluated through audio, text, and visual data analysis. Based on the evaluation results, the importance of each scene is scored. Specifically, a machine learning algorithm is applied using scikit-learn.

[0898] 5. Prioritization: Based on the scoring results, scenes are prioritized in order of importance.

[0899] Terminal behavior:

[0900] 1. Playlist generation: The user uses the device application to select videos and set time-saving playback settings. Based on this, the device retrieves prioritized scene information from the server and generates a playlist.

[0901] 2. Work efficiency suggestions: Based on the generated playlist, suggestions for work efficiency are displayed, allowing users to quickly review only the important scenes.

[0902] 3. Updates based on user interaction: The playlist is updated in real time based on user actions.

[0903] Specific examples

[0904] For example, consider a video recording of assembly work in a factory. When a user uploads the video to the server, the server divides the video into frames and classifies them into scenes using K-means clustering. Each scene is evaluated through audio data analysis, text data analysis, and visual data analysis, and the importance of the scene is scored based on the results. Scenes with higher importance are given higher priority and displayed on the user's device based on the generated playlist. This allows users to quickly check the efficiency of work in the factory and receive suggestions for improvement.

[0905] Prompt Sentence Examples

[0906] "Analyze videos recorded in a factory to determine which scenes are important and which parts can be made more efficient. Then, generate suggestions for improving efficiency for each scene."

[0907] In this way, a system that dramatically improves work efficiency within a factory can be realized.

[0908] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0909] Step 1:

[0910] The server receives the work videos taken by the user in the factory and stores the video data uploaded to the server.

[0911] Input: Video file uploaded from the user's device

[0912] Data processing: Save the video data on the server.

[0913] Output: Video data on the server

[0914] Step 2:

[0915] The server divides the video data into frames.

[0916] Input: Saved video data

[0917] Data processing: Video data is divided into individual frames using OpenCV.

[0918] Output: List of frame images

[0919] Step 3:

[0920] The server clusters each of the divided frames.

[0921] Input: List of frame images

[0922] Data operations: Use K-means clustering to group frames that belong to the same scene. Vectorize the frame images and cluster them in feature space.

[0923] Output: Clustered groups of frames

[0924] Step 4:

[0925] The server analyzes the clustered scenes and performs scoring.

[0926] Input: A group of clustered frames

[0927] Data calculation: Evaluate the importance of each scene through audio data analysis, text data analysis, and visual data analysis. Specifically, audio analysis uses the Google Cloud Speech-to-Text API, text analysis uses NLP (natural language processing) models, and visual data analysis applies machine learning algorithms.

[0928] Output: Importance score for each scene

[0929] Step 5:

[0930] The server prioritizes the scenes based on the scoring results.

[0931] Input: Importance score for each scene

[0932] Data calculation: sorting scenes in order of importance based on their scores.

[0933] Output: A prioritized list of scenes

[0934] Step 6:

[0935] The terminal generates a playlist based on the user's time-saving playback settings.

[0936] Input: A prioritized list of scenes and the user's time-saving playback settings

[0937] Data processing: Generate a playlist containing only important scenes, taking into account user preferences.

[0938] Output: Playlist

[0939] Step 7:

[0940] The terminal plays the videos based on the playlist.

[0941] Input: Playlist

[0942] Data processing: Play only the scenes included in the playlist in sequence.

[0943] Output: Video playback

[0944] Step 8:

[0945] The terminal updates the playlist in real time in response to user interactions.

[0946] Input: User actions (selecting a specific scene, skipping, etc.)

[0947] Data processing: Regenerate the playlist based on user actions to reflect updated information.

[0948] Output: Updated playlist and video playback changes

[0949] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[0950] This invention is a system that allows users to effectively shorten the time spent watching videos and efficiently watch only the interesting or important parts. This system consists of four main components: a server, a terminal, a user, and an emotion engine.

[0951] Server Operation

[0952] The server receives video data uploaded by users. After receiving the video data, the server divides the video into frames and clusters each frame into scenes. The clustered scenes are evaluated through audio data analysis, text data analysis, and visual data analysis, and each scene is scored for its interest and importance. The scored scenes are prioritized based on the evaluation results and stored in a database on the server.

[0953] Device behavior

[0954] When a user wants to watch a video, they first launch the application on their device and select a video. Users can set viewing preferences for time-saving playback, such as "play only the important parts for 50%" or "play only the interesting parts for 30%." These preferences are sent from the device to the server, which then obtains scoring information for prioritized scenes. The device generates a playlist based on this information and plays the video according to the preferences. If the user selects a specific scene or wants to view a skipped scene while watching, the playlist is updated in real time. The server reflects the user's actions, and a regenerated playlist is sent to the device.

[0955] User behavior

[0956] When a user wants to watch a video, they first upload the video file to the server. Next, they launch the application on their device and select the video they want to watch. The user then configures time-saving playback settings to suit their convenience and interests. Once playback settings are configured, the device automatically obtains scene scoring information from the server and generates a playlist. If the user wants to check a scene that was skipped during viewing or cancel time-saving playback, they can press a button to update the playlist in real time, allowing them to view content according to their needs.

[0957] Emotion Engine Operation

[0958] The device is equipped with an emotion engine that monitors the user's emotions while watching videos. The engine analyzes the user's facial expressions, tone of voice, and eye movements to recognize the user's emotional state in real time. The emotion data recognized by the emotion engine is fed back to the device and used to adjust the playlist.

[0959] For example, if a user is enjoying a scene, the score for that scene will increase, and similar scenes will be added to the list with priority. Emotional data will also be accumulated and used as learning data to reflect in future scoring and prioritization. This function allows video viewing to be more tailored to the user's preferences.

[0960] Specific examples

[0961] For example, consider a case where a user wants to watch a 30-minute talk show but only has 50% of the time. The user sets the device application to "play only 50% of the important parts." When the video is uploaded to the server, the server divides the video into frames and uses an AI model to score the importance of each scene. Based on the scoring results, the server prioritizes the scenes and sends them back to the device. Based on the settings, the device lists only the important scenes and plays them so that the user can watch efficiently.

[0962] While watching, the emotion engine monitors the user's reactions and, if they are enjoying the content, adjusts the priority of the corresponding scene in real time. If a scene that interests the user is skipped while watching, they can press a dedicated button to jump to that scene. This system not only allows users to efficiently view important information and interesting parts within a limited time, but also provides a viewing experience customized to their emotions.

[0963] The processing flow will be explained below.

[0964] Server Processing

[0965] Step 1:

[0966] The server receives video data from the user and stores it in temporary storage.

[0967] Step 2:

[0968] The server divides the video data into frames and converts the video into individual still images.

[0969] Step 3:

[0970] The server clusters frames by scene, which groups frames with similar characteristics together.

[0971] Step 4:

[0972] To evaluate each scene, the server performs audio data analysis, text data analysis, and visual data analysis, such as analyzing the speaker's content from the audio data and analyzing subtitles and descriptions from the text data.

[0973] Step 5:

[0974] The server scores each scene based on its interest and importance, and assigns a score to each scene based on the evaluation results.

[0975] Step 6:

[0976] The server prioritizes scenes based on the evaluated scores and stores the priority information and scoring results in a database.

[0977] Terminal handling

[0978] Step 1:

[0979] The device receives the user's time-saving playback settings, such as "play only the important parts at 50%."

[0980] Step 2:

[0981] The device retrieves scene scoring information, including prioritization information, from the server.

[0982] Step 3:

[0983] The device generates a playlist based on the user's settings, prioritizing important or interesting scenes.

[0984] Step 4:

[0985] The device plays videos based on the playlist, and unnecessary scenes are skipped as the user watches.

[0986] Step 5:

[0987] If the user wants to view a scene that was skipped during viewing, the device will accept a dedicated button and readjust the playlist based on that information.

[0988] Step 6:

[0989] The device sends the updated playlist to the server and continues playing the video based on the regenerated playlist.

[0990] User Action

[0991] Step 1:

[0992] The user selects a video and uploads it to the server.

[0993] Step 2:

[0994] The user launches the application on the device, selects the video they want to watch, and enters the time-saving playback settings.

[0995] Step 3:

[0996] The user watches a video that has been efficiently edited based on the settings.

[0997] Step 4:

[0998] If a scene of interest to the user is skipped while watching, the user can press a dedicated button to jump to that scene.

[0999] Emotion engine processing

[1000] Step 1:

[1001] The device's built-in emotion engine monitors the user's facial expressions, tone of voice, gaze, and other factors in real time.

[1002] Step 2:

[1003] The emotion engine analyzes the user's emotional state and generates emotion data based on that.

[1004] Step 3:

[1005] The generated emotion data is received by the device and used to adjust the playlist, for example by increasing the score of scenes that are enjoyed.

[1006] Step 4:

[1007] Emotional data is accumulated and sent to the server to be used as learning data for future scoring and prioritization.

[1008] Specific examples

[1009] For example, consider a case where a user wants to watch a 30-minute talk show but only has 50% of the time. The user sets the device application to "play only 50% of the important parts." When the video is uploaded to the server, the server divides the video into frames and uses an AI model to score the importance of each scene. Based on the scoring results, the server prioritizes the scenes and sends them back to the device. The device then lists only the important scenes based on the settings and plays them back so that the user can watch efficiently.

[1010] While watching, the emotion engine monitors the user's reactions and, if they are enjoying the content, adjusts the priority of the corresponding scene in real time. If a scene that interests the user is skipped while watching, they can press a dedicated button to jump to that scene. This system not only allows users to efficiently view important information and interesting parts within a limited time, but also provides a viewing experience that is customized according to their emotions.

[1011] Example 2

[1012] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1013] Conventional video viewing systems make it difficult for users to efficiently grasp important or interesting parts of a video. They also lack a means to meet the need to view only the most important information within a limited time. Furthermore, content optimization based on user emotions is not performed, preventing consistent satisfaction.

[1014] The identification process by the identification processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means. In this invention, the server includes a means for uploading video data to the information processing device, a means for dividing the video data into frames and classifying them by scene, a means for evaluating and scoring the interest and importance of each scene, a means for prioritizing the scenes based on the evaluated scores, a means for generating a playlist based on the user's playback settings and playing the video, a means for updating the playlist in response to the user's operations, and a means for recognizing the user's emotions during viewing and adjusting the playlist. This makes it possible to streamline the video viewing experience and provide customized playback according to the user's interests and emotions.

[1015] "Video data" means a digitally represented video file containing visual and audio information.

[1016] An "information processing device" is a computer system for processing, storing, and analyzing data.

[1017] A "user" is a person who operates the system and watches videos.

[1018] A "frame" is an individual still image that makes up a video.

[1019] A "scene" is a set of frames that represent a certain continuous time and space within a video.

[1020] "Classification" is the process of grouping frames that have the same characteristics or features.

[1021] "Scoring" is the process of quantifying the interest and importance of a scene based on certain criteria.

[1022] "Prioritization" means determining the order in which scenes are played based on the evaluated scores.

[1023] A "playlist" is a list that defines the order in which scenes are to be played.

[1024] "Playback" is the process of visually and aurally representing a moving image.

[1025] "Playback settings" are settings that specify the playback method and range of the video that the user wants to watch.

[1026] "Emotions" are the psychological and physiological reactions of users while watching.

[1027] "Recognition" refers to grasping a user's emotional state by analyzing their facial expressions, tone of voice, eye movements, etc.

[1028] "Adjusting" means changing scene priorities and playlists in real time.

[1029] "Operation" refers to the instructions or input that a user gives to the system.

[1030] This invention is a system that allows users to effectively shorten the time it takes to watch a video and efficiently watch only the interesting or important parts. This system consists of four main components: a server, a terminal, a user, and an emotion engine.

[1031] Server Operation

[1032] The server receives video data uploaded by users. Once the video data is received, the server divides the video into frames using FFmpeg. The divided frames are then clustered by scene using the FrameGenius API. The clustered scenes are analyzed for audio, text, and visual data using Google Cloud's Speech-to-Text API and NLP Cloud. Based on the analyzed data, the interestingness and importance of each scene is evaluated and scored using a proprietary scoring algorithm. The scenes are prioritized based on the scoring results, and the results are stored in a MongoDB database.

[1033] Device behavior

[1034] When a user wants to watch a video, they launch a dedicated application on their device and select the video they want to watch. Users can set time-saving playback settings such as "play only the important parts at 50%" or "play only the interesting parts at 30%." This setting information is sent from the device to the server, which then obtains scoring information for prioritized scenes. The device generates a playlist based on this information and plays the video using HLS (HTTP Live Streaming) technology. If the user selects a specific scene while watching or wants to watch a skipped scene, the playlist is updated in real time and the information is reflected on the server.

[1035] User actions

[1036] Users select the video they want to watch from their device and upload it to the server. Next, they launch a dedicated application on their device and select the video they want to watch. After selecting the video they want to watch, users can set up time-saving playback according to their interests and convenience. Once playback settings are set, the device automatically obtains scene scoring information from the server and generates a playlist. If users want to check scenes that were skipped during viewing or cancel time-saving playback, they can operate a dedicated button to update the playlist in real time, allowing them to watch according to their needs.

[1037] Emotion Engine Operation

[1038] The device is equipped with an emotion engine that monitors the user's emotions while watching videos using the Momento AI engine. This engine analyzes the user's facial expressions, vocal tone, eye movements, and other factors to recognize the user's emotional state in real time. The recognized emotion data is fed back to the device and used to adjust the playlist. For example, if the user is enjoying a scene very much, the score for that scene will increase, and similar scenes will be added to the playlist with priority. Emotion data is also accumulated and used as learning data to reflect in future scoring and prioritization. This function enables video viewing that is more tailored to the user's preferences.

[1039] Specific examples

[1040] For example, consider the case where a user wants to watch a 30-minute talk show but only has 15 minutes to watch. The user sets the device application to "play only the important parts at 50%." When the video is uploaded to the server, the server uses FFmpeg to divide the video into frames and clusters each scene using the FrameGenius API. The server then analyzes the video using Google Cloud's Speech-to-Text API and NLP Cloud to score its importance. Prioritization is performed based on the scoring results, and the results are sent to the device. The device then lists only the important scenes in a playlist and plays the video using HLS technology.

[1041] While watching, the emotional engine monitors the user's reactions using the Momento AI engine, and if the user is enjoying the video, the priority of the corresponding scene is adjusted in real time. Users can jump to the relevant scene by pressing the "Watch Skipped Scene" button. This system not only allows users to efficiently view important information and interesting parts within a limited time, but also provides a viewing experience customized to their emotions.

[1042] Prompt Sentence Examples

[1043] "I want to set up time-saving playback so that I can watch only the important parts of a 30-minute talk show video in 50% of the viewing time."

[1044] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1045] Step 1: Upload your video

[1046] The server receives video data from the user. The user selects the video they want to watch on their device and uploads it to the server using a dedicated application. Input: Video file sent from the user's device. Output: Video data saved on the server.

[1047] Specific behavior:

[1048] A user opens the application and clicks the "Upload Video" button.

[1049] The user selects the video they want to watch from the file selection dialog and clicks "Open."

[1050] The video file is sent to the server and the upload is complete.

[1051] Step 2: Video analysis and frame segmentation

[1052] The server splits the received video into frames using FFmpeg. Input: Video data stored on the server. Output: Still images for each frame.

[1053] Specific behavior:

[1054] The server splits the received video data into frames using the FFmpeg tool.

[1055] The split frames are saved in a temporary folder.

[1056] Step 3: Scene Clustering

[1057] The server uses the FrameGenius API to classify frames into scenes. Input: Segmented frame data. Output: Clustered scene data.

[1058] Specific behavior:

[1059] The FrameGenius API is used to analyze the visual characteristics of each frame.

[1060] Frames are grouped into clusters based on similarity and organized into scenes.

[1061] Step 4: Analyze and score the scene

[1062] The server uses Google Cloud's Speech-to-Text API and NLP Cloud to analyze the audio, text, and visual data for each scene. Input: Clustered scene data. Output: Analysis results and scoring data for each scene.

[1063] Specific behavior:

[1064] The audio data contained in each clustered scene is converted into text data using Google Cloud's Speech-to-Text API.

[1065] Text data is analyzed using NLP Cloud to extract keywords and emotions.

[1066] The visual data of each scene is also analyzed to evaluate its importance and interest.

[1067] Step 5: Prioritize and save your scenes

[1068] The server uses a proprietary algorithm to determine the priority of scenes based on the analysis results and stores the results in a MongoDB database. Input: Scoring data for each scene. Output: A prioritized scene list.

[1069] Specific behavior:

[1070] Scenes are sorted by priority based on the scoring data for each scene.

[1071] Store the prioritized scene list in a MongoDB database.

[1072] Step 6: Submit viewing preferences

[1073] The user configures the time-saving playback settings using the application on the device and sends this setting information to the server. Input: User's viewing settings. Output: Viewing setting data sent to the server.

[1074] Specific behavior:

[1075] The user opens the application and selects the video they want to watch.

[1076] Users can select viewing settings such as "play only the important parts at 50%" or "play only the interesting parts at 30%."

[1077] The configuration information is sent to the server.

[1078] Step 7: Generate and play a playlist

[1079] The server sends prioritized scene scoring information based on the received viewing preferences to the device, which then generates a playlist based on that information and plays the video using HLS technology. Input: User's viewing preferences, prioritized scene list. Output: Playlist, video to be played.

[1080] Specific behavior:

[1081] The server filters out priority scenes according to viewing settings and transmits them to the terminal.

[1082] The device uses HLS technology to generate a playlist and plays the video in scene order.

[1083] Step 8: Emotion Engine Feedback

[1084] The emotion engine built into the device monitors the user's emotional state in real time and adjusts the playlist. Input: Real-time emotion data while the user is watching. Output: Adjusted playlist.

[1085] Specific behavior:

[1086] The emotion engine analyzes the user's facial expressions, tone of voice, and eye movements.

[1087] Using the Momento AI engine, it recognizes the user's emotional state and provides feedback in real time.

[1088] Adjust the playlist to include scenes that users enjoy or find interesting.

[1089] (Application example 2)

[1090] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1091] While modern consumers have access to a wealth of video content, it is difficult for them to efficiently obtain the most interesting content and important information within a limited viewing time. Furthermore, the ability to adjust playback content in real time based on the viewer's emotions and reactions is also lacking. In response to this problem, the present invention aims to provide a system that provides a video viewing experience that maximizes user engagement within a limited time.

[1092] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[1093] In this invention, the server includes means for uploading video data to the server, means for dividing the video data into frames and clustering them by scene, means for evaluating and scoring the interest and importance of each scene, means for prioritizing the scenes based on the evaluated scores, means for generating a playlist based on a user's time-saving playback settings and playing the videos, means for updating the playlist in response to user interactions, and means for analyzing the user's facial expressions, tone of voice, and eye movements and adjusting the playlist based on emotion data. This allows users to efficiently watch only interesting or important parts even within a limited time, and further enables a personalized viewing experience based on emotions.

[1094] "Means for uploading video data to a server" refers to a function that allows a user to send video files from their own device to a server.

[1095] The "means of dividing video data into frames and clustering by scene" is a process of dividing a video uploaded to a server into frames and grouping similar frames.

[1096] The "means for evaluating and scoring the interest and importance of each scene" is a process of analyzing the clustered scenes and assigning a numerical evaluation of their interest and importance.

[1097] The "means for prioritizing scenes based on the evaluated scores" is a mechanism for sorting scored scenes in order of importance based on their scores.

[1098] "Means for generating a playlist based on the user's time-saving playback settings and playing videos" refers to a function that creates an optimized playlist based on the user's specified playback time and priority scene conditions, and plays videos according to that list.

[1099] The "means for updating the playlist in response to user interaction" refers to a system that dynamically changes the playlist based on the operations performed by the user while viewing (e.g., selecting or skipping a specific scene).

[1100] "Means of analyzing the user's facial expressions, tone of voice, and eye movements, and adjusting the playlist based on emotional data" refers to a function that analyzes the user's emotions and reactions while watching, and adjusts the playlist on the spot based on the results.

[1101] This invention provides a system that enables users to effectively shorten the time spent watching videos and efficiently watch only the interesting or important parts. This system consists of four main components: a server, a terminal, a user, and an emotion engine.

[1102] Server Operation

[1103] The server receives video data uploaded by users. After receiving the video data, the server divides the video into frames and clusters each frame into scenes. The clustered scenes are evaluated through audio data analysis, text data analysis, and visual data analysis to score each scene's interest and importance. The scenes are prioritized based on the evaluation results and stored in a database.

[1104] Device behavior

[1105] When a user wants to watch a video, they first launch the application on their device and select a video. Users can set viewing preferences for time-saving playback, such as "play only the important parts for 50%" or "play only the interesting parts for 30%." These preferences are sent from the device to the server, which then obtains scoring information for prioritized scenes. The device generates a playlist based on this information and plays the video according to the preferences. If the user selects a specific scene or wants to view a skipped scene while watching, the playlist is updated in real time. The user's actions are reflected on the server, and a regenerated playlist is sent to the device.

[1106] User behavior

[1107] When a user wants to watch a video, they first upload the video file to the server. Next, they launch the application on their device and select the video they want to watch. The user then configures time-saving playback settings to suit their convenience and interests. Once playback settings are configured, the device automatically obtains scene scoring information from the server and generates a playlist. If the user wants to check a scene that was skipped during viewing or cancel time-saving playback, they can press a button to update the playlist in real time, allowing them to watch according to their needs.

[1108] Emotion Engine Operation

[1109] The device is equipped with an emotion engine that monitors the user's emotions while watching videos. This engine analyzes the user's facial expressions, vocal tone, eye movements, and other factors to recognize the user's emotional state in real time. The emotion data recognized by the emotion engine is fed back to the device and used to adjust the playlist. For example, if the user is enjoying a video very much, the score for that scene will increase and similar scenes will be added to the playlist with priority. Emotion data is also accumulated and used as learning data to reflect in future scoring and prioritization. This function enables video viewing that is more tailored to the user's preferences.

[1110] Specific examples

[1111] Consider a scenario where a user wants to watch a 30-minute talk show but only has 50% of the time available. The user selects "Play only 50% of the important parts" in the device application. When the video is uploaded to the server, the server divides it into frames and uses a generative AI model to score the importance of each scene. Based on the scoring results, the server prioritizes the scenes and sends them back to the device. The device then lists only the important scenes based on the user's settings and plays them for efficient viewing. While watching, the emotion engine monitors the user's reactions. If the user is enjoying the video, the priority of the corresponding scene is adjusted in real time. If a scene that interests the user is skipped while watching, the user can press a dedicated button to jump back to that scene. This system not only allows users to efficiently view important information and interesting parts within a limited time, but also provides a customized viewing experience tailored to their emotions.

[1112] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1113] Step 1:

[1114] The user uploads a video file to the server. The user selects a video file using an application on their device and presses the upload button, sending the video data to the server. The input is the video file selected by the user, and the output is the video data saved on the server.

[1115] Step 2:

[1116] The server divides the uploaded video data into frames. When the server receives the video data, it divides the video into frames and stores each frame in memory. The input is video data, and the output is image data divided into frames.

[1117] Step 3:

[1118] The server clusters the image data for each frame by scene. The server uses a clustering algorithm to classify similar frames into the same scene cluster. The input is the divided frame image data, and the output is the clustered scene groups.

[1119] Step 4:

[1120] The server evaluates and scores the interest and importance of each scene group. The server performs audio, text, and visual analysis, and uses a generative AI model to evaluate the importance and interest of each scene as a numerical value. The input is the clustered scene data, and the output is the evaluation score.

[1121] Step 5:

[1122] The server prioritizes scenes based on the evaluated scores. The server prioritizes the playback order of each scene based on the scoring results. The input is the evaluation scores, and the output is a list of prioritized scene clusters.

[1123] Step 6:

[1124] The device sends the user's time-saving playback settings to the server and generates a playlist. The user configures playback settings within the application, and the device sends them to the server. The server creates a playlist based on the prioritized scene list and matches the user's settings. The input is the user's playback settings, and the output is an optimized playlist.

[1125] Step 7:

[1126] The device plays videos based on the generated playlist. The device receives the playlist sent from the server and plays videos according to that list. The input is the playlist and the output is the playback of the video.

[1127] Step 8:

[1128] The device updates the playlist according to the user's interactions. When the user skips or selects a specific scene while watching a video, the device sends that information to the server in real time and updates the playlist. The input is the user's interaction data, and the output is the updated playlist.

[1129] Step 9:

[1130] The device uses an emotion engine to monitor the user's emotional state and adjust the playlist accordingly. The emotion engine built into the device analyzes the user's facial expressions, tone of voice, and eye movements to obtain emotional data in real time. The obtained emotional data is sent to the server, which then adjusts the playlist based on that data. The input is the user's emotional data, and the output is a playlist adjusted based on the emotions.

[1131] Step 10:

[1132] The server sends the adjusted playlist to the terminal, and the terminal updates the playlist. The input is the adjusted playlist, and the output is the playlist update.

[1133] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[1134] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1135] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.

[1136] [Fourth embodiment]

[1137] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.

[1138] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[1139] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[1140] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.

[1141] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[1142] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[1143] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[1144] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.

[1145] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[1146] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[1147] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[1148] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[1149] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1150] This invention is a system that allows users to effectively shorten the time spent watching videos and efficiently watch only the interesting or important parts. This system consists of three main components: a server, a terminal, and a user.

[1151] Server Operation

[1152] The server receives video data uploaded by users. After receiving the video data, the server divides the video into frames and clusters each frame into scenes. The clustered scenes are evaluated through audio data analysis, text data analysis, and visual data analysis, and each scene is scored for its interest and importance. The scored scenes are prioritized based on the evaluation results and stored in a database on the server.

[1153] Device behavior

[1154] When a user wants to watch a video, they first launch the application on their device and select a video. Users can set viewing preferences for time-saving playback, such as "play only the important parts for 50%" or "play only the interesting parts for 30%." These preferences are sent from the device to the server, which then obtains scoring information for prioritized scenes. The device generates a playlist based on this information and plays the video according to the preferences. If the user selects a specific scene or wants to view a skipped scene while watching, the playlist is updated in real time. The server reflects the user's actions, and a regenerated playlist is sent to the device.

[1155] User behavior

[1156] When a user wants to watch a video, they first upload the video file to the server. Next, they launch the application on their device and select the video they want to watch. The user then configures time-saving playback settings to suit their convenience and interests. Once playback settings are configured, the device automatically obtains scene scoring information from the server and generates a playlist. If the user wants to check a scene that was skipped during viewing or cancel time-saving playback, they can press a button to update the playlist in real time, allowing them to view content according to their needs.

[1157] Specific examples

[1158] For example, consider a case where a user wants to watch a 30-minute talk show but only has 50% of the time. The user sets the device's application to "play only 50% of the important parts." When the video is uploaded to the server, the server divides the video into frames and uses an AI model to score the importance of each scene. Based on the scoring results, the server prioritizes the scenes and sends them back to the device. The device then lists only the important scenes based on the settings and plays them back so that the user can watch efficiently. If a scene of interest is skipped while watching, the user can press a dedicated button to jump to that scene. This system allows users to watch within a limited time without missing any important information.

[1159] The processing flow will be explained below.

[1160] Server Processing

[1161] Step 1:

[1162] The server receives the video data from the user and stores it in temporary storage.

[1163] Step 2:

[1164] The server divides the video data into frames, i.e., converts the video into multiple still images.

[1165] Step 3:

[1166] The server clusters frames by scene, which groups frames with similar characteristics together.

[1167] Step 4:

[1168] To evaluate each scene, the server analyzes audio data, text data, and visual data. For example, it analyzes the speaker's content from audio data and subtitles and descriptions from text data.

[1169] Step 5:

[1170] The server scores each scene based on its interest and importance, and assigns a score to each scene based on the evaluation results.

[1171] Step 6:

[1172] The server prioritizes scenes based on the evaluated scores, and the priority information and scoring results are stored in a database.

[1173] Terminal handling

[1174] Step 1:

[1175] The device receives the user's time-saving playback settings, such as "play only the important parts at 50%."

[1176] Step 2:

[1177] The device retrieves scene scoring information from the server, including scene priority information.

[1178] Step 3:

[1179] The device will generate a playlist based on the user's settings, prioritizing important or interesting scenes.

[1180] Step 4:

[1181] The device plays videos based on the playlist, and unnecessary scenes are skipped as the user watches.

[1182] Step 5:

[1183] If a user wants to view a scene that was skipped during viewing, the device will accept a dedicated button and will readjust the playlist based on that information.

[1184] Step 6:

[1185] The device sends the updated playlist to the server and continues playing the video based on the regenerated playlist.

[1186] User Action

[1187] Step 1:

[1188] The user selects a video and uploads it to the server.

[1189] Step 2:

[1190] The user launches the application on the device, selects the video they want to watch, and enters the time-saving playback settings.

[1191] Step 3:

[1192] The user watches a video that has been efficiently edited based on the settings.

[1193] Step 4:

[1194] If a scene of interest to the user is skipped while watching, the user can press a dedicated button to jump to that scene.

[1195] By following the above steps, the user can efficiently view only the important or interesting parts of the content, thereby saving time.

[1196] Example 1

[1197] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1198] Previously, it was difficult to efficiently watch only the important or interesting parts of a long video. In particular, to obtain the information you wanted to watch within a limited time, you had to manually skip or fast-forward, which was inconvenient. Furthermore, there was no function to update the playlist in real time in response to user interactions, which meant that a flexible viewing experience was not provided.

[1199] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[1200] In this invention, the server includes means for uploading video data to the server, means for dividing the video data into frames and clustering them by scene, means for evaluating and scoring the interest and importance of each scene, means for evaluating the scenes using audio data analysis, text data analysis, and visual data analysis, means for prioritizing the scenes based on the evaluated scores, means for generating a playlist based on a user's time-saving playback setting and playing the video, and means for updating the playlist in response to user interactions. This allows a user to efficiently view important or interesting parts within a limited time and update the playlist in real time.

[1201] "Video data" is data in the form of dynamic media that includes visual and audio information.

[1202] A "server" is a system that provides services to other computers and devices on a network and processes and stores data.

[1203] A "frame" refers to each still image that makes up a video, and when played back in succession, they form a moving image.

[1204] "Clustering" is a process of classifying data into groups with common characteristics, and is a technique used to divide scenes in videos.

[1205] A "scene" is a group of consecutive frames that are semantically consistent within a video, and is a part that expresses a story or event.

[1206] "Fun" refers to elements that interest and entertain viewers, and is one of the criteria for evaluating a scene or the entire video.

[1207] "Importance" is an index that indicates how valuable the content of a video is to the viewer, and is a criterion for determining the priority of scenes.

[1208] "Scoring" is the process of assigning a numerical score to each scene based on an evaluation standard.

[1209] "Audio data analysis" is a technology that converts the audio portion of a video into text and analyzes its content.

[1210] "Text data analysis" is a technique for analyzing the content of text and evaluating its meaning and importance.

[1211] "Visual data analysis" is a technology that analyzes the content of images and videos and identifies objects and people.

[1212] A "playlist" is a list created to determine the playback order and content of videos.

[1213] "Time-saving playback settings" are options that users can set to shorten the playback time of videos under certain conditions.

[1214] "User interaction" refers to the operations and inputs that a user makes to a system.

[1215] The present invention provides a system that enables users to efficiently watch videos. The system mainly comprises a server, a terminal, and a user.

[1216] The server receives video data uploaded by users. After receiving the video data, the server uses common video processing tools such as FFmpeg to divide the video into frames. For example, a 30-minute video at 30 FPS would be divided into 54,000 frames. Next, each frame is clustered into scenes. This process involves using OpenCV to detect changes in color and motion, and then using a clustering algorithm (e.g., K-means) to divide the scenes. Each scene is evaluated using audio data analysis (e.g., Google Cloud Speech-to-Text, which converts speech to text), text data analysis (e.g., OpenAI GPT-4, which analyzes the meaning of text), and visual data analysis (e.g., OpenCV, which identifies objects and people). Each scene is scored based on the evaluation results and stored in a database. This scoring information is used to generate a playlist when users watch the video.

[1217] The device provides the user with an interface to select the video they want to watch. The user can set time-saving playback settings, such as "play only the important parts for 50%" or "play only the interesting parts for 30%." These settings are sent to the server, which provides the device with data to generate a playlist based on the stored scoring information. The device generates a playlist based on this data and plays the video according to the settings selected by the user. If the user wants to select a specific scene or view a skipped scene while watching, the playlist is updated in real time and the information is synchronized with the server.

[1218] To watch a video, a user first uploads the video file to the server. Next, they launch a dedicated application on their device and select the video they want to watch. The user configures time-saving playback settings based on their interests and convenience. Once this configuration is complete, the server obtains the scene scoring information and sends it to the device. The device then generates a playlist based on this information and plays the video according to the settings. As the user interacts with the video while watching, the playlist is updated in real time, allowing for flexible viewing.

[1219] For example, suppose a user wants to watch a 30-minute talk show, but only has 15 minutes available. The user sets the device app to "play only 50% of the important parts." Use this prompt:

[1220] "I would like to upload the following 30-minute talk show video and play it based on the setting of 'play only 50% of the important parts'. Please rate the importance of each scene using a system that uses Google Cloud Speech-to-Text for audio analysis, OpenAI GPT-4 for text analysis, and OpenCV for visual analysis. Based on the evaluation results, please make a list of only the important scenes, so that it can be viewed in 15 minutes."

[1221] As described above, the present invention provides a system that allows users to efficiently view important or interesting parts within a limited time. The advanced analysis technology of the server and the user-friendly interface of the terminal realize a flexible and effective viewing experience.

[1222] The flow of the identification process in the first embodiment will be described with reference to FIG.

[1223] Step 1:

[1224] The server receives the video data. The user selects a video file from the device and uploads it to the server. The input is the video file provided by the user, and the output is a video file saved in a temporary directory on the server. At this stage, data is sent and received using the HTTP protocol.

[1225] Step 2:

[1226] The server splits the video data into frames. The input is the video file saved in step 1, and the output is a set of still images for each frame. This process uses FFmpeg to split the video based on the number of seconds and frame rate. Specifically, a 30-minute video (30 FPS) will generate 54,000 still frames.

[1227] Step 3:

[1228] The server clusters each frame by scene. The input is the set of still images generated in step 2, and the output is a list of frames clustered by scene. OpenCV is used to analyze the color and motion of the images, and a clustering algorithm (e.g., K-means) is used to divide the scenes. Specifically, identical scenes are grouped together based on the similarity of the frames.

[1229] Step 4:

[1230] The server evaluates the scenes. The input is the list of scenes clustered in Step 3, and the output is an evaluation score for each scene. Google Cloud Speech-to-Text is used for audio data analysis, OpenAI GPT-4 for text data analysis, and OpenCV is used again for visual data analysis. Specifically, the server converts audio into text, analyzes its content, and then analyzes visual information to evaluate the importance and interest of the scene.

[1231] Step 5:

[1232] The server scores and prioritizes scenes based on the evaluation results. The input is the evaluation data obtained in step 4, and the output is a score list for each scene. Scoring is performed on a scale of 0 to 100, and scenes with higher scores have a higher viewing priority. Specifically, the evaluation score for each scene is quantified, and the scenes are sorted based on that.

[1233] Step 6:

[1234] The server saves the scored scenes in a database. The input is a list of scenes scored in step 5, and the output is the scoring information saved in the database. Specifically, data such as the scene frame number, start and end time, and score are stored in the database.

[1235] Step 7:

[1236] The user sets viewing preferences on the device. The user launches a dedicated application and selects the video they want to watch. The input is the user's viewing preferences (e.g., "play only the important parts at 50%), and the output is the viewing preferences data sent to the server. This setting is made via the application's UI.

[1237] Step 8:

[1238] The device retrieves the scoring information from the server. The input is the viewing preference data sent in step 7, and the output is the scoring information received from the server. The server filters the scoring information based on the user-specified preferences and sends it back to the device.

[1239] Step 9:

[1240] The device generates a playlist. The input is the scoring information obtained in step 8, and the output is the generated playlist. The playlist is ordered by scenes with the highest scores based on the user's viewing preferences, and is adjusted to match the configured playback percentage.

[1241] Step 10:

[1242] As the user watches the video, the playlist is updated in real time. The input is the user's viewing actions and interactions, and the output is the updated playlist and the video being played. If the user wants to select a specific scene or rewatch a skipped scene, they can do so using the device interface, and the content is synchronized with the server in real time. This allows users to flexibly customize their viewing experience.

[1243] (Application example 1)

[1244] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1245] Currently, many factories are taking various measures to improve work efficiency, but optimal suggestions and real-time improvements are still often done manually. Therefore, there is a need for a system that can automatically generate suggestions for efficiency and quickly identify and play back important scenes in videos. Currently, there is no effective method to solve this problem.

[1246] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[1247] In this invention, the server includes a means for uploading video data to the server, a means for dividing the video data into frames and clustering them by scene, a means for evaluating and scoring the interest and importance of each scene, a means for prioritizing the scenes based on the evaluated scores, and a means for analyzing work videos in an industrial environment and generating suggestions for improving efficiency, thereby enabling automatic and effective improvement of work efficiency in a factory.

[1248] "Video data" means digital files containing moving images and audio.

[1249] "Uploading" is the act of a user sending data from their device to a server.

[1250] A "scene" is a part of a video that consists of a certain series of frames and represents a series of movements or situations.

[1251] "Clustering" is a data processing technique that classifies many frames into groups with similar attributes.

[1252] "Evaluation" is the act of quantitatively or qualitatively judging the importance and interest of each scene.

[1253] "Scoring" is the process of assigning a numerical value to each scene based on the evaluation results.

[1254] "Prioritization" is the act of ranking scored scenes according to their importance and interest.

[1255] "Industrial environment" means the place or conditions in which manufacturing or production takes place.

[1256] "Efficiency" refers to reducing waste in work and processes and improving productivity.

[1257] A "proposal" refers to a solution or improvement to a specific problem.

[1258] System Overview

[1259] This invention is a video analysis system for improving work efficiency in an industrial environment. The system mainly includes a means for uploading video data to a server, a means for dividing the video data into frames and clustering them by scene, a means for evaluating and scoring the interest and importance of each scene, a means for prioritizing scenes based on the evaluated scores, and a means for analyzing work videos in an industrial environment and generating suggestions for improving efficiency.

[1260] Hardware and software used

[1261] Hardware: Robots used in factories (e.g. drones for monitoring work)

[1262] Software: Python, OpenCV, scikit-learn, cloud AI services (e.g., Google Cloud AI)

[1263] Data processing and calculation process

[1264] Server behavior:

[1265] 1. Uploading video data: Users upload videos of work they have taken in the factory to the server. The uploaded video data is stored on the server.

[1266] 2. Video Frame Segmentation: The server splits the received video data into frames. OpenCV is used to split the video data into individual frames.

[1267] 3. Frame Clustering: Each frame is then classified into scenes using K-means clustering, which groups similar scenes together.

[1268] 4. Scene Scoring: The clustered scenes are evaluated through audio, text, and visual data analysis. Based on the evaluation results, the importance of each scene is scored. Specifically, a machine learning algorithm is applied using scikit-learn.

[1269] 5. Prioritization: Based on the scoring results, scenes are prioritized in order of importance.

[1270] Terminal behavior:

[1271] 1. Playlist generation: The user uses the device application to select videos and set time-saving playback settings. Based on this, the device retrieves prioritized scene information from the server and generates a playlist.

[1272] 2. Work efficiency suggestions: Based on the generated playlist, suggestions for work efficiency are displayed, allowing users to quickly review only the important scenes.

[1273] 3. Updates based on user interaction: The playlist is updated in real time based on user actions.

[1274] Specific examples

[1275] For example, consider a video recording of assembly work in a factory. When a user uploads the video to the server, the server divides the video into frames and classifies them into scenes using K-means clustering. Each scene is evaluated through audio data analysis, text data analysis, and visual data analysis, and the importance of the scene is scored based on the results. Scenes with higher importance are given higher priority and displayed on the user's device based on the generated playlist. This allows users to quickly check the efficiency of work in the factory and receive suggestions for improvement.

[1276] Prompt Sentence Examples

[1277] "Analyze videos recorded in a factory to determine which scenes are important and which parts can be made more efficient. Then, generate suggestions for improving efficiency for each scene."

[1278] In this way, a system that dramatically improves work efficiency within a factory can be realized.

[1279] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[1280] Step 1:

[1281] The server receives the work videos taken by the user in the factory and stores the video data uploaded to the server.

[1282] Input: Video file uploaded from the user's device

[1283] Data processing: Save the video data on the server.

[1284] Output: Video data on the server

[1285] Step 2:

[1286] The server divides the video data into frames.

[1287] Input: Saved video data

[1288] Data processing: Video data is divided into individual frames using OpenCV.

[1289] Output: List of frame images

[1290] Step 3:

[1291] The server clusters each of the divided frames.

[1292] Input: List of frame images

[1293] Data operations: Use K-means clustering to group frames that belong to the same scene. Vectorize the frame images and cluster them in feature space.

[1294] Output: Clustered groups of frames

[1295] Step 4:

[1296] The server analyzes the clustered scenes and performs scoring.

[1297] Input: A group of clustered frames

[1298] Data calculation: Evaluate the importance of each scene through audio data analysis, text data analysis, and visual data analysis. Specifically, audio analysis uses the Google Cloud Speech-to-Text API, text analysis uses NLP (natural language processing) models, and visual data analysis applies machine learning algorithms.

[1299] Output: Importance score for each scene

[1300] Step 5:

[1301] The server prioritizes the scenes based on the scoring results.

[1302] Input: Importance score for each scene

[1303] Data calculation: sorting scenes in order of importance based on their scores.

[1304] Output: A prioritized list of scenes

[1305] Step 6:

[1306] The terminal generates a playlist based on the user's time-saving playback settings.

[1307] Input: A prioritized list of scenes and the user's time-saving playback settings

[1308] Data processing: Generate a playlist containing only important scenes, taking into account user preferences.

[1309] Output: Playlist

[1310] Step 7:

[1311] The terminal plays the videos based on the playlist.

[1312] Input: Playlist

[1313] Data processing: Play only the scenes included in the playlist in sequence.

[1314] Output: Video playback

[1315] Step 8:

[1316] The terminal updates the playlist in real time in response to user interactions.

[1317] Input: User actions (selecting a specific scene, skipping, etc.)

[1318] Data processing: Regenerate the playlist based on user actions to reflect updated information.

[1319] Output: Updated playlist and video playback changes

[1320] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[1321] This invention is a system that allows users to effectively shorten the time spent watching videos and efficiently watch only the interesting or important parts. This system consists of four main components: a server, a terminal, a user, and an emotion engine.

[1322] Server Operation

[1323] The server receives video data uploaded by users. After receiving the video data, the server divides the video into frames and clusters each frame into scenes. The clustered scenes are evaluated through audio data analysis, text data analysis, and visual data analysis, and each scene is scored for its interest and importance. The scored scenes are prioritized based on the evaluation results and stored in a database on the server.

[1324] Device behavior

[1325] When a user wants to watch a video, they first launch the application on their device and select a video. Users can set viewing preferences for time-saving playback, such as "play only the important parts for 50%" or "play only the interesting parts for 30%." These preferences are sent from the device to the server, which then obtains scoring information for prioritized scenes. The device generates a playlist based on this information and plays the video according to the preferences. If the user selects a specific scene or wants to view a skipped scene while watching, the playlist is updated in real time. The server reflects the user's actions, and a regenerated playlist is sent to the device.

[1326] User behavior

[1327] When a user wants to watch a video, they first upload the video file to the server. Next, they launch the application on their device and select the video they want to watch. The user then configures time-saving playback settings to suit their convenience and interests. Once playback settings are configured, the device automatically obtains scene scoring information from the server and generates a playlist. If the user wants to check a scene that was skipped during viewing or cancel time-saving playback, they can press a button to update the playlist in real time, allowing them to view content according to their needs.

[1328] Emotion Engine Operation

[1329] The device is equipped with an emotion engine that monitors the user's emotions while watching videos. The engine analyzes the user's facial expressions, tone of voice, and eye movements to recognize the user's emotional state in real time. The emotion data recognized by the emotion engine is fed back to the device and used to adjust the playlist.

[1330] For example, if a user is enjoying a scene, the score for that scene will increase, and similar scenes will be added to the list with priority. Emotional data will also be accumulated and used as learning data to reflect in future scoring and prioritization. This function allows video viewing to be more tailored to the user's preferences.

[1331] Specific examples

[1332] For example, consider a case where a user wants to watch a 30-minute talk show but only has 50% of the time. The user sets the device application to "play only 50% of the important parts." When the video is uploaded to the server, the server divides the video into frames and uses an AI model to score the importance of each scene. Based on the scoring results, the server prioritizes the scenes and sends them back to the device. Based on the settings, the device lists only the important scenes and plays them so that the user can watch efficiently.

[1333] While watching, the emotion engine monitors the user's reactions and, if they are enjoying the content, adjusts the priority of the corresponding scene in real time. If a scene that interests the user is skipped while watching, they can press a dedicated button to jump to that scene. This system not only allows users to efficiently view important information and interesting parts within a limited time, but also provides a viewing experience customized to their emotions.

[1334] The processing flow will be explained below.

[1335] Server Processing

[1336] Step 1:

[1337] The server receives video data from the user and stores it in temporary storage.

[1338] Step 2:

[1339] The server divides the video data into frames and converts the video into individual still images.

[1340] Step 3:

[1341] The server clusters frames by scene, which groups frames with similar characteristics together.

[1342] Step 4:

[1343] To evaluate each scene, the server performs audio data analysis, text data analysis, and visual data analysis, such as analyzing the speaker's content from the audio data and analyzing subtitles and descriptions from the text data.

[1344] Step 5:

[1345] The server scores each scene based on its interest and importance, and assigns a score to each scene based on the evaluation results.

[1346] Step 6:

[1347] The server prioritizes scenes based on the evaluated scores and stores the priority information and scoring results in a database.

[1348] Terminal handling

[1349] Step 1:

[1350] The device receives the user's time-saving playback settings, such as "play only the important parts at 50%."

[1351] Step 2:

[1352] The device retrieves scene scoring information, including prioritization information, from the server.

[1353] Step 3:

[1354] The device generates a playlist based on the user's settings, prioritizing important or interesting scenes.

[1355] Step 4:

[1356] The device plays videos based on the playlist, and unnecessary scenes are skipped as the user watches.

[1357] Step 5:

[1358] If the user wants to view a scene that was skipped during viewing, the device will accept a dedicated button and readjust the playlist based on that information.

[1359] Step 6:

[1360] The device sends the updated playlist to the server and continues playing the video based on the regenerated playlist.

[1361] User Action

[1362] Step 1:

[1363] The user selects a video and uploads it to the server.

[1364] Step 2:

[1365] The user launches the application on the device, selects the video they want to watch, and enters the time-saving playback settings.

[1366] Step 3:

[1367] The user watches a video that has been efficiently edited based on the settings.

[1368] Step 4:

[1369] If a scene of interest to the user is skipped while watching, the user can press a dedicated button to jump to that scene.

[1370] Emotion engine processing

[1371] Step 1:

[1372] The device's built-in emotion engine monitors the user's facial expressions, tone of voice, gaze, and other factors in real time.

[1373] Step 2:

[1374] The emotion engine analyzes the user's emotional state and generates emotion data based on that.

[1375] Step 3:

[1376] The generated emotion data is received by the device and used to adjust the playlist, for example by increasing the score of scenes that are enjoyed.

[1377] Step 4:

[1378] Emotional data is accumulated and sent to the server to be used as learning data for future scoring and prioritization.

[1379] Specific examples

[1380] For example, consider a case where a user wants to watch a 30-minute talk show but only has 50% of the time. The user sets the device application to "play only 50% of the important parts." When the video is uploaded to the server, the server divides the video into frames and uses an AI model to score the importance of each scene. Based on the scoring results, the server prioritizes the scenes and sends them back to the device. The device then lists only the important scenes based on the settings and plays them back so that the user can watch efficiently.

[1381] While watching, the emotion engine monitors the user's reactions and, if they are enjoying the content, adjusts the priority of the corresponding scene in real time. If a scene that interests the user is skipped while watching, they can press a dedicated button to jump to that scene. This system not only allows users to efficiently view important information and interesting parts within a limited time, but also provides a viewing experience that is customized according to their emotions.

[1382] Example 2

[1383] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1384] Conventional video viewing systems make it difficult for users to efficiently grasp important or interesting parts of a video. They also lack a means to meet the need to view only the most important information within a limited time. Furthermore, content optimization based on user emotions is not performed, preventing consistent satisfaction.

[1385] The identification process by the identification processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means. In this invention, the server includes a means for uploading video data to the information processing device, a means for dividing the video data into frames and classifying them by scene, a means for evaluating and scoring the interest and importance of each scene, a means for prioritizing the scenes based on the evaluated scores, a means for generating a playlist based on the user's playback settings and playing the video, a means for updating the playlist in response to the user's operations, and a means for recognizing the user's emotions during viewing and adjusting the playlist. This makes it possible to streamline the video viewing experience and provide customized playback according to the user's interests and emotions.

[1386] "Video data" means a digitally represented video file containing visual and audio information.

[1387] An "information processing device" is a computer system for processing, storing, and analyzing data.

[1388] A "user" is a person who operates the system and watches videos.

[1389] A "frame" is an individual still image that makes up a video.

[1390] A "scene" is a set of frames that represent a certain continuous time and space within a video.

[1391] "Classification" is the process of grouping frames that have the same characteristics or features.

[1392] "Scoring" is the process of quantifying the interest and importance of a scene based on certain criteria.

[1393] "Prioritization" means determining the order in which scenes are played based on the evaluated scores.

[1394] A "playlist" is a list that defines the order in which scenes are to be played.

[1395] "Playback" is the process of visually and aurally representing a moving image.

[1396] "Playback settings" are settings that specify the playback method and range of the video that the user wants to watch.

[1397] "Emotions" are the psychological and physiological reactions of users while watching.

[1398] "Recognition" refers to grasping a user's emotional state by analyzing their facial expressions, tone of voice, eye movements, etc.

[1399] "Adjusting" means changing scene priorities and playlists in real time.

[1400] "Operation" refers to the instructions or input that a user gives to the system.

[1401] This invention is a system that allows users to effectively shorten the time it takes to watch a video and efficiently watch only the interesting or important parts. This system consists of four main components: a server, a terminal, a user, and an emotion engine.

[1402] Server Operation

[1403] The server receives video data uploaded by users. Once the video data is received, the server divides the video into frames using FFmpeg. The divided frames are then clustered by scene using the FrameGenius API. The clustered scenes are analyzed for audio, text, and visual data using Google Cloud's Speech-to-Text API and NLP Cloud. Based on the analyzed data, the interestingness and importance of each scene is evaluated and scored using a proprietary scoring algorithm. The scenes are prioritized based on the scoring results, and the results are stored in a MongoDB database.

[1404] Device behavior

[1405] When a user wants to watch a video, they launch a dedicated application on their device and select the video they want to watch. Users can set time-saving playback settings such as "play only the important parts at 50%" or "play only the interesting parts at 30%." This setting information is sent from the device to the server, which then obtains scoring information for prioritized scenes. The device generates a playlist based on this information and plays the video using HLS (HTTP Live Streaming) technology. If the user selects a specific scene while watching or wants to watch a skipped scene, the playlist is updated in real time and the information is reflected on the server.

[1406] User actions

[1407] Users select the video they want to watch from their device and upload it to the server. Next, they launch a dedicated application on their device and select the video they want to watch. After selecting the video they want to watch, users can set up time-saving playback according to their interests and convenience. Once playback settings are set, the device automatically obtains scene scoring information from the server and generates a playlist. If users want to check scenes that were skipped during viewing or cancel time-saving playback, they can operate a dedicated button to update the playlist in real time, allowing them to watch according to their needs.

[1408] Emotion Engine Operation

[1409] The device is equipped with an emotion engine that monitors the user's emotions while watching videos using the Momento AI engine. This engine analyzes the user's facial expressions, vocal tone, eye movements, and other factors to recognize the user's emotional state in real time. The recognized emotion data is fed back to the device and used to adjust the playlist. For example, if the user is enjoying a scene very much, the score for that scene will increase, and similar scenes will be added to the playlist with priority. Emotion data is also accumulated and used as learning data to reflect in future scoring and prioritization. This function enables video viewing that is more tailored to the user's preferences.

[1410] Specific examples

[1411] For example, consider the case where a user wants to watch a 30-minute talk show but only has 15 minutes to watch. The user sets the device application to "play only the important parts at 50%." When the video is uploaded to the server, the server uses FFmpeg to divide the video into frames and clusters each scene using the FrameGenius API. The server then analyzes the video using Google Cloud's Speech-to-Text API and NLP Cloud to score its importance. Prioritization is performed based on the scoring results, and the results are sent to the device. The device then lists only the important scenes in a playlist and plays the video using HLS technology.

[1412] While watching, the emotional engine monitors the user's reactions using the Momento AI engine, and if the user is enjoying the video, the priority of the corresponding scene is adjusted in real time. Users can jump to the relevant scene by pressing the "Watch Skipped Scene" button. This system not only allows users to efficiently view important information and interesting parts within a limited time, but also provides a viewing experience customized to their emotions.

[1413] Prompt Sentence Examples

[1414] "I want to set up time-saving playback so that I can watch only the important parts of a 30-minute talk show video in 50% of the viewing time."

[1415] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1416] Step 1: Upload your video

[1417] The server receives video data from the user. The user selects the video they want to watch on their device and uploads it to the server using a dedicated application. Input: Video file sent from the user's device. Output: Video data saved on the server.

[1418] Specific behavior:

[1419] A user opens the application and clicks the "Upload Video" button.

[1420] The user selects the video they want to watch from the file selection dialog and clicks "Open."

[1421] The video file is sent to the server and the upload is complete.

[1422] Step 2: Video analysis and frame segmentation

[1423] The server splits the received video into frames using FFmpeg. Input: Video data stored on the server. Output: Still images for each frame.

[1424] Specific behavior:

[1425] The server splits the received video data into frames using the FFmpeg tool.

[1426] The split frames are saved in a temporary folder.

[1427] Step 3: Scene Clustering

[1428] The server uses the FrameGenius API to classify frames into scenes. Input: Segmented frame data. Output: Clustered scene data.

[1429] Specific behavior:

[1430] The FrameGenius API is used to analyze the visual characteristics of each frame.

[1431] Frames are grouped into clusters based on similarity and organized into scenes.

[1432] Step 4: Analyze and score the scene

[1433] The server uses Google Cloud's Speech-to-Text API and NLP Cloud to analyze the audio, text, and visual data for each scene. Input: Clustered scene data. Output: Analysis results and scoring data for each scene.

[1434] Specific behavior:

[1435] The audio data contained in each clustered scene is converted into text data using Google Cloud's Speech-to-Text API.

[1436] Text data is analyzed using NLP Cloud to extract keywords and emotions.

[1437] The visual data of each scene is also analyzed to evaluate its importance and interest.

[1438] Step 5: Prioritize and save your scenes

[1439] The server uses a proprietary algorithm to determine the priority of scenes based on the analysis results and stores the results in a MongoDB database. Input: Scoring data for each scene. Output: A prioritized scene list.

[1440] Specific behavior:

[1441] Scenes are sorted by priority based on the scoring data for each scene.

[1442] Store the prioritized scene list in a MongoDB database.

[1443] Step 6: Submit viewing preferences

[1444] The user configures the time-saving playback settings using the application on the device and sends this setting information to the server. Input: User's viewing settings. Output: Viewing setting data sent to the server.

[1445] Specific behavior:

[1446] The user opens the application and selects the video they want to watch.

[1447] Users can select viewing settings such as "play only the important parts at 50%" or "play only the interesting parts at 30%."

[1448] The configuration information is sent to the server.

[1449] Step 7: Generate and play a playlist

[1450] The server sends prioritized scene scoring information based on the received viewing preferences to the device, which then generates a playlist based on that information and plays the video using HLS technology. Input: User's viewing preferences, prioritized scene list. Output: Playlist, video to be played.

[1451] Specific behavior:

[1452] The server filters out priority scenes according to viewing settings and transmits them to the terminal.

[1453] The device uses HLS technology to generate a playlist and plays the video in scene order.

[1454] Step 8: Emotion Engine Feedback

[1455] The emotion engine built into the device monitors the user's emotional state in real time and adjusts the playlist. Input: Real-time emotion data while the user is watching. Output: Adjusted playlist.

[1456] Specific behavior:

[1457] The emotion engine analyzes the user's facial expressions, tone of voice, and eye movements.

[1458] Using the Momento AI engine, it recognizes the user's emotional state and provides feedback in real time.

[1459] Adjust the playlist to include scenes that users enjoy or find interesting.

[1460] (Application example 2)

[1461] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1462] While modern consumers have access to a wealth of video content, it is difficult for them to efficiently obtain the most interesting content and important information within a limited viewing time. Furthermore, the ability to adjust playback content in real time based on the viewer's emotions and reactions is also lacking. In response to this problem, the present invention aims to provide a system that provides a video viewing experience that maximizes user engagement within a limited time.

[1463] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[1464] In this invention, the server includes means for uploading video data to the server, means for dividing the video data into frames and clustering them by scene, means for evaluating and scoring the interest and importance of each scene, means for prioritizing the scenes based on the evaluated scores, means for generating a playlist based on a user's time-saving playback settings and playing the videos, means for updating the playlist in response to user interactions, and means for analyzing the user's facial expressions, tone of voice, and eye movements and adjusting the playlist based on emotion data. This allows users to efficiently watch only interesting or important parts even within a limited time, and further enables a personalized viewing experience based on emotions.

[1465] "Means for uploading video data to a server" refers to a function that allows a user to send video files from their own device to a server.

[1466] The "means of dividing video data into frames and clustering by scene" is a process of dividing a video uploaded to a server into frames and grouping similar frames.

[1467] The "means for evaluating and scoring the interest and importance of each scene" is a process of analyzing the clustered scenes and assigning a numerical evaluation of their interest and importance.

[1468] The "means for prioritizing scenes based on the evaluated scores" is a mechanism for sorting scored scenes in order of importance based on their scores.

[1469] "Means for generating a playlist based on the user's time-saving playback settings and playing videos" refers to a function that creates an optimized playlist based on the user's specified playback time and priority scene conditions, and plays videos according to that list.

[1470] The "means for updating the playlist in response to user interaction" refers to a system that dynamically changes the playlist based on the operations performed by the user while viewing (e.g., selecting or skipping a specific scene).

[1471] "Means of analyzing the user's facial expressions, tone of voice, and eye movements, and adjusting the playlist based on emotional data" refers to a function that analyzes the user's emotions and reactions while watching, and adjusts the playlist on the spot based on the results.

[1472] This invention provides a system that enables users to effectively shorten the time spent watching videos and efficiently watch only the interesting or important parts. This system consists of four main components: a server, a terminal, a user, and an emotion engine.

[1473] Server Operation

[1474] The server receives video data uploaded by users. After receiving the video data, the server divides the video into frames and clusters each frame into scenes. The clustered scenes are evaluated through audio data analysis, text data analysis, and visual data analysis to score each scene's interest and importance. The scenes are prioritized based on the evaluation results and stored in a database.

[1475] Device behavior

[1476] When a user wants to watch a video, they first launch the application on their device and select a video. Users can set viewing preferences for time-saving playback, such as "play only the important parts for 50%" or "play only the interesting parts for 30%." These preferences are sent from the device to the server, which then obtains scoring information for prioritized scenes. The device generates a playlist based on this information and plays the video according to the preferences. If the user selects a specific scene or wants to view a skipped scene while watching, the playlist is updated in real time. The user's actions are reflected on the server, and a regenerated playlist is sent to the device.

[1477] User behavior

[1478] When a user wants to watch a video, they first upload the video file to the server. Next, they launch the application on their device and select the video they want to watch. The user then configures time-saving playback settings to suit their convenience and interests. Once playback settings are configured, the device automatically obtains scene scoring information from the server and generates a playlist. If the user wants to check a scene that was skipped during viewing or cancel time-saving playback, they can press a button to update the playlist in real time, allowing them to watch according to their needs.

[1479] Emotion Engine Operation

[1480] The device is equipped with an emotion engine that monitors the user's emotions while watching videos. This engine analyzes the user's facial expressions, vocal tone, eye movements, and other factors to recognize the user's emotional state in real time. The emotion data recognized by the emotion engine is fed back to the device and used to adjust the playlist. For example, if the user is enjoying a video very much, the score for that scene will increase and similar scenes will be added to the playlist with priority. Emotion data is also accumulated and used as learning data to reflect in future scoring and prioritization. This function enables video viewing that is more tailored to the user's preferences.

[1481] Specific examples

[1482] Consider a scenario where a user wants to watch a 30-minute talk show but only has 50% of the time available. The user selects "Play only 50% of the important parts" in the device application. When the video is uploaded to the server, the server divides it into frames and uses a generative AI model to score the importance of each scene. Based on the scoring results, the server prioritizes the scenes and sends them back to the device. The device then lists only the important scenes based on the user's settings and plays them for efficient viewing. While watching, the emotion engine monitors the user's reactions. If the user is enjoying the video, the priority of the corresponding scene is adjusted in real time. If a scene that interests the user is skipped while watching, the user can press a dedicated button to jump back to that scene. This system not only allows users to efficiently view important information and interesting parts within a limited time, but also provides a customized viewing experience tailored to their emotions.

[1483] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1484] Step 1:

[1485] The user uploads a video file to the server. The user selects a video file using an application on their device and presses the upload button, sending the video data to the server. The input is the video file selected by the user, and the output is the video data saved on the server.

[1486] Step 2:

[1487] The server divides the uploaded video data into frames. When the server receives the video data, it divides the video into frames and stores each frame in memory. The input is video data, and the output is image data divided into frames.

[1488] Step 3:

[1489] The server clusters the image data for each frame by scene. The server uses a clustering algorithm to classify similar frames into the same scene cluster. The input is the divided frame image data, and the output is the clustered scene groups.

[1490] Step 4:

[1491] The server evaluates and scores the interest and importance of each scene group. The server performs audio, text, and visual analysis, and uses a generative AI model to evaluate the importance and interest of each scene as a numerical value. The input is the clustered scene data, and the output is the evaluation score.

[1492] Step 5:

[1493] The server prioritizes scenes based on the evaluated scores. The server prioritizes the playback order of each scene based on the scoring results. The input is the evaluation scores, and the output is a list of prioritized scene clusters.

[1494] Step 6:

[1495] The device sends the user's time-saving playback settings to the server and generates a playlist. The user configures playback settings within the application, and the device sends them to the server. The server creates a playlist based on the prioritized scene list and matches the user's settings. The input is the user's playback settings, and the output is an optimized playlist.

[1496] Step 7:

[1497] The device plays videos based on the generated playlist. The device receives the playlist sent from the server and plays videos according to that list. The input is the playlist and the output is the playback of the video.

[1498] Step 8:

[1499] The device updates the playlist according to the user's interactions. When the user skips or selects a specific scene while watching a video, the device sends that information to the server in real time and updates the playlist. The input is the user's interaction data, and the output is the updated playlist.

[1500] Step 9:

[1501] The device uses an emotion engine to monitor the user's emotional state and adjust the playlist accordingly. The emotion engine built into the device analyzes the user's facial expressions, tone of voice, and eye movements to obtain emotional data in real time. The obtained emotional data is sent to the server, which then adjusts the playlist based on that data. The input is the user's emotional data, and the output is a playlist adjusted based on the emotions.

[1502] Step 10:

[1503] The server sends the adjusted playlist to the terminal, and the terminal updates the playlist. The input is the adjusted playlist, and the output is the playlist update.

[1504] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.

[1505] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1506] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.

[1507] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[1508] FIG. 9 is a diagram illustrating an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and actions arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.

[1509] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.

[1510] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).

[1511] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.

[1512] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."

[1513] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values ​​indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.

[1514] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).

[1515] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.

[1516] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.

[1517] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[1518] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.

[1519] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.

[1520] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.

[1521] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.

[1522] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.

[1523] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.

[1524] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.

[1525] The following is further disclosed regarding the above embodiment.

[1526] (Claim 1)

[1527] A means for uploading video data to a server;

[1528] A means for dividing video data into frames and clustering each scene;

[1529] A method for evaluating and scoring the funniness and importance of each scene,

[1530] a means for prioritizing scenes based on the evaluated scores;

[1531] means for generating a playlist based on a user's time-saving playback settings and playing videos;

[1532] means for updating the playlist in response to user interactions;

[1533] A system including:

[1534] (Claim 2)

[1535] 2. The system of claim 1, wherein the means for evaluating the interestingness and importance of a scene scores the scene using audio data analysis, text data analysis, and visual data analysis.

[1536] (Claim 3)

[1537] 10. The system of claim 1, including means for allowing a user to select and play specific scenes during viewing, and to quickly jump to skipped scenes.

[1538] "Example 1"

[1539] (Claim 1)

[1540] A means for uploading video data to a server;

[1541] A means for dividing video data into frames and clustering each scene;

[1542] A method for evaluating and scoring the funniness and importance of each scene,

[1543] means for assessing a scene using audio data analysis, text data analysis, and visual data analysis;

[1544] a means for prioritizing scenes based on the evaluated scores;

[1545] means for generating a playlist based on a user's time-saving playback settings and playing videos;

[1546] means for updating the playlist in response to user interactions;

[1547] A system including:

[1548] (Claim 2)

[1549] 10. The system of claim 1, wherein the system scores a scene using audio data analysis, text data analysis, and visual data analysis.

[1550] (Claim 3)

[1551] 10. The system of claim 1, including means for allowing a user to select and play specific scenes during viewing, and to quickly jump to skipped scenes.

[1552] "Application Example 1"

[1553] (Claim 1)

[1554] A means for uploading video data to a server;

[1555] A means for dividing video data into frames and clustering each scene;

[1556] A method for evaluating and scoring the funniness and importance of each scene,

[1557] a means for prioritizing scenes based on the evaluated scores;

[1558] A means for analyzing work videos in an industrial environment and generating efficiency suggestions;

[1559] means for generating a playlist based on a user's time-saving playback settings and playing videos;

[1560] means for updating the playlist in response to user interactions;

[1561] A system including:

[1562] (Claim 2)

[1563] 2. The system of claim 1, wherein the means for evaluating the interestingness and importance of a scene scores the scene using audio data analysis, text data analysis, and visual data analysis.

[1564] (Claim 3)

[1565] 10. The system of claim 1, including means for allowing a user to select and play specific scenes during viewing, and to quickly jump to skipped scenes.

[1566] (Claim 4)

[1567] 10. The system of claim 1, further comprising suggestions for improving work efficiency within a factory.

[1568] "Example 2: Combining Emotion Engines"

[1569] (Claim 1)

[1570] means for uploading video data to an information processing device;

[1571] A means for dividing video data into frames and classifying them by scene;

[1572] A method for evaluating and scoring the funniness and importance of each scene,

[1573] a means for prioritizing scenes based on the evaluated scores;

[1574] means for generating a playlist and playing videos based on user preferences;

[1575] A means for updating the playlist in response to a user's operation;

[1576] a means for recognizing a user's emotions while listening and adjusting the playlist;

[1577] A system including:

[1578] (Claim 2)

[1579] 2. The system of claim 1, wherein the means for evaluating the interestingness and importance of a scene scores the scene using audio data analysis, text data analysis, and visual data analysis.

[1580] (Claim 3)

[1581] 10. The system of claim 1, further comprising means for allowing a user to select and play specific scenes during viewing, and to quickly jump to skipped scenes.

[1582] (Claim 4)

[1583] 10. The system of claim 1, wherein the system monitors user emotions while watching and adjusts the playlist in real time based on that data.

[1584] "Application example 2 when combining emotion engines"

[1585] (Claim 1)

[1586] A means for uploading video data to a server;

[1587] A means for dividing video data into frames and clustering each scene;

[1588] A method for evaluating and scoring the funniness and importance of each scene,

[1589] a means for prioritizing scenes based on the evaluated scores;

[1590] means for generating a playlist based on a user's time-saving playback settings and playing videos;

[1591] means for updating the playlist in response to user interactions;

[1592] A means for analyzing a user's facial expressions, tone of voice, and eye movements and adjusting the playlist based on the emotional data;

[1593] A system including:

[1594] (Claim 2)

[1595] 2. The system of claim 1, wherein the means for evaluating the interestingness and importance of a scene scores the scene using audio data analysis, text data analysis, and visual data analysis.

[1596] (Claim 3)

[1597] 10. The system of claim 1, including means for allowing a user to select and play specific scenes during viewing, and to quickly jump to skipped scenes. [Explanation of symbols]

[1598] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Device 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robot< / url:> < / url:> < / url:> < / url:>

Claims

1. A means for uploading video data to a server; A means for dividing video data into frames and clustering each scene; A method for evaluating and scoring the funniness and importance of each scene, a means for prioritizing scenes based on the evaluated scores; means for generating a playlist based on a user's time-saving playback settings and playing videos; means for updating the playlist in response to user interactions; A system including:

2. 2. The system according to claim 1, wherein the means for evaluating the interestingness and importance of a scene scores the scene using audio data analysis, text data analysis, and visual data analysis.

3. 10. The system of claim 1, further comprising means for allowing a user to select and play a particular scene during viewing, and to quickly jump to a skipped scene.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A