system
The system addresses the challenge of information overload in videos by converting audio to text, generating summaries, and creating chapters, allowing for rapid access to relevant content.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- SOFTBANK GROUP CORP
- Filing Date
- 2024-10-21
- Publication Date
- 2026-05-07
AI Technical Summary
Users face challenges in efficiently obtaining necessary information from vast amounts of video content, as it requires significant time and effort to search through lengthy videos.
A system that extracts audio from video data, converts it to text, identifies important points, generates summaries, and creates chapters, enabling quick access to relevant sections through a reverse search function.
Enables users to quickly and efficiently acquire information by reducing viewing time and improving information retrieval efficiency.
Smart Images

Figure 2026074997000001_ABST
Abstract
Description
Technical Field
[0004] , , , ,
[0005] , , , , ,
[0003] , , ,
[0001] The technology of the present disclosure relates to a system.
Background Art
[0002] Patent Document 1 discloses a method for controlling a persona chatbot, which is performed by at least one processor, including steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] In recent years, with the increase in video content, there is a problem that it is difficult for users to efficiently obtain necessary information from videos containing a vast amount of information. In particular, it takes time and effort to watch all of a long video to search for the target information, so there is a demand for improving the efficiency of information acquisition.
Means for Solving the Problems
[0005] This invention generates summary information by extracting audio from input video data, converting it to text, and using that text information to extract important points. Furthermore, it automatically generates chapters based on important information within the video and provides a reverse search function that supports keyword searches by the user, thereby providing a system that allows users to quickly and easily obtain necessary information from videos. This enables users to reduce viewing time and obtain information efficiently.
[0006] "Video data" refers to digital data that includes video and audio, and can be played back on computer systems and digital devices.
[0007] "Audio information" refers to sound information contained in video data, including human speech, music, and ambient sounds.
[0008] "Text information" refers to data that converts audio information into text, and is subject to natural language processing.
[0009] "Summary information" is information that concisely expresses the content of the original video data, and is composed of information that extracts important points and topics.
[0010] A "chapter" is a segment of video data divided based on a specific topic or content, allowing viewers to directly access the parts that interest them.
[0011] "Reverse search" is a function that allows users to efficiently find related information or chapters by entering specific keywords.
[0012] A "timestamp" is a digital marker that indicates a specific time within video data, and is used to mark the start of a chapter or event.
[0013] A "user interface" is a screen or means of interaction between a user and a computer system or digital device for exchanging information. [Brief explanation of the drawing]
[0014] [Figure 1] This is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] This is a conceptual diagram showing an example of the essential functions of a data processing device and a smart device according to the first embodiment. [Figure 3] This is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] This is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] This is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] This is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] This is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] This is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] This shows an emotion map where multiple emotions are mapped. [Figure 10] This shows an emotion map where multiple emotions are mapped. [Figure 11] This is a sequence diagram showing the processing flow of the data processing system in Example 1. [Figure 12] This is a sequence diagram showing the processing flow of the data processing system in Application Example 1. [Figure 13] This is a sequence diagram showing the processing flow of the data processing system in Example 2, which incorporates an emotion engine. [Figure 14]It is a sequence diagram showing the processing flow of a data processing system in Application Example 2 when a sentiment engine is combined.
Embodiments for Carrying Out the Invention
[0015] Hereinafter, an example of an embodiment of a system according to the technology of the present disclosure will be described with reference to the accompanying drawings.
[0016] First, the terms used in the following description will be explained.
[0017] In the following embodiments, a numbered processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of multiple arithmetic units. Also, the processor may be a single type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), an APU (Accelerated Processing Unit), and the like.
[0018] In the following embodiments, a numbered RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a work memory by the processor.
[0019] In the following embodiments, a numbered storage is one or more non-volatile storage devices that store various programs and various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes, etc.
[0020] In the following embodiments, the signed communication interface (I / F) is an interface that includes a communication processor and an antenna, etc. The communication interface manages communication between multiple computers. Examples of communication standards applicable to the communication interface include wireless communication standards such as 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark).
[0021] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."
[0022] [First Embodiment]
[0023] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.
[0024] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0025] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0026] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.
[0027] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0028] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0029] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.
[0030] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0031] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.
[0032] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0033] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0034] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".
[0035] This invention provides a system for efficiently summarizing video content and enabling quick access to sections of interest. This system is particularly effective for quickly obtaining necessary information from information-rich videos.
[0036] The system is realized through the interaction of servers, terminals, and users. The following describes the processing of the program.
[0037] Video analysis and summarization
[0038] The server receives videos uploaded by users and extracts the audio track. This audio track is converted into text using speech recognition technology. Then, important topics are identified through natural language processing, and summary information is generated. This summary information concisely represents the entire lengthy video.
[0039] Chapter generation and search function
[0040] The server generates video chapters based on identified topics and key points. This allows users to view the video divided into specific segments. Furthermore, a reverse search index is created based on the generated summary and chapter information, making it easier to search for specific information using keywords.
[0041] Providing a user interface
[0042] The device visually displays summary and chapter information delivered from the server to the user. Through this interface, the user can select a section of interest and directly watch the video related to that segment.
[0043] Specific example
[0044] For example, suppose there is a video of a one-hour business seminar. When a user uploads this video to the system, the server processes the video, extracts the important topics, and generates a summary of about 10 minutes. Furthermore, the video is divided into chapters by topic, for example, a chapter titled "Marketing Strategy" is generated. The terminal displays this information on its interface, allowing the user to immediately watch the "Marketing Strategy" chapter.
[0045] This system allows users to quickly and efficiently acquire the information they need in this age of information overload. It is extremely useful not only for personal use but also for business purposes.
[0046] The following describes the processing flow.
[0047] Step 1:
[0048] The user uploads a video file to the system. The terminal receives the file and transfers it to the server.
[0049] Step 2:
[0050] The server retrieves the video file, checks its format and encoding, and converts it to a standard format if necessary.
[0051] Step 3:
[0052] The server extracts the audio track from the converted video. This audio information is then converted into text using speech recognition (ASR) technology.
[0053] Step 4:
[0054] The server applies natural language processing (NLP) to the acquired text information to extract important keywords and topics. This process generates a summary of the video.
[0055] Step 5:
[0056] The server generates chapters for the video based on the extracted key topics. Each chapter is assigned a timestamp indicating its start time.
[0057] Step 6:
[0058] The server organizes summary and chapter information and creates an index to facilitate reverse searching by users. Search is used when users are looking for specific information.
[0059] Step 7:
[0060] The device displays summary and chapter information transmitted from the server on the user interface. This allows the user to select and directly view chapters of interest.
[0061] Step 8:
[0062] Users can use the search function to find chapters containing the information they need using keywords. Search results quickly navigate to relevant summaries and chapters.
[0063] (Example 1)
[0064] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0065] With the increasing volume of video content, it has become difficult for users to quickly and efficiently obtain important information from long videos. In particular, with information-rich videos, there is a need for methods to determine which parts are important to the user and access them efficiently. Furthermore, the inability to quickly obtain necessary information increases the user's time cost.
[0066] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0067] In this invention, the server includes means for acquiring video information, means for converting audio data into text information, and means for identifying important information based on the text information and generating summary data. This enables the user to efficiently access important information and quickly understand the video.
[0068] "Video and image information" refers to digital data consisting of video and audio.
[0069] "Audio data" refers to the sound elements included as part of video information, and is data expressed as an audio signal.
[0070] "Textual information" refers to text-formatted data converted from acoustic data.
[0071] "Summary data" refers to information obtained by analyzing textual information, extracting the important parts, and summarizing them concisely.
[0072] A "chapter" is a segment that structures video information and divides it based on identified important topics.
[0073] A "time marker" is information that represents a specific time within the video data associated with a chapter.
[0074] "Reverse search" is a method for efficiently finding relevant information based on the entered search terms.
[0075] This invention is a system for streamlining the summarization and searching of video content. Specific embodiments of this system are described below.
[0076] The server first receives the video file uploaded by the user. The server uses a media processing library such as FFmpeg to extract audio data from the video. This audio data is then converted into text information using speech recognition services such as Google® Cloud Speech-to-Text or IBM Watson® Speech to Text.
[0077] Based on the converted text information, the server utilizes natural language processing technologies such as the Natural Language Toolkit (NLTK) and spaCy to identify important information and generate summary data. Furthermore, it structures the video based on the identified topics and generates chapters using libraries such as OpenCV. At this stage, each chapter is assigned a corresponding time marker.
[0078] The generated chapter information and summary data are stored in a database with reverse search functionality on the server, enabling quick information retrieval based on user search terms.
[0079] The terminal presents the user with summary data and chapter information retrieved from the server through a user interface. An intuitive interface using React and Vue.js allows users to easily select the information they expect.
[0080] As a concrete example, when a user uploads a one-hour business seminar video, the server digitally processes it, extracts key topics, and presents a summary of approximately 10 minutes. Chapters such as "Marketing Strategy" and "Customer Management" are generated. The terminal presents this information to the user through its interface, supporting quick viewing of, for example, the "Marketing Strategy" section.
[0081] This system utilizes a generative AI model and can use the following prompt: "Summarize a one-hour business seminar video and divide it into chapters, one for each topic."
[0082] In this way, the present invention provides an effective means for users to quickly obtain important information from videos.
[0083] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0084] Step 1:
[0085] The user uploads a video file to the system. The input is a video file, which the server receives. The server checks the video format and verifies that it is a compatible format (e.g., MP4, AVI).
[0086] Step 2:
[0087] The server extracts audio data from the video. The input is a video file, and the output is audio data. The server uses media processing tools such as FFmpeg to separate the audio track from the video and create an audio file. This process makes the audio portion analyzable.
[0088] Step 3:
[0089] The server converts audio data into text information. The input is audio data, and the output is text information. The server uses a speech recognition service (e.g., Google Cloud Speech-to-Text) to convert the audio to text. This allows the audio content to be treated as text.
[0090] Step 4:
[0091] The server analyzes textual information, identifies important information, and generates summary data. The input is textual information, and the output is summary data. Natural language processing techniques (e.g., Natural Language Toolkit (NLTK)) are used to extract important topics and compress the information. This process condenses lengthy content into a concise summary.
[0092] Step 5:
[0093] The server generates chapters based on identified topics. Input is summary data and topic information, and output is chapters. Using OpenCV or similar tools, video segments are formed and time markers are set. In this way, the video is divided into parts according to its syntax.
[0094] Step 6:
[0095] The terminal presents the user with summary data and chapter information received from the server. The input is the summary data and chapter information. The output is information visualized through a user interface. An intuitively operable interface is built using frameworks such as React. This helps users quickly access important information.
[0096] (Application Example 1)
[0097] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0098] In today's information-saturated world, users face the challenge of efficiently obtaining information of interest from a vast amount of video content. Furthermore, there is a lack of means to quickly acquire information based on specified topics across multiple information platforms. This results in inefficient and time-consuming information retrieval and viewing for users, and a solution to this problem is needed.
[0099] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0100] In this invention, the server includes means for acquiring input video information, means for extracting audio information from the video information and converting the audio information into text information, and means for extracting important information based on the text information and generating summary information. This enables efficient grasping of the key points of a video and allows information acquisition based on a specified topic across multiple information platforms.
[0101] "Video and image information" refers to media data that includes audio and video, and is intended to allow users to obtain information visually and aurally.
[0102] "Acoustic information" refers to audio data extracted from video and image information, and serves as the basis for its conversion into text information.
[0103] "Text information" refers to string data converted using speech recognition technology based on acoustic information, and is used for extracting important information and generating summaries.
[0104] "Summary information" refers to information that concisely summarizes the key points extracted from text information, and is intended to allow for quick understanding of the content of video information.
[0105] A "segment" is a specific section of video information generated based on summary information, designed to allow users to efficiently view the parts that interest them.
[0106] A "time stamp" is a timestamp used to identify the time of video information corresponding to a segment, and is intended to allow for the quick retrieval of specific information.
[0107] A "user interface" is a computer processing environment that allows users to visually view and select summary and segment information.
[0108] An "information platform" is a fundamental technological infrastructure for providing and distributing media content such as video and image information.
[0109] A "search instruction" is a keyword or question that a user enters to find specific data within video or image information.
[0110] The system for realizing this invention efficiently processes video information specified by the user and provides a summary of important information. The system consists of a cloud server and the user's terminal and functions as follows:
[0111] The server first receives video information from the user. It extracts audio information from this video and converts it into text using the Google Cloud Speech-to-Text API or similar tools. Next, it uses natural language processing libraries such as spaCy or NLTK to extract important information from the converted text and generate a summary. This summary is then used to generate segments from the video information.
[0112] Furthermore, based on the generated summary information, the video information is segmented, and each segment is assigned a corresponding timestamp. Users can view this information on an interface through an application installed on their smartphone or other device. When a user selects a segment of interest, the corresponding video information is played.
[0113] For example, if a user uploads a one-hour educational lecture, the system generates segments based on important topics such as "educational philosophy" and "effectiveness of the lesson." The user can then quickly retrieve the necessary information from these segments.
[0114] An example of a prompt message is, "Tell us what topics in this video interest you. We will generate a summary and segments based on that and make them available for you to watch immediately."
[0115] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0116] Step 1:
[0117] The server receives video information specified by the user. The input is a video file uploaded by the user, and the server prepares this file for processing. The video information is sent to the server via streaming or batch processing.
[0118] Step 2:
[0119] The server extracts audio information from the uploaded video data. The input is the video file obtained in step 1, and the output is audio data. The server analyzes the audio portion of the video file to extract the audio information and separates the audio track.
[0120] Step 3:
[0121] The server converts acoustic information into text information. The input is the audio data obtained in the previous step, and the output is the converted text data. The server uses the Google Cloud Speech-to-Text API to perform speech recognition and generate the text information.
[0122] Step 4:
[0123] The server extracts important information from text data and generates a summary. The input is the text data generated in step 3, and the output is the summary. The server uses natural language processing with libraries such as spaCy and NLTK to identify important keywords and sentences and create a summary.
[0124] Step 5:
[0125] The server segments video information based on summary information and adds timestamps. The input is summary information, and the output is segmented video information and timestamps. Each segment corresponds to the summarized content and is timestamped to allow the user to easily find the parts of interest.
[0126] Step 6:
[0127] The terminal receives segment information and timestamps from the server and displays them on the user interface. The input is segment information from the server, and the output is a user-operable GUI display. The terminal provides a visual interface to the user, making it easier to select videos for each segment.
[0128] Step 7:
[0129] The user selects a segment of interest from the interface on their device and plays the corresponding video information. The input is the user's selection, and the output is the video playback of the selected segment. The user can efficiently view different segments and obtain the necessary information in a short amount of time.
[0130] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0131] This invention combines a system designed for efficient summarization and viewing support of video content with an emotion engine that recognizes user emotions. This system is particularly effective in providing a personalized video experience tailored to the user's interests and emotions.
[0132] The system operates through the collaboration of servers, terminals, and users. The process is described below.
[0133] Integration of video analysis and emotion recognition
[0134] The server receives video data uploaded by users and converts the audio information into text. Then, it generates a summary of the video using natural language processing and extracts important topics. It also creates a reverse search index based on the generated summary and chapter information. Simultaneously, it utilizes an emotion engine to analyze the user's facial expressions and audio data and recognize emotions in real time.
[0135] Emotion-based content optimization
[0136] The server adjusts the priority of video segments based on the user's emotional data recognized by the emotion engine. This process displays video content that is appropriate for the user's interests and current emotional state. The system also personalizes chapter presentation within videos based on emotional data, making it easier for users to access content they prefer in their specific emotional state.
[0137] Providing a user interface
[0138] The device displays summary information, chapter information, and emotion-based content suggestions sent from the server in its user interface. Using this interface, users can select and watch video segments optimized for their emotions.
[0139] Specific example
[0140] For example, suppose educational videos are available. In this system, if a user's mood deteriorates while watching the video, the server detects this using an emotion engine and prioritizes displaying content or chapters that boost motivation. Conversely, if the user is excited, it can suggest more detailed and technical segments. The device provides a user-friendly interface, ensuring an optimal viewing experience tailored to the user's emotional state.
[0141] This invention enables efficient and highly satisfying information acquisition by providing videos that reflect the user's emotional state.
[0142] The following describes the processing flow.
[0143] Step 1:
[0144] The user uploads video data to their device. The device then prepares to send the file to the server.
[0145] Step 2:
[0146] The server analyzes the received video file and extracts the audio track. Using speech recognition technology, this audio is converted into text data, and the information within the video is transcribed into text.
[0147] Step 3:
[0148] The server performs natural language processing based on the text information to extract important keywords and topics. This information is then used to generate a summary of the video.
[0149] Step 4:
[0150] The server generates video chapters based on summary information. Each chapter is timestamped, allowing viewers to directly access segments of interest.
[0151] Step 5:
[0152] The device captures the user's facial expressions and voice through its camera and microphone, and sends this data to a server.
[0153] Step 6:
[0154] The server uses an emotion engine to recognize the user's emotional state in real time from received facial expression and voice data. It then analyzes the emotional data to determine the user's current emotions.
[0155] Step 7:
[0156] The server reconstructs chapters and summary information optimized for the user based on the emotions it perceives. It also adjusts the priority of the content presented according to the emotions.
[0157] Step 8:
[0158] The device displays optimized summaries and chapter information in the user interface. Based on this information, users can select and watch video segments that match their emotional responses.
[0159] This entire process allows users to have an effective and engaging video viewing experience with content that matches their emotional state.
[0160] (Example 2)
[0161] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".
[0162] The increasing volume of digital video content makes it difficult for users to efficiently obtain information that matches their interests and circumstances. Furthermore, the lack of personalized video experiences tailored to individual user emotions and interests leads to stress and decreased satisfaction. Solving these challenges and enabling users to enjoy efficient and highly satisfying video viewing experiences is essential.
[0163] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0164] In this invention, the server includes means for acquiring input video data, means for extracting audio data from the video data and converting the audio data into text data, and means for extracting important information based on the text data and generating summary information. This enables the provision of customized content for each user and efficient video viewing tailored to their emotional state.
[0165] "Video data" refers to files containing visual information that are stored or transmitted in electronic format, and may include videos and animations.
[0166] "Audio data" refers to digital information used to electronically store or transmit audio, and includes the corresponding audio portion within video data.
[0167] "Text data" refers to information in text format converted from audio data, representing the content of the audio in a linguistic way.
[0168] "Summary information" refers to information that expresses the key points and main topics extracted from text data in a shortened form.
[0169] A "chapter" refers to a specific segment within video data that has been organized to make it easily accessible to users.
[0170] A "recognition engine" refers to a set of algorithms and software used to analyze a user's emotional state, identifying emotions based on the input data.
[0171] A "segment" refers to a section of video data that is divided into parts, and each segment may contain different content or topics.
[0172] "Timestamp" refers to timestamp information used to identify the time when a particular event or data point occurred.
[0173] A "user interface" refers to a screen or input device that allows a user to interact with a system, enabling operations and information provision.
[0174] This invention is a system that optimizes video data according to the user's emotional state to provide a personalized viewing experience. Its embodiments are described in detail below.
[0175] The server receives video data transmitted digitally from the user. HTTPS, a common data transfer protocol, is used for communication to receive the data. Audio data is separated from the received video data, and the audio data is converted into text data using a speech recognition API. Examples of APIs used here include the Google Cloud Speech-to-Text API.
[0176] The server uses natural language processing (NLP) techniques to extract important information based on text data acquired through speech recognition. This process utilizes NLP libraries such as spaCy and NLTK. It generates summary information from the important data and then uses that information to generate chapters for the video data.
[0177] In parallel, the server extracts frames from the user's video data and performs emotion recognition using an image processing library (e.g., OpenCV) and an audio emotion analysis tool (e.g., DeepFace). Based on the emotion recognition results, the order and display of the video data segments are adjusted. This process dynamically provides content optimized for the user's emotional state.
[0178] The device provides users with personalized content suggestions based on summary information, chapter information, and sentiment transmitted from the server. The user interface used here is built using JavaScript frameworks such as React and Vue.js and is designed to accurately reflect the user's choices.
[0179] For example, consider the case of providing educational videos. If a user begins to lose interest while watching, the server uses a recognition engine to detect the user's emotional change and prioritizes displaying chapters that restore their motivation. In this way, content that fits the user's situation and emotions can be smoothly delivered. An example of a prompt would be: "Suggest content to display when the emotion engine recognizes that the user is relaxed."
[0180] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0181] Step 1:
[0182] The server receives video data uploaded by the user from their device. In this process, data transfer occurs when the user selects a file and sends it to the server. The input data consists of video files selected by the user. The output is the video data stored on the server.
[0183] Step 2:
[0184] The server separates the audio data from the received video data. This is done by analyzing the media file using a video processing library. The input is the stored video data. The output is stored as audio data.
[0185] Step 3:
[0186] The server sends the extracted audio data to a speech recognition API and converts it into text data. Specifically, it calls the API to send audio and receives a text data response. The input is audio data, and the output is converted text data.
[0187] Step 4:
[0188] The server passes the text data to a natural language processing engine, which generates summary information and extracts important topics. This is achieved by using NLP techniques to extract key phrases and summarize the text. The input is text data, and the output is summary information and topic data.
[0189] Step 5:
[0190] The server analyzes facial expression data from the user to recognize emotions. This process uses frames from video data as input to an emotion recognition model. The input is the user's facial expression data, and the output is the recognized emotion information.
[0191] Step 6:
[0192] The server adjusts the chapter priority of video data based on sentiment information and selects content to provide the optimal viewing experience. It analyzes sentiment data and chapter information and changes the display priority. The input is sentiment information and chapter information, and the output is a list of optimized content.
[0193] Step 7:
[0194] The terminal displays summary information, chapter information, and content suggestions sent from the server on the user interface. The user selects a suggested chapter through the interface and starts playback. The input is information data from the server, and the output is the user's chapter selection and playback.
[0195] (Application Example 2)
[0196] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as a "server" and the smart device 14 as a "terminal".
[0197] Viewers are exposed to a vast amount of video content every day, making it difficult to select the right content. Furthermore, providing content that caters to the emotions and interests of individual viewers is challenging, resulting in a uniform viewing experience and a lack of user satisfaction.
[0198] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0199] In this invention, the server includes means for acquiring input video information, means for extracting audio data from the video information and converting it into text information, means for extracting important information based on the text information and generating summary information, means for analyzing the user's facial expression data and audio data to recognize emotions, and means for adjusting the display order of video chapters based on emotions. This makes it possible to provide a more personalized viewing experience that responds to the viewer's emotional state.
[0200] "Video information" refers to digital data that represents the content of a video, and includes audio and video.
[0201] "Audio data" refers to digital data related to sound extracted from video and image information, and typically includes information that is perceived by the human ear as sound.
[0202] "Text information" refers to information obtained by converting audio data into written characters, and is converted into a format that can be understood by humans as natural language.
[0203] "Summary information" is a condensed version of important information extracted from text, designed to help viewers quickly grasp the main content of a video.
[0204] A "chapter" refers to a specific segment within video information, indicating a particular part of the video generated based on summary information.
[0205] "Facial expression data" refers to digital data that indicates a user's emotional state, obtained from the user's facial movements and other factors.
[0206] "Means of recognizing emotions" refers to technologies that analyze a user's emotional state from their facial expressions and voice, and recognize it as digital information.
[0207] "Means of adjusting the display order" refers to the process of optimizing the order in which video chapters are presented to the user based on summary information and sentiment data.
[0208] To realize this application, the program builds a system that provides viewers with personalized video content tailored to their emotional state. The main hardware used is a server for processing video information, which communicates with terminals that collect user facial expressions and audio data.
[0209] The server acquires video and audio data, extracts audio data, and converts it into text. This uses the Google Cloud Speech-to-Text API as its speech recognition technology. Next, it uses the natural language processing library spaCy to generate a summary from the text, and then generates video segments based on that summary.
[0210] Simultaneously, the device uses its built-in camera and microphone to collect user facial expression data and voice, and analyzes this data in real time to recognize emotions. This utilizes image processing libraries such as OpenCV and specific algorithms useful for emotion analysis.
[0211] Based on user sentiment data, the server optimizes content and presents video segments in an order that suits the user. The user interface on the device displays video segments tailored to the user's sentiment, allowing the user to select content to watch according to their interests.
[0212] For example, if a user is feeling stressed, the server will prioritize displaying relaxing content. For instance, a prompt such as "When the user is deemed tired, prioritize displaying relaxing videos" will be generated, and content suggestions will be made to help the user relieve stress.
[0213] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0214] Step 1:
[0215] The server receives video information from the user and extracts audio data. At this stage, the input is the video file from the user, and the output is the audio data separated from the video. The process involves converting the video file to a digital media format and extracting the audio track.
[0216] Step 2:
[0217] The server uses speech recognition technology to convert audio data into text. The input here is the audio data obtained in step 1, and the output is the text information derived from it. The Google Cloud Speech-to-Text API is used to analyze the audio waveform and output the audio as text.
[0218] Step 3:
[0219] The server analyzes text information using natural language processing and generates summary information. In this step, text information is input, and information summarizing the main content of the video is output. The natural language processing library spaCy is used to extract important topics from the text and create a summary.
[0220] Step 4:
[0221] The server generates video chapters based on the summary information. The input is the summary information obtained in step 3, and the output is the chapter-structured video segments. The timeline is divided based on the summary information, and the information related to each segment is organized as chapters.
[0222] Step 5:
[0223] The device collects and analyzes the user's facial expression data and audio in real time. The input here is the user's camera video and audio, and the output is digital data representing the user's emotions. Camera video is captured using OpenCV, and emotion analysis is performed using a specific algorithm.
[0224] Step 6:
[0225] The server optimizes the display order of video segments based on sentiment data. The starting points are the chapter information generated in step 4 and the sentiment data obtained in step 5. The output is a list presenting chapters in an optimized viewing order for the user. The sentiment data is analyzed to prioritize videos based on the user's current state.
[0226] Step 7:
[0227] The device displays emotion-based video segment suggestions on the user interface and accepts user selections. The input is a server-prepared list of chapters, and the output is a list of video segment candidates displayed on the user interface. It visually displays a list of recommended content based on emotion and plays the video corresponding to the user's selection.
[0228] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0229] Data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0230] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.
[0231] [Second Embodiment]
[0232] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.
[0233] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0234] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0235] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.
[0236] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0237] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0238] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0239] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0240] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0241] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0242] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0243] Next, the identification processing performed by the identification processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0244] This invention provides a system for efficiently summarizing video content and enabling quick access to sections of interest. This system is particularly effective for quickly obtaining necessary information from information-rich videos.
[0245] The system is realized through the interaction of servers, terminals, and users. The following describes the processing of the program.
[0246] Video analysis and summarization
[0247] The server receives videos uploaded by users and extracts the audio track. This audio track is converted into text using speech recognition technology. Then, important topics are identified through natural language processing, and summary information is generated. This summary information concisely represents the entire lengthy video.
[0248] Chapter generation and search function
[0249] The server generates video chapters based on identified topics and key points. This allows users to view the video divided into specific segments. Furthermore, a reverse search index is created based on the generated summary and chapter information, making it easier to search for specific information using keywords.
[0250] Providing a user interface
[0251] The device visually displays summary and chapter information delivered from the server to the user. Through this interface, the user can select a section of interest and directly watch the video related to that segment.
[0252] Specific example
[0253] For example, suppose there is a video of a one-hour business seminar. When a user uploads this video to the system, the server processes the video, extracts the important topics, and generates a summary of about 10 minutes. Furthermore, the video is divided into chapters by topic, for example, a chapter titled "Marketing Strategy" is generated. The terminal displays this information on its interface, allowing the user to immediately watch the "Marketing Strategy" chapter.
[0254] This system allows users to quickly and efficiently acquire the information they need in this age of information overload. It is extremely useful not only for personal use but also for business purposes.
[0255] The following describes the processing flow.
[0256] Step 1:
[0257] The user uploads a video file to the system. The terminal receives the file and transfers it to the server.
[0258] Step 2:
[0259] The server retrieves the video file, checks its format and encoding, and converts it to a standard format if necessary.
[0260] Step 3:
[0261] The server extracts the audio track from the converted video. This audio information is then converted into text using speech recognition (ASR) technology.
[0262] Step 4:
[0263] The server applies natural language processing (NLP) to the acquired text information to extract important keywords and topics. This process generates a summary of the video.
[0264] Step 5:
[0265] The server generates chapters for the video based on the extracted key topics. Each chapter is assigned a timestamp indicating its start time.
[0266] Step 6:
[0267] The server organizes summary and chapter information and creates an index to facilitate reverse searching by users. Search is used when users are looking for specific information.
[0268] Step 7:
[0269] The device displays summary and chapter information transmitted from the server on the user interface. This allows the user to select and directly view chapters of interest.
[0270] Step 8:
[0271] Users can use the search function to find chapters containing the information they need using keywords. Search results quickly navigate to relevant summaries and chapters.
[0272] (Example 1)
[0273] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0274] With the increasing volume of video content, it has become difficult for users to quickly and efficiently obtain important information from long videos. In particular, with information-rich videos, there is a need for methods to determine which parts are important to the user and access them efficiently. Furthermore, the inability to quickly obtain necessary information increases the user's time cost.
[0275] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0276] In this invention, the server includes means for acquiring video information, means for converting audio data into text information, and means for identifying important information based on the text information and generating summary data. This enables the user to efficiently access important information and quickly understand the video.
[0277] "Video and image information" refers to digital data consisting of video and audio.
[0278] "Audio data" refers to the sound elements included as part of video information, and is data expressed as an audio signal.
[0279] "Character information" refers to data in text format converted based on acoustic data.
[0280] "Summary data" refers to information obtained by analyzing character information, extracting important parts, and presenting a concise summary.
[0281] "Chapter" refers to segments obtained by structuring moving image information and dividing it based on identified important topics.
[0282] "Time marker" refers to information representing a specific time within moving image information related to a chapter.
[0283] "Reverse search" is a method for efficiently searching for relevant information based on an input search term.
[0284] The present invention is a system for improving the summarization and search of video content. Specific embodiments of this system will be described below.
[0285] The server first receives a video file uploaded by a user. The server extracts acoustic data from the video using a media processing library such as FFmpeg. This acoustic data is converted into character information using a speech recognition service such as Google Cloud Speech-to-Text or IBM Watson Speech to Text.
[0286] Based on the converted character information, the server utilizes natural language processing technologies such as the Natural Language Toolkit (NLTK) and spaCy to identify important information and generate summary data. Furthermore, the server structures the video based on the identified topics and generates chapters using a library such as OpenCV. At this time, corresponding time markers are assigned to each chapter.
[0287] The generated chapter information and summary data are stored in a database with reverse search functionality on the server, enabling quick information retrieval based on user search terms.
[0288] The terminal presents the user with summary data and chapter information retrieved from the server through a user interface. An intuitive interface using React and Vue.js allows users to easily select the information they expect.
[0289] As a concrete example, when a user uploads a one-hour business seminar video, the server digitally processes it, extracts key topics, and presents a summary of approximately 10 minutes. Chapters such as "Marketing Strategy" and "Customer Management" are generated. The terminal presents this information to the user through its interface, supporting quick viewing of, for example, the "Marketing Strategy" section.
[0290] This system utilizes a generative AI model and can use the following prompt: "Summarize a one-hour business seminar video and divide it into chapters, one for each topic."
[0291] In this way, the present invention provides an effective means for users to quickly obtain important information from videos.
[0292] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0293] Step 1:
[0294] The user uploads a video file to the system. The input is a video file, which the server receives. The server checks the video format and verifies that it is a compatible format (e.g., MP4, AVI).
[0295] Step 2:
[0296] The server extracts audio data from the video. The input is a video file, and the output is audio data. The server uses media processing tools such as FFmpeg to separate the audio track from the video and create an audio file. This process makes the audio portion analyzable.
[0297] Step 3:
[0298] The server converts audio data into text information. The input is audio data, and the output is text information. The server uses a speech recognition service (e.g., Google Cloud Speech-to-Text) to convert the audio to text. This allows the audio content to be treated as text.
[0299] Step 4:
[0300] The server analyzes textual information, identifies important information, and generates summary data. The input is textual information, and the output is summary data. Natural language processing techniques (e.g., Natural Language Toolkit (NLTK)) are used to extract important topics and compress the information. This process condenses lengthy content into a concise summary.
[0301] Step 5:
[0302] The server generates chapters based on identified topics. Input is summary data and topic information, and output is chapters. Using OpenCV or similar tools, video segments are formed and time markers are set. In this way, the video is divided into parts according to its syntax.
[0303] Step 6:
[0304] The terminal presents the summary data and chapter information received from the server to the user. The input is the summary data and chapter information. The output is the information visualized via the user interface. Through a framework such as React, a user-friendly interface that can be intuitively operated is constructed. This supports the user to quickly access important content.
[0305] (Application Example 1)
[0306] Next, Application Example 1 will be described. In the following description, the data processing device 12 is referred to as the "server", and the smart glasses 214 are referred to as the "terminal".
[0307] In modern times when the amount of information is increasing, there is a problem that it is difficult for users to efficiently obtain the information they are interested in from a large number of video contents. In addition, there is a lack of means to quickly obtain information based on a specified topic across multiple information platforms. This makes the user's information search and viewing inefficient, and it is necessary to solve the problem of taking time and effort.
[0308] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0309] In this invention, the server includes means for acquiring the input moving image information, means for extracting acoustic information from the moving image information and converting the acoustic information into text information, and means for extracting important information based on the text information and generating summary information. This enables efficient grasping of the key points of the video and acquisition of information based on a specified topic across multiple information platforms.
[0310] "Moving image information" is media data including audio and video, which is for users to obtain information visually and auditorily.
[0311] "Acoustic information" is audio data extracted from moving image information and serves as the basis for conversion into text information.
[0312] "Text information" refers to string data converted using speech recognition technology based on acoustic information, and is used for extracting important information and generating summaries.
[0313] "Summary information" refers to information that concisely summarizes the key points extracted from text information, and is intended to allow for quick understanding of the content of video and image information.
[0314] A "segment" is a specific section of video information generated based on summary information, designed to allow users to efficiently view the parts that interest them.
[0315] A "time stamp" is a timestamp used to identify the time of video information corresponding to a segment, and is intended to allow for the quick retrieval of specific information.
[0316] A "user interface" is a computer processing environment that allows users to visually view and select summary and segment information.
[0317] An "information platform" is a fundamental technological infrastructure for providing and distributing media content such as video and image information.
[0318] A "search instruction" is a keyword or question that a user enters to find specific data within video or image information.
[0319] The system for realizing this invention efficiently processes video information specified by the user and provides a summary of important information. The system consists of a cloud server and the user's terminal and functions as follows:
[0320] The server first receives video information from the user. It extracts audio information from this video and converts it into text using the Google Cloud Speech-to-Text API or similar tools. Next, it uses natural language processing libraries such as spaCy or NLTK to extract important information from the converted text and generate a summary. This summary is then used to generate segments from the video information.
[0321] Furthermore, based on the generated summary information, the video information is segmented, and each segment is assigned a corresponding timestamp. Users can view this information on an interface through an application installed on their smartphone or other device. When a user selects a segment of interest, the corresponding video information is played.
[0322] For example, if a user uploads a one-hour educational lecture, the system generates segments based on important topics such as "educational philosophy" and "effectiveness of the lesson." The user can then quickly retrieve the necessary information from these segments.
[0323] An example of a prompt message is, "Tell us what topics in this video interest you. We will generate a summary and segments based on that and make them available for you to watch immediately."
[0324] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0325] Step 1:
[0326] The server receives video information specified by the user. The input is a video file uploaded by the user, and the server prepares this file for processing. The video information is sent to the server via streaming or batch processing.
[0327] Step 2:
[0328] The server extracts audio information from the uploaded video data. The input is the video file obtained in step 1, and the output is audio data. The server analyzes the audio portion of the video file to extract the audio information and separates the audio track.
[0329] Step 3:
[0330] The server converts acoustic information into text information. The input is the audio data obtained in the previous step, and the output is the converted text data. The server uses the Google Cloud Speech-to-Text API to perform speech recognition and generate the text information.
[0331] Step 4:
[0332] The server extracts important information from text data and generates a summary. The input is the text data generated in step 3, and the output is the summary. The server uses natural language processing with libraries such as spaCy and NLTK to identify important keywords and sentences and create a summary.
[0333] Step 5:
[0334] The server segments video information based on summary information and adds timestamps. The input is summary information, and the output is segmented video information and timestamps. Each segment corresponds to the summarized content and is timestamped to allow the user to easily find the parts of interest.
[0335] Step 6:
[0336] The terminal receives segment information and timestamps from the server and displays them on the user interface. The input is segment information from the server, and the output is a user-operable GUI display. The terminal provides a visual interface to the user, making it easier to select videos for each segment.
[0337] Step 7:
[0338] The user selects a segment of interest from the interface on their device and plays the corresponding video information. The input is the user's selection, and the output is the video playback of the selected segment. The user can efficiently view different segments and obtain the necessary information in a short amount of time.
[0339] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0340] This invention combines a system designed for efficient summarization and viewing support of video content with an emotion engine that recognizes user emotions. This system is particularly effective in providing a personalized video experience tailored to the user's interests and emotions.
[0341] The system operates through the collaboration of servers, terminals, and users. The process is described below.
[0342] Integration of video analysis and emotion recognition
[0343] The server receives video data uploaded by users and converts the audio information into text. Then, it generates a summary of the video using natural language processing and extracts important topics. It also creates a reverse search index based on the generated summary and chapter information. Simultaneously, it utilizes an emotion engine to analyze the user's facial expressions and audio data and recognize emotions in real time.
[0344] Emotion-based content optimization
[0345] The server adjusts the priority of video segments based on the user's emotional data recognized by the emotion engine. This process displays video content that is appropriate for the user's interests and current emotional state. The system also personalizes chapter presentation within videos based on emotional data, making it easier for users to access content they prefer in their specific emotional state.
[0346] Providing a user interface
[0347] The device displays summary information, chapter information, and emotion-based content suggestions sent from the server in its user interface. Using this interface, users can select and watch video segments optimized for their emotions.
[0348] Specific example
[0349] For example, suppose educational videos are available. In this system, if a user's mood deteriorates while watching the video, the server detects this using an emotion engine and prioritizes displaying content or chapters that boost motivation. Conversely, if the user is excited, it can suggest more detailed and technical segments. The device provides a user-friendly interface, ensuring an optimal viewing experience tailored to the user's emotional state.
[0350] This invention enables efficient and highly satisfying information acquisition by providing videos that reflect the user's emotional state.
[0351] The following describes the processing flow.
[0352] Step 1:
[0353] The user uploads video data to their device. The device then prepares to send the file to the server.
[0354] Step 2:
[0355] The server analyzes the received video file and extracts the audio track. Using speech recognition technology, this audio is converted into text data, and the information within the video is transcribed into text.
[0356] Step 3:
[0357] The server performs natural language processing based on the text information to extract important keywords and topics. This information is then used to generate a summary of the video.
[0358] Step 4:
[0359] The server generates video chapters based on summary information. Each chapter is timestamped, allowing viewers to directly access segments of interest.
[0360] Step 5:
[0361] The device captures the user's facial expressions and voice through its camera and microphone, and sends this data to a server.
[0362] Step 6:
[0363] The server uses an emotion engine to recognize the user's emotional state in real time from received facial expression and voice data. It then analyzes the emotional data to determine the user's current emotions.
[0364] Step 7:
[0365] The server reconstructs chapters and summaries optimized for the user based on the emotions it perceives. It also adjusts the priority of the content presented according to the user's emotions.
[0366] Step 8:
[0367] The device displays optimized summaries and chapter information in the user interface. Based on this information, users can select and watch video segments that match their emotional responses.
[0368] This entire process allows users to have an effective and engaging video viewing experience with content that matches their emotional state.
[0369] (Example 2)
[0370] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0371] The increasing volume of digital video content makes it difficult for users to efficiently obtain information that matches their interests and circumstances. Furthermore, the lack of personalized video experiences tailored to individual user emotions and interests leads to stress and decreased satisfaction. Solving these challenges and enabling users to enjoy efficient and highly satisfying video viewing experiences is essential.
[0372] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0373] In this invention, the server includes means for acquiring input video data, means for extracting audio data from the video data and converting the audio data into text data, and means for extracting important information based on the text data and generating summary information. This enables the provision of customized content for each user and efficient video viewing tailored to their emotional state.
[0374] "Video data" refers to files containing visual information that are stored or transmitted in electronic format, and may include videos and animations.
[0375] "Audio data" refers to digital information used to electronically store or transmit audio, and includes the corresponding audio portion within video data.
[0376] "Text data" refers to information in text format converted from audio data, representing the content of the audio in a linguistic way.
[0377] "Summary information" refers to information that expresses the key points and main topics extracted from text data in a shortened form.
[0378] A "chapter" refers to a specific segment within video data that has been organized to make it easily accessible to users.
[0379] A "recognition engine" refers to a set of algorithms and software used to analyze a user's emotional state, identifying emotions based on the input data.
[0380] A "segment" refers to a section of video data that is divided into parts, and each segment may contain different content or topics.
[0381] "Timestamp" refers to timestamp information used to identify the time when a particular event or data point occurred.
[0382] A "user interface" refers to a screen or input device that allows a user to interact with a system, enabling operations and information provision.
[0383] This invention is a system that optimizes video data according to the user's emotional state to provide a personalized viewing experience. Its embodiments are described in detail below.
[0384] The server receives video data transmitted digitally from the user. HTTPS, a common data transfer protocol, is used for communication to receive the data. Audio data is separated from the received video data, and the audio data is converted into text data using a speech recognition API. Examples of APIs used here include the Google Cloud Speech-to-Text API.
[0385] The server uses natural language processing (NLP) techniques to extract important information based on text data acquired through speech recognition. This process utilizes NLP libraries such as spaCy and NLTK. It generates summary information from the important data and then uses that information to generate chapters for the video data.
[0386] In parallel, the server extracts frames from the user's video data and performs emotion recognition using an image processing library (e.g., OpenCV) and an audio emotion analysis tool (e.g., DeepFace). Based on the emotion recognition results, the order and display of the video data segments are adjusted. This process dynamically provides content optimized for the user's emotional state.
[0387] The device provides users with personalized content suggestions based on summary information, chapter information, and sentiment sent from the server. The user interface used here is built using JavaScript frameworks such as React and Vue.js and is designed to accurately reflect the user's choices.
[0388] For example, consider the case of providing educational videos. If a user begins to lose interest while watching, the server uses a recognition engine to detect the user's emotional change and prioritizes displaying chapters that restore their motivation. In this way, content that fits the user's situation and emotions can be smoothly delivered. An example of a prompt would be: "Suggest content to display when the emotion engine recognizes that the user is relaxed."
[0389] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0390] Step 1:
[0391] The server receives video data uploaded by the user from their device. In this process, data transfer occurs when the user selects a file and sends it to the server. The input data consists of video files selected by the user. The output is the video data stored on the server.
[0392] Step 2:
[0393] The server separates the audio data from the received video data. This is done by analyzing the media file using a video processing library. The input is the stored video data. The output is stored as audio data.
[0394] Step 3:
[0395] The server sends the extracted audio data to a speech recognition API and converts it into text data. Specifically, it calls the API to send audio and receives a text data response. The input is audio data, and the output is converted text data.
[0396] Step 4:
[0397] The server passes the text data to a natural language processing engine, which generates summary information and extracts important topics. This is achieved by using NLP techniques to extract key phrases and summarize the text. The input is text data, and the output is summary information and topic data.
[0398] Step 5:
[0399] The server analyzes facial expression data from the user to recognize emotions. This process uses frames from video data as input to an emotion recognition model. The input is the user's facial expression data, and the output is the recognized emotion information.
[0400] Step 6:
[0401] The server adjusts the chapter priority of video data based on sentiment information and selects content to provide the optimal viewing experience. It analyzes sentiment data and chapter information and changes the display priority. The input is sentiment information and chapter information, and the output is a list of optimized content.
[0402] Step 7:
[0403] The terminal displays summary information, chapter information, and content suggestions sent from the server on the user interface. The user selects a suggested chapter through the interface and starts playback. The input is information data from the server, and the output is the user's chapter selection and playback.
[0404] (Application Example 2)
[0405] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0406] Viewers are exposed to a vast amount of video content every day, making it difficult to select the right content. Furthermore, providing content that caters to the emotions and interests of individual viewers is challenging, resulting in a uniform viewing experience and a lack of user satisfaction.
[0407] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0408] In this invention, the server includes means for acquiring input video information, means for extracting audio data from the video information and converting it into text information, means for extracting important information based on the text information and generating summary information, means for analyzing the user's facial expression data and audio data to recognize emotions, and means for adjusting the display order of video chapters based on emotions. This makes it possible to provide a more personalized viewing experience that responds to the viewer's emotional state.
[0409] "Video information" refers to digital data that represents the content of a video, and includes audio and video.
[0410] "Audio data" refers to digital data related to sound extracted from video and image information, and typically includes information that is perceived by the human ear as sound.
[0411] "Text information" refers to information obtained by converting audio data into written characters, and is converted into a format that can be understood by humans as natural language.
[0412] "Summary information" is a condensed version of important information extracted from text, designed to help viewers quickly grasp the main content of a video.
[0413] A "chapter" refers to a specific segment within video information, indicating a particular part of the video generated based on summary information.
[0414] "Facial expression data" refers to digital data that indicates a user's emotional state, obtained from the user's facial movements and other factors.
[0415] "Means of recognizing emotions" refers to technologies that analyze a user's emotional state from their facial expressions and voice, and recognize it as digital information.
[0416] "Means of adjusting the display order" refers to the process of optimizing the order in which video chapters are presented to the user based on summary information and sentiment data.
[0417] To realize this application, the program builds a system that provides viewers with personalized video content tailored to their emotional state. The main hardware used is a server for processing video information, which communicates with terminals that collect user facial expressions and audio data.
[0418] The server acquires video and audio data, extracts audio data, and converts it into text. This uses the Google Cloud Speech-to-Text API as its speech recognition technology. Next, it uses the natural language processing library spaCy to generate a summary from the text, and then generates video segments based on that summary.
[0419] Simultaneously, the device uses its built-in camera and microphone to collect user facial expression data and voice, and analyzes this data in real time to recognize emotions. This utilizes image processing libraries such as OpenCV and specific algorithms useful for emotion analysis.
[0420] Based on user sentiment data, the server optimizes content and presents video segments in an order that suits the user. The user interface on the device displays video segments tailored to the user's sentiment, allowing the user to select content to watch according to their interests.
[0421] For example, if a user is feeling stressed, the server will prioritize displaying relaxing content. For instance, a prompt such as "When the user is deemed tired, prioritize displaying relaxing videos" will be generated, and content suggestions will be made to help the user relieve stress.
[0422] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0423] Step 1:
[0424] The server receives video information from the user and extracts audio data. At this stage, the input is the video file from the user, and the output is the audio data separated from the video. The process involves converting the video file to a digital media format and extracting the audio track.
[0425] Step 2:
[0426] The server uses speech recognition technology to convert audio data into text. The input here is the audio data obtained in step 1, and the output is the text information derived from it. The Google Cloud Speech-to-Text API is used to analyze the audio waveform and output the audio as text.
[0427] Step 3:
[0428] The server analyzes text information using natural language processing and generates summary information. In this step, text information is input, and information summarizing the main content of the video is output. The natural language processing library spaCy is used to extract important topics from the text and create a summary.
[0429] Step 4:
[0430] The server generates video chapters based on the summary information. The input is the summary information obtained in step 3, and the output is the chapter-structured video segments. The timeline is divided based on the summary information, and the information related to each segment is organized as chapters.
[0431] Step 5:
[0432] The device collects and analyzes the user's facial expression data and audio in real time. The input here is the user's camera video and audio, and the output is digital data representing the user's emotions. Camera video is captured using OpenCV, and emotion analysis is performed using a specific algorithm.
[0433] Step 6:
[0434] The server optimizes the display order of video segments based on sentiment data. The starting points are the chapter information generated in step 4 and the sentiment data obtained in step 5. The output is a list presenting chapters in an optimized viewing order for the user. The sentiment data is analyzed to prioritize videos based on the user's current state.
[0435] Step 7:
[0436] The device displays emotion-based video segment suggestions on the user interface and accepts user selections. The input is a server-prepared list of chapters, and the output is a list of video segment candidates displayed on the user interface. It visually displays a list of recommended content based on emotion and plays the video corresponding to the user's selection.
[0437] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0438] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0439] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.
[0440] [Third Embodiment]
[0441] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.
[0442] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.
[0443] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0444] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.
[0445] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0446] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0447] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0448] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0449] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0450] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0451] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0452] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the headset terminal 314 will be referred to as the "terminal".
[0453] This invention provides a system for efficiently summarizing video content and enabling quick access to sections of interest. This system is particularly effective for quickly obtaining necessary information from information-rich videos.
[0454] The system is realized through the interaction of servers, terminals, and users. The following describes the processing of the program.
[0455] Video analysis and summarization
[0456] The server receives videos uploaded by users and extracts the audio track. This audio track is converted into text using speech recognition technology. Then, important topics are identified through natural language processing, and summary information is generated. This summary information concisely represents the entire lengthy video.
[0457] Chapter generation and search function
[0458] The server generates video chapters based on identified topics and key points. This allows users to view the video divided into specific segments. Furthermore, a reverse search index is created based on the generated summary and chapter information, making it easier to search for specific information using keywords.
[0459] Providing a user interface
[0460] The device visually displays summary and chapter information delivered from the server to the user. Through this interface, the user can select a section of interest and directly watch the video related to that segment.
[0461] Specific example
[0462] For example, suppose there is a video of a one-hour business seminar. When a user uploads this video to the system, the server processes the video, extracts the important topics, and generates a summary of about 10 minutes. Furthermore, the video is divided into chapters by topic, for example, a chapter titled "Marketing Strategy" is generated. The terminal displays this information on its interface, allowing the user to immediately watch the "Marketing Strategy" chapter.
[0463] This system allows users to quickly and efficiently acquire the information they need in this age of information overload. It is extremely useful not only for personal use but also for business purposes.
[0464] The following describes the processing flow.
[0465] Step 1:
[0466] The user uploads a video file to the system. The terminal receives the file and transfers it to the server.
[0467] Step 2:
[0468] The server retrieves the video file, checks its format and encoding, and converts it to a standard format if necessary.
[0469] Step 3:
[0470] The server extracts the audio track from the converted video. This audio information is then converted into text using speech recognition (ASR) technology.
[0471] Step 4:
[0472] The server applies natural language processing (NLP) to the acquired text information to extract important keywords and topics. This process generates a summary of the video.
[0473] Step 5:
[0474] The server generates chapters for the video based on the extracted key topics. Each chapter is assigned a timestamp indicating its start time.
[0475] Step 6:
[0476] The server organizes summary and chapter information and creates an index to facilitate reverse searching by users. Search is used when users are looking for specific information.
[0477] Step 7:
[0478] The device displays summary and chapter information transmitted from the server on the user interface. This allows the user to select and directly view chapters of interest.
[0479] Step 8:
[0480] Users can use the search function to find chapters containing the information they need using keywords. Search results quickly navigate to relevant summaries and chapters.
[0481] (Example 1)
[0482] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0483] With the increasing volume of video content, it has become difficult for users to quickly and efficiently obtain important information from long videos. In particular, with information-rich videos, there is a need for methods to determine which parts are important to the user and access them efficiently. Furthermore, the inability to quickly obtain necessary information increases the user's time cost.
[0484] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0485] In this invention, the server includes means for acquiring video information, means for converting audio data into text information, and means for identifying important information based on the text information and generating summary data. This enables the user to efficiently access important information and quickly understand the video.
[0486] "Video and image information" refers to digital data consisting of video and audio.
[0487] "Audio data" refers to the sound elements included as part of video information, and is data expressed as an audio signal.
[0488] "Textual information" refers to text-formatted data converted from acoustic data.
[0489] "Summary data" refers to information obtained by analyzing textual information, extracting the important parts, and summarizing them concisely.
[0490] A "chapter" is a segment that structures video information and divides it based on identified important topics.
[0491] A "time marker" is information that represents a specific time within the video data associated with a chapter.
[0492] "Reverse search" is a method for efficiently finding relevant information based on the entered search terms.
[0493] This invention is a system for streamlining the summarization and searching of video content. Specific embodiments of this system are described below.
[0494] The server first receives the video file uploaded by the user. The server uses a media processing library such as FFmpeg to extract audio data from the video. This audio data is then converted into text using a speech recognition service such as Google Cloud Speech-to-Text or IBM Watson Speech to Text.
[0495] Based on the converted text information, the server utilizes natural language processing technologies such as the Natural Language Toolkit (NLTK) and spaCy to identify important information and generate summary data. Furthermore, it structures the video based on the identified topics and generates chapters using libraries such as OpenCV. At this stage, each chapter is assigned a corresponding time marker.
[0496] The generated chapter information and summary data are stored in a database with reverse search functionality on the server, enabling quick information retrieval based on user search terms.
[0497] The terminal presents the user with summary data and chapter information retrieved from the server through a user interface. An intuitive interface using React and Vue.js allows users to easily select the information they expect.
[0498] As a concrete example, when a user uploads a one-hour business seminar video, the server digitally processes it, extracts key topics, and presents a summary of approximately 10 minutes. Chapters such as "Marketing Strategy" and "Customer Management" are generated. The terminal presents this information to the user through its interface, supporting quick viewing of, for example, the "Marketing Strategy" section.
[0499] This system utilizes a generative AI model and can use the following prompt: "Summarize a one-hour business seminar video and divide it into chapters, one for each topic."
[0500] In this way, the present invention provides an effective means for users to quickly obtain important information from videos.
[0501] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0502] Step 1:
[0503] The user uploads a video file to the system. The input is a video file, which the server receives. The server checks the video format and verifies that it is a compatible format (e.g., MP4, AVI).
[0504] Step 2:
[0505] The server extracts audio data from the video. The input is a video file, and the output is audio data. The server uses media processing tools such as FFmpeg to separate the audio track from the video and create an audio file. This process makes the audio portion analyzable.
[0506] Step 3:
[0507] The server converts audio data into text information. The input is audio data, and the output is text information. The server uses a speech recognition service (e.g., Google Cloud Speech-to-Text) to convert the audio to text. This allows the audio content to be treated as text.
[0508] Step 4:
[0509] The server analyzes textual information, identifies important information, and generates summary data. The input is textual information, and the output is summary data. Natural language processing techniques (e.g., Natural Language Toolkit (NLTK)) are used to extract important topics and compress the information. This process condenses lengthy content into a concise summary.
[0510] Step 5:
[0511] The server generates chapters based on identified topics. Input is summary data and topic information, and output is chapters. Using OpenCV or similar tools, video segments are formed and time markers are set. In this way, the video is divided into parts according to its syntax.
[0512] Step 6:
[0513] The terminal presents the user with summary data and chapter information received from the server. The input is the summary data and chapter information. The output is information visualized through a user interface. An intuitively operable interface is built using frameworks such as React. This helps users quickly access important information.
[0514] (Application Example 1)
[0515] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0516] In today's information-saturated world, users face the challenge of efficiently obtaining information of interest from a vast amount of video content. Furthermore, there is a lack of means to quickly acquire information based on specified topics across multiple information platforms. This results in inefficient and time-consuming information retrieval and viewing for users, and a solution to this problem is needed.
[0517] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0518] In this invention, the server includes means for acquiring input video information, means for extracting audio information from the video information and converting the audio information into text information, and means for extracting important information based on the text information and generating summary information. This enables efficient grasping of the key points of a video and allows information acquisition based on a specified topic across multiple information platforms.
[0519] "Video and image information" refers to media data that includes audio and video, and is intended to allow users to obtain information visually and aurally.
[0520] "Acoustic information" refers to audio data extracted from video and image information, and serves as the basis for its conversion into text information.
[0521] "Text information" refers to string data converted using speech recognition technology based on acoustic information, and is used for extracting important information and generating summaries.
[0522] "Summary information" refers to information that concisely summarizes the key points extracted from text information, and is intended to allow for quick understanding of the content of video and image information.
[0523] A "segment" is a specific section of video information generated based on summary information, designed to allow users to efficiently view the parts that interest them.
[0524] A "time stamp" is a timestamp used to identify the time of video information corresponding to a segment, and is intended to allow for the quick retrieval of specific information.
[0525] A "user interface" is a computer processing environment that allows users to visually view and select summary and segment information.
[0526] An "information platform" is a fundamental technological infrastructure for providing and distributing media content such as video and image information.
[0527] A "search instruction" is a keyword or question that a user enters to find specific data within video or image information.
[0528] The system for realizing this invention efficiently processes video information specified by the user and provides a summary of important information. The system consists of a cloud server and the user's terminal and functions as follows:
[0529] The server first receives video information from the user. It extracts audio information from this video and converts it into text using the Google Cloud Speech-to-Text API or similar tools. Next, it uses natural language processing libraries such as spaCy or NLTK to extract important information from the converted text and generate a summary. This summary is then used to generate segments from the video information.
[0530] Furthermore, based on the generated summary information, the video information is segmented, and each segment is assigned a corresponding timestamp. Users can view this information on an interface through an application installed on their smartphone or other device. When a user selects a segment of interest, the corresponding video information is played.
[0531] For example, if a user uploads a one-hour educational lecture, the system generates segments based on important topics such as "educational philosophy" and "effectiveness of the lesson." The user can then quickly retrieve the necessary information from these segments.
[0532] An example of a prompt message is, "Tell us what topics in this video interest you. We will generate a summary and segments based on that and make them available for you to watch immediately."
[0533] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0534] Step 1:
[0535] The server receives video information specified by the user. The input is a video file uploaded by the user, and the server prepares this file for processing. The video information is sent to the server via streaming or batch processing.
[0536] Step 2:
[0537] The server extracts audio information from the uploaded video data. The input is the video file obtained in step 1, and the output is audio data. The server analyzes the audio portion of the video file to extract the audio information and separates the audio track.
[0538] Step 3:
[0539] The server converts acoustic information into text information. The input is the audio data obtained in the previous step, and the output is the converted text data. The server uses the Google Cloud Speech-to-Text API to perform speech recognition and generate the text information.
[0540] Step 4:
[0541] The server extracts important information from text data and generates a summary. The input is the text data generated in step 3, and the output is the summary. The server uses natural language processing with libraries such as spaCy and NLTK to identify important keywords and sentences and create a summary.
[0542] Step 5:
[0543] The server segments video information based on summary information and adds timestamps. The input is summary information, and the output is segmented video information and timestamps. Each segment corresponds to the summarized content and is timestamped to allow the user to easily find the parts of interest.
[0544] Step 6:
[0545] The terminal receives segment information and timestamps from the server and displays them on the user interface. The input is segment information from the server, and the output is a user-operable GUI display. The terminal provides a visual interface to the user, making it easier to select videos for each segment.
[0546] Step 7:
[0547] The user selects a segment of interest from the interface on their device and plays the corresponding video information. The input is the user's selection, and the output is the video playback of the selected segment. The user can efficiently view different segments and obtain the necessary information in a short amount of time.
[0548] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0549] This invention combines a system designed for efficient summarization and viewing support of video content with an emotion engine that recognizes user emotions. This system is particularly effective in providing a personalized video experience tailored to the user's interests and emotions.
[0550] The system operates through the collaboration of servers, terminals, and users. The process is described below.
[0551] Integration of video analysis and emotion recognition
[0552] The server receives video data uploaded by users and converts the audio information into text. Then, it generates a summary of the video using natural language processing and extracts important topics. It also creates a reverse search index based on the generated summary and chapter information. Simultaneously, it utilizes an emotion engine to analyze the user's facial expressions and audio data and recognize emotions in real time.
[0553] Emotion-based content optimization
[0554] The server adjusts the priority of video segments based on the user's emotional data recognized by the emotion engine. This process displays video content that is appropriate for the user's interests and current emotional state. The system also personalizes chapter presentation within videos based on emotional data, making it easier for users to access content they prefer in their specific emotional state.
[0555] Providing a user interface
[0556] The device displays summary information, chapter information, and emotion-based content suggestions sent from the server in its user interface. Using this interface, users can select and watch video segments optimized for their emotions.
[0557] Specific example
[0558] For example, suppose educational videos are available. In this system, if a user's mood deteriorates while watching the video, the server detects this using an emotion engine and prioritizes displaying content or chapters that boost motivation. Conversely, if the user is excited, it can suggest more detailed and technical segments. The device provides a user-friendly interface, ensuring an optimal viewing experience tailored to the user's emotional state.
[0559] This invention enables efficient and highly satisfying information acquisition by providing videos that reflect the user's emotional state.
[0560] The following describes the processing flow.
[0561] Step 1:
[0562] The user uploads video data to their device. The device then prepares to send the file to the server.
[0563] Step 2:
[0564] The server analyzes the received video file and extracts the audio track. Using speech recognition technology, this audio is converted into text data, and the information within the video is transcribed into text.
[0565] Step 3:
[0566] The server performs natural language processing based on the text information to extract important keywords and topics. This information is then used to generate a summary of the video.
[0567] Step 4:
[0568] The server generates video chapters based on summary information. Each chapter is timestamped, allowing viewers to directly access segments of interest.
[0569] Step 5:
[0570] The device captures the user's facial expressions and voice through its camera and microphone, and sends this data to a server.
[0571] Step 6:
[0572] The server uses an emotion engine to recognize the user's emotional state in real time from received facial expression and voice data. It then analyzes the emotional data to determine the user's current emotions.
[0573] Step 7:
[0574] The server reconstructs chapters and summaries optimized for the user based on the emotions it perceives. It also adjusts the priority of the content presented according to the user's emotions.
[0575] Step 8:
[0576] The device displays optimized summaries and chapter information in the user interface. Based on this information, users can select and watch video segments that match their emotional responses.
[0577] This entire process allows users to have an effective and engaging video viewing experience with content that matches their emotional state.
[0578] (Example 2)
[0579] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0580] The increasing volume of digital video content makes it difficult for users to efficiently obtain information that matches their interests and circumstances. Furthermore, the lack of personalized video experiences tailored to individual user emotions and interests leads to stress and decreased satisfaction. Solving these challenges and enabling users to enjoy efficient and highly satisfying video viewing experiences is essential.
[0581] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0582] In this invention, the server includes means for acquiring input video data, means for extracting audio data from the video data and converting the audio data into text data, and means for extracting important information based on the text data and generating summary information. This enables the provision of customized content for each user and efficient video viewing tailored to their emotional state.
[0583] "Video data" refers to files containing visual information that are stored or transmitted in electronic format, and may include videos and animations.
[0584] "Audio data" refers to digital information used to electronically store or transmit audio, and includes the corresponding audio portion within video data.
[0585] "Text data" refers to information in text format converted from audio data, representing the content of the audio in a linguistic way.
[0586] "Summary information" refers to information that expresses the key points and main topics extracted from text data in a shortened form.
[0587] A "chapter" refers to a specific segment within video data that has been organized to make it easily accessible to users.
[0588] A "recognition engine" refers to a set of algorithms and software used to analyze a user's emotional state, identifying emotions based on the input data.
[0589] A "segment" refers to a section of video data that is divided into parts, and each segment may contain different content or topics.
[0590] "Timestamp" refers to timestamp information used to identify the time when a particular event or data point occurred.
[0591] A "user interface" refers to a screen or input device that allows a user to interact with a system, enabling operations and information provision.
[0592] This invention is a system that optimizes video data according to the user's emotional state to provide a personalized viewing experience. Its embodiments are described in detail below.
[0593] The server receives video data transmitted digitally from the user. HTTPS, a common data transfer protocol, is used for communication to receive the data. Audio data is separated from the received video data, and the audio data is converted into text data using a speech recognition API. Examples of APIs used here include the Google Cloud Speech-to-Text API.
[0594] The server uses natural language processing (NLP) techniques to extract important information based on text data acquired through speech recognition. This process utilizes NLP libraries such as spaCy and NLTK. It generates summary information from the important data and then uses that information to generate chapters for the video data.
[0595] In parallel, the server extracts frames from the user's video data and performs emotion recognition using an image processing library (e.g., OpenCV) and an audio emotion analysis tool (e.g., DeepFace). Based on the emotion recognition results, the order and display of the video data segments are adjusted. This process dynamically provides content optimized for the user's emotional state.
[0596] The device provides users with personalized content suggestions based on summary information, chapter information, and sentiment sent from the server. The user interface used here is built using JavaScript frameworks such as React and Vue.js and is designed to accurately reflect the user's choices.
[0597] For example, consider the case of providing educational videos. If a user begins to lose interest while watching, the server uses a recognition engine to detect the user's emotional change and prioritizes displaying chapters that restore their motivation. In this way, content that fits the user's situation and emotions can be smoothly delivered. An example of a prompt would be: "Suggest content to display when the emotion engine recognizes that the user is relaxed."
[0598] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0599] Step 1:
[0600] The server receives video data uploaded by the user from their device. In this process, data transfer occurs when the user selects a file and sends it to the server. The input data consists of video files selected by the user. The output is the video data stored on the server.
[0601] Step 2:
[0602] The server separates the audio data from the received video data. This is done by analyzing the media file using a video processing library. The input is the stored video data. The output is stored as audio data.
[0603] Step 3:
[0604] The server sends the extracted audio data to a speech recognition API and converts it into text data. Specifically, it calls the API to send audio and receives a text data response. The input is audio data, and the output is converted text data.
[0605] Step 4:
[0606] The server passes the text data to a natural language processing engine, which generates summary information and extracts important topics. This is achieved by using NLP techniques to extract key phrases and summarize the text. The input is text data, and the output is summary information and topic data.
[0607] Step 5:
[0608] The server analyzes facial expression data from the user to recognize emotions. This process uses frames from video data as input to an emotion recognition model. The input is the user's facial expression data, and the output is the recognized emotion information.
[0609] Step 6:
[0610] The server adjusts the chapter priority of video data based on sentiment information and selects content to provide the optimal viewing experience. It analyzes sentiment data and chapter information and changes the display priority. The input is sentiment information and chapter information, and the output is a list of optimized content.
[0611] Step 7:
[0612] The terminal displays summary information, chapter information, and content suggestions sent from the server on the user interface. The user selects a suggested chapter through the interface and starts playback. The input is information data from the server, and the output is the user's chapter selection and playback.
[0613] (Application Example 2)
[0614] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0615] Viewers are exposed to a vast amount of video content every day, making it difficult to select the right content. Furthermore, providing content that caters to the emotions and interests of individual viewers is challenging, resulting in a uniform viewing experience and a lack of user satisfaction.
[0616] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0617] In this invention, the server includes means for acquiring input video information, means for extracting audio data from the video information and converting it into text information, means for extracting important information based on the text information and generating summary information, means for analyzing the user's facial expression data and audio data to recognize emotions, and means for adjusting the display order of video chapters based on emotions. This makes it possible to provide a more personalized viewing experience that responds to the viewer's emotional state.
[0618] "Video information" refers to digital data that represents the content of a video, and includes audio and video.
[0619] "Audio data" refers to digital data related to sound extracted from video and image information, and typically includes information that is perceived by the human ear as sound.
[0620] "Text information" refers to information obtained by converting audio data into written characters, and is converted into a format that can be understood by humans as natural language.
[0621] "Summary information" is a condensed version of important information extracted from text, designed to help viewers quickly grasp the main content of a video.
[0622] A "chapter" refers to a specific segment within video information, indicating a particular part of the video generated based on summary information.
[0623] "Facial expression data" refers to digital data that indicates a user's emotional state, obtained from the user's facial movements and other factors.
[0624] "Means of recognizing emotions" refers to technologies that analyze a user's emotional state from their facial expressions and voice, and recognize it as digital information.
[0625] "Means of adjusting the display order" refers to the process of optimizing the order in which video chapters are presented to the user based on summary information and sentiment data.
[0626] To realize this application, the program builds a system that provides viewers with personalized video content tailored to their emotional state. The main hardware used is a server for processing video information, which communicates with terminals that collect user facial expressions and audio data.
[0627] The server acquires video and audio data, extracts audio data, and converts it into text. This uses the Google Cloud Speech-to-Text API as its speech recognition technology. Next, it uses the natural language processing library spaCy to generate a summary from the text, and then generates video segments based on that summary.
[0628] Simultaneously, the device uses its built-in camera and microphone to collect user facial expression data and voice, and analyzes this data in real time to recognize emotions. This utilizes image processing libraries such as OpenCV and specific algorithms useful for emotion analysis.
[0629] Based on user sentiment data, the server optimizes content and presents video segments in an order that suits the user. The user interface on the device displays video segments tailored to the user's sentiment, allowing the user to select content to watch according to their interests.
[0630] For example, if a user is feeling stressed, the server will prioritize displaying relaxing content. For instance, a prompt such as "When the user is deemed tired, prioritize displaying relaxing videos" will be generated, and content suggestions will be made to help the user relieve stress.
[0631] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0632] Step 1:
[0633] The server receives video information from the user and extracts audio data. At this stage, the input is the video file from the user, and the output is the audio data separated from the video. The process involves converting the video file to a digital media format and extracting the audio track.
[0634] Step 2:
[0635] The server uses speech recognition technology to convert audio data into text. The input here is the audio data obtained in step 1, and the output is the text information derived from it. The Google Cloud Speech-to-Text API is used to analyze the audio waveform and output the audio as text.
[0636] Step 3:
[0637] The server analyzes text information using natural language processing and generates summary information. In this step, text information is input, and information summarizing the main content of the video is output. The natural language processing library spaCy is used to extract important topics from the text and create a summary.
[0638] Step 4:
[0639] The server generates video chapters based on the summary information. The input is the summary information obtained in step 3, and the output is the chapter-structured video segments. The timeline is divided based on the summary information, and the information related to each segment is organized as chapters.
[0640] Step 5:
[0641] The device collects and analyzes the user's facial expression data and audio in real time. The input here is the user's camera video and audio, and the output is digital data representing the user's emotions. Camera video is captured using OpenCV, and emotion analysis is performed using a specific algorithm.
[0642] Step 6:
[0643] The server optimizes the display order of video segments based on sentiment data. The starting points are the chapter information generated in step 4 and the sentiment data obtained in step 5. The output is a list presenting chapters in an optimized viewing order for the user. The sentiment data is analyzed to prioritize videos based on the user's current state.
[0644] Step 7:
[0645] The device displays emotion-based video segment suggestions on the user interface and accepts user selections. The input is a server-prepared list of chapters, and the output is a list of video segment candidates displayed on the user interface. It visually displays a list of recommended content based on emotion and plays the video corresponding to the user's selection.
[0646] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0647] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0648] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.
[0649] [Fourth Embodiment]
[0650] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.
[0651] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[0652] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0653] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.
[0654] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0655] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0656] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0657] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.
[0658] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0659] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0660] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0661] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0662] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0663] This invention provides a system for efficiently summarizing video content and enabling quick access to sections of interest. This system is particularly effective for quickly obtaining necessary information from information-rich videos.
[0664] The system is realized through the interaction of servers, terminals, and users. The following describes the processing of the program.
[0665] Video analysis and summarization
[0666] The server receives videos uploaded by users and extracts the audio track. This audio track is converted into text using speech recognition technology. Then, important topics are identified through natural language processing, and summary information is generated. This summary information concisely represents the entire lengthy video.
[0667] Chapter generation and search function
[0668] The server generates video chapters based on identified topics and key points. This allows users to view the video divided into specific segments. Furthermore, a reverse search index is created based on the generated summary and chapter information, making it easier to search for specific information using keywords.
[0669] Providing a user interface
[0670] The device visually displays summary and chapter information delivered from the server to the user. Through this interface, the user can select a section of interest and directly watch the video related to that segment.
[0671] Specific example
[0672] For example, suppose there is a video of a one-hour business seminar. When a user uploads this video to the system, the server processes the video, extracts the important topics, and generates a summary of about 10 minutes. Furthermore, the video is divided into chapters by topic, for example, a chapter titled "Marketing Strategy" is generated. The terminal displays this information on its interface, allowing the user to immediately watch the "Marketing Strategy" chapter.
[0673] This system allows users to quickly and efficiently acquire the information they need in this age of information overload. It is extremely useful not only for personal use but also for business purposes.
[0674] The following describes the processing flow.
[0675] Step 1:
[0676] The user uploads a video file to the system. The terminal receives the file and transfers it to the server.
[0677] Step 2:
[0678] The server retrieves the video file, checks its format and encoding, and converts it to a standard format if necessary.
[0679] Step 3:
[0680] The server extracts the audio track from the converted video. This audio information is then converted into text using speech recognition (ASR) technology.
[0681] Step 4:
[0682] The server applies natural language processing (NLP) to the acquired text information to extract important keywords and topics. This process generates a summary of the video.
[0683] Step 5:
[0684] The server generates chapters for the video based on the extracted key topics. Each chapter is assigned a timestamp indicating its start time.
[0685] Step 6:
[0686] The server organizes summary and chapter information and creates an index to facilitate reverse searching by users. Search is used when users are looking for specific information.
[0687] Step 7:
[0688] The device displays summary and chapter information transmitted from the server on the user interface. This allows the user to select and directly view chapters of interest.
[0689] Step 8:
[0690] Users can use the search function to find chapters containing the information they need using keywords. Search results quickly navigate to relevant summaries and chapters.
[0691] (Example 1)
[0692] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0693] With the increasing volume of video content, it has become difficult for users to quickly and efficiently obtain important information from long videos. In particular, with information-rich videos, there is a need for methods to determine which parts are important to the user and access them efficiently. Furthermore, the inability to quickly obtain necessary information increases the user's time cost.
[0694] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0695] In this invention, the server includes means for acquiring video information, means for converting audio data into text information, and means for identifying important information based on the text information and generating summary data. This enables the user to efficiently access important information and quickly understand the video.
[0696] "Video and image information" refers to digital data consisting of video and audio.
[0697] "Audio data" refers to the sound elements included as part of video information, and is data expressed as an audio signal.
[0698] "Textual information" refers to text-formatted data converted from acoustic data.
[0699] "Summary data" refers to information obtained by analyzing textual information, extracting the important parts, and summarizing them concisely.
[0700] A "chapter" is a segment that structures video information and divides it based on identified important topics.
[0701] A "time marker" is information that represents a specific time within the video data associated with a chapter.
[0702] "Reverse search" is a method for efficiently finding relevant information based on the entered search terms.
[0703] This invention is a system for streamlining the summarization and searching of video content. Specific embodiments of this system are described below.
[0704] The server first receives the video file uploaded by the user. The server uses a media processing library such as FFmpeg to extract audio data from the video. This audio data is then converted into text using a speech recognition service such as Google Cloud Speech-to-Text or IBM Watson Speech to Text.
[0705] Based on the converted text information, the server utilizes natural language processing technologies such as the Natural Language Toolkit (NLTK) and spaCy to identify important information and generate summary data. Furthermore, it structures the video based on the identified topics and generates chapters using libraries such as OpenCV. At this stage, each chapter is assigned a corresponding time marker.
[0706] The generated chapter information and summary data are stored in a database with reverse search functionality on the server, enabling quick information retrieval based on user search terms.
[0707] The terminal presents the user with summary data and chapter information retrieved from the server through a user interface. An intuitive interface using React and Vue.js allows users to easily select the information they expect.
[0708] As a concrete example, when a user uploads a one-hour business seminar video, the server digitally processes it, extracts key topics, and presents a summary of approximately 10 minutes. Chapters such as "Marketing Strategy" and "Customer Management" are generated. The terminal presents this information to the user through its interface, supporting quick viewing of, for example, the "Marketing Strategy" section.
[0709] This system utilizes a generative AI model and can use the following prompt: "Summarize a one-hour business seminar video and divide it into chapters, one for each topic."
[0710] In this way, the present invention provides an effective means for users to quickly obtain important information from videos.
[0711] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0712] Step 1:
[0713] The user uploads a video file to the system. The input is a video file, which the server receives. The server checks the video format and verifies that it is a compatible format (e.g., MP4, AVI).
[0714] Step 2:
[0715] The server extracts audio data from the video. The input is a video file, and the output is audio data. The server uses media processing tools such as FFmpeg to separate the audio track from the video and create an audio file. This process makes the audio portion analyzable.
[0716] Step 3:
[0717] The server converts audio data into text information. The input is audio data, and the output is text information. The server uses a speech recognition service (e.g., Google Cloud Speech-to-Text) to convert the audio to text. This allows the audio content to be treated as text.
[0718] Step 4:
[0719] The server analyzes textual information, identifies important information, and generates summary data. The input is textual information, and the output is summary data. Natural language processing techniques (e.g., Natural Language Toolkit (NLTK)) are used to extract important topics and compress the information. This process condenses lengthy content into a concise summary.
[0720] Step 5:
[0721] The server generates chapters based on identified topics. Input is summary data and topic information, and output is chapters. Using OpenCV or similar tools, video segments are formed and time markers are set. In this way, the video is divided into parts according to its syntax.
[0722] Step 6:
[0723] The terminal presents the user with summary data and chapter information received from the server. The input is the summary data and chapter information. The output is information visualized through a user interface. An intuitively operable interface is built using frameworks such as React. This helps users quickly access important information.
[0724] (Application Example 1)
[0725] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0726] In today's information-saturated world, users face the challenge of efficiently obtaining information of interest from a vast amount of video content. Furthermore, there is a lack of means to quickly acquire information based on specified topics across multiple information platforms. This results in inefficient and time-consuming information retrieval and viewing for users, and a solution to this problem is needed.
[0727] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0728] In this invention, the server includes means for acquiring input video information, means for extracting audio information from the video information and converting the audio information into text information, and means for extracting important information based on the text information and generating summary information. This enables efficient grasping of the key points of a video and allows information acquisition based on a specified topic across multiple information platforms.
[0729] "Video and image information" refers to media data that includes audio and video, and is intended to allow users to obtain information visually and aurally.
[0730] "Acoustic information" refers to audio data extracted from video and image information, and serves as the basis for its conversion into text information.
[0731] "Text information" refers to string data converted using speech recognition technology based on acoustic information, and is used for extracting important information and generating summaries.
[0732] "Summary information" refers to information that concisely summarizes the key points extracted from text information, and is intended to allow for quick understanding of the content of video and image information.
[0733] A "segment" is a specific section of video information generated based on summary information, designed to allow users to efficiently view the parts that interest them.
[0734] A "time stamp" is a timestamp used to identify the time of video information corresponding to a segment, and is intended to allow for the quick retrieval of specific information.
[0735] A "user interface" is a computer processing environment that allows users to visually view and select summary and segment information.
[0736] An "information platform" is a fundamental technological infrastructure for providing and distributing media content such as video and image information.
[0737] A "search instruction" is a keyword or question that a user enters to find specific data within video or image information.
[0738] The system for realizing this invention efficiently processes video information specified by the user and provides a summary of important information. The system consists of a cloud server and the user's terminal and functions as follows:
[0739] The server first receives video information from the user. It extracts audio information from this video and converts it into text using the Google Cloud Speech-to-Text API or similar tools. Next, it uses natural language processing libraries such as spaCy or NLTK to extract important information from the converted text and generate a summary. This summary is then used to generate segments from the video information.
[0740] Furthermore, based on the generated summary information, the video information is segmented, and each segment is assigned a corresponding timestamp. Users can view this information on an interface through an application installed on their smartphone or other device. When a user selects a segment of interest, the corresponding video information is played.
[0741] For example, if a user uploads a one-hour educational lecture, the system generates segments based on important topics such as "educational philosophy" and "effectiveness of the lesson." The user can then quickly retrieve the necessary information from these segments.
[0742] An example of a prompt message is, "Tell us what topics in this video interest you. We will generate a summary and segments based on that and make them available for you to watch immediately."
[0743] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0744] Step 1:
[0745] The server receives video information specified by the user. The input is a video file uploaded by the user, and the server prepares this file for processing. The video information is sent to the server via streaming or batch processing.
[0746] Step 2:
[0747] The server extracts audio information from the uploaded video data. The input is the video file obtained in step 1, and the output is audio data. The server analyzes the audio portion of the video file to extract the audio information and separates the audio track.
[0748] Step 3:
[0749] The server converts acoustic information into text information. The input is the audio data obtained in the previous step, and the output is the converted text data. The server uses the Google Cloud Speech-to-Text API to perform speech recognition and generate the text information.
[0750] Step 4:
[0751] The server extracts important information from text data and generates a summary. The input is the text data generated in step 3, and the output is the summary. The server uses natural language processing with libraries such as spaCy and NLTK to identify important keywords and sentences and create a summary.
[0752] Step 5:
[0753] The server segments video information based on summary information and adds timestamps. The input is summary information, and the output is segmented video information and timestamps. Each segment corresponds to the summarized content and is timestamped to allow the user to easily find the parts of interest.
[0754] Step 6:
[0755] The terminal receives segment information and timestamps from the server and displays them on the user interface. The input is segment information from the server, and the output is a user-operable GUI display. The terminal provides a visual interface to the user, making it easier to select videos for each segment.
[0756] Step 7:
[0757] The user selects a segment of interest from the interface on their device and plays the corresponding video information. The input is the user's selection, and the output is the video playback of the selected segment. The user can efficiently view different segments and obtain the necessary information in a short amount of time.
[0758] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0759] This invention combines a system designed for efficient summarization and viewing support of video content with an emotion engine that recognizes user emotions. This system is particularly effective in providing a personalized video experience tailored to the user's interests and emotions.
[0760] The system operates through the collaboration of servers, terminals, and users. The process is described below.
[0761] Integration of video analysis and emotion recognition
[0762] The server receives video data uploaded by users and converts the audio information into text. Then, it generates a summary of the video using natural language processing and extracts important topics. It also creates a reverse search index based on the generated summary and chapter information. Simultaneously, it utilizes an emotion engine to analyze the user's facial expressions and audio data and recognize emotions in real time.
[0763] Emotion-based content optimization
[0764] The server adjusts the priority of video segments based on the user's emotional data recognized by the emotion engine. This process displays video content that is appropriate for the user's interests and current emotional state. The system also personalizes chapter presentation within videos based on emotional data, making it easier for users to access content they prefer in their specific emotional state.
[0765] Providing a user interface
[0766] The device displays summary information, chapter information, and emotion-based content suggestions sent from the server in its user interface. Using this interface, users can select and watch video segments optimized for their emotions.
[0767] Specific example
[0768] For example, suppose educational videos are available. In this system, if a user's mood deteriorates while watching the video, the server detects this using an emotion engine and prioritizes displaying content or chapters that boost motivation. Conversely, if the user is excited, it can suggest more detailed and technical segments. The device provides a user-friendly interface, ensuring an optimal viewing experience tailored to the user's emotional state.
[0769] This invention enables efficient and highly satisfying information acquisition by providing videos that reflect the user's emotional state.
[0770] The following describes the processing flow.
[0771] Step 1:
[0772] The user uploads video data to their device. The device then prepares to send the file to the server.
[0773] Step 2:
[0774] The server analyzes the received video file and extracts the audio track. Using speech recognition technology, this audio is converted into text data, and the information within the video is transcribed into text.
[0775] Step 3:
[0776] The server performs natural language processing based on the text information to extract important keywords and topics. This information is then used to generate a summary of the video.
[0777] Step 4:
[0778] The server generates video chapters based on summary information. Each chapter is timestamped, allowing viewers to directly access segments of interest.
[0779] Step 5:
[0780] The device captures the user's facial expressions and voice through its camera and microphone, and sends this data to a server.
[0781] Step 6:
[0782] The server uses an emotion engine to recognize the user's emotional state in real time from received facial expression and voice data. It then analyzes the emotional data to determine the user's current emotions.
[0783] Step 7:
[0784] The server reconstructs chapters and summaries optimized for the user based on the emotions it perceives. It also adjusts the priority of the content presented according to the user's emotions.
[0785] Step 8:
[0786] The device displays optimized summaries and chapter information in the user interface. Based on this information, users can select and watch video segments that match their emotional responses.
[0787] This entire process allows users to have an effective and engaging video viewing experience with content that matches their emotional state.
[0788] (Example 2)
[0789] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0790] The increasing volume of digital video content makes it difficult for users to efficiently obtain information that matches their interests and circumstances. Furthermore, the lack of personalized video experiences tailored to individual user emotions and interests leads to stress and decreased satisfaction. Solving these challenges and enabling users to enjoy efficient and highly satisfying video viewing experiences is essential.
[0791] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0792] In this invention, the server includes means for acquiring input video data, means for extracting audio data from the video data and converting the audio data into text data, and means for extracting important information based on the text data and generating summary information. This enables the provision of customized content for each user and efficient video viewing tailored to their emotional state.
[0793] "Video data" refers to files containing visual information that are stored or transmitted in electronic format, and may include videos and animations.
[0794] "Audio data" refers to digital information used to electronically store or transmit audio, and includes the corresponding audio portion within video data.
[0795] "Text data" refers to information in text format converted from audio data, representing the content of the audio in a linguistic way.
[0796] "Summary information" refers to information that expresses the key points and main topics extracted from text data in a shortened form.
[0797] A "chapter" refers to a specific segment within video data that has been organized to make it easily accessible to users.
[0798] A "recognition engine" refers to a set of algorithms and software used to analyze a user's emotional state, identifying emotions based on the input data.
[0799] A "segment" refers to a section of video data that is divided into parts, and each segment may contain different content or topics.
[0800] "Timestamp" refers to timestamp information used to identify the time when a particular event or data point occurred.
[0801] A "user interface" refers to a screen or input device that allows a user to interact with a system, enabling operations and information provision.
[0802] This invention is a system that optimizes video data according to the user's emotional state to provide a personalized viewing experience. Its embodiments are described in detail below.
[0803] The server receives video data transmitted digitally from the user. HTTPS, a common data transfer protocol, is used for communication to receive the data. Audio data is separated from the received video data, and the audio data is converted into text data using a speech recognition API. Examples of APIs used here include the Google Cloud Speech-to-Text API.
[0804] The server uses natural language processing (NLP) techniques to extract important information based on text data acquired through speech recognition. This process utilizes NLP libraries such as spaCy and NLTK. It generates summary information from the important data and then uses that information to generate chapters for the video data.
[0805] In parallel, the server extracts frames from the user's video data and performs emotion recognition using an image processing library (e.g., OpenCV) and an audio emotion analysis tool (e.g., DeepFace). Based on the emotion recognition results, the order and display of the video data segments are adjusted. This process dynamically provides content optimized for the user's emotional state.
[0806] The device provides users with personalized content suggestions based on summary information, chapter information, and sentiment sent from the server. The user interface used here is built using JavaScript frameworks such as React and Vue.js and is designed to accurately reflect the user's choices.
[0807] For example, consider the case of providing educational videos. If a user begins to lose interest while watching, the server uses a recognition engine to detect the user's emotional change and prioritizes displaying chapters that restore their motivation. In this way, content that fits the user's situation and emotions can be smoothly delivered. An example of a prompt would be: "Suggest content to display when the emotion engine recognizes that the user is relaxed."
[0808] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0809] Step 1:
[0810] The server receives video data uploaded by the user from their device. In this process, data transfer occurs when the user selects a file and sends it to the server. The input data consists of video files selected by the user. The output is the video data stored on the server.
[0811] Step 2:
[0812] The server separates the audio data from the received video data. This is done by analyzing the media file using a video processing library. The input is the stored video data. The output is stored as audio data.
[0813] Step 3:
[0814] The server sends the extracted audio data to a speech recognition API and converts it into text data. Specifically, it calls the API to send audio and receives a text data response. The input is audio data, and the output is converted text data.
[0815] Step 4:
[0816] The server passes the text data to a natural language processing engine, which generates summary information and extracts important topics. This is achieved by using NLP techniques to extract key phrases and summarize the text. The input is text data, and the output is summary information and topic data.
[0817] Step 5:
[0818] The server analyzes facial expression data from the user to recognize emotions. This process uses frames from video data as input to an emotion recognition model. The input is the user's facial expression data, and the output is the recognized emotion information.
[0819] Step 6:
[0820] The server adjusts the chapter priority of video data based on sentiment information and selects content to provide the optimal viewing experience. It analyzes sentiment data and chapter information and changes the display priority. The input is sentiment information and chapter information, and the output is a list of optimized content.
[0821] Step 7:
[0822] The terminal displays summary information, chapter information, and content suggestions sent from the server on the user interface. The user selects a suggested chapter through the interface and starts playback. The input is information data from the server, and the output is the user's chapter selection and playback.
[0823] (Application Example 2)
[0824] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0825] Viewers are exposed to a vast amount of video content every day, making it difficult to select the right content. Furthermore, providing content that caters to the emotions and interests of individual viewers is challenging, resulting in a uniform viewing experience and a lack of user satisfaction.
[0826] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0827] In this invention, the server includes means for acquiring input video information, means for extracting audio data from the video information and converting it into text information, means for extracting important information based on the text information and generating summary information, means for analyzing the user's facial expression data and audio data to recognize emotions, and means for adjusting the display order of video chapters based on emotions. This makes it possible to provide a more personalized viewing experience that responds to the viewer's emotional state.
[0828] "Video information" refers to digital data that represents the content of a video, and includes audio and video.
[0829] "Audio data" refers to digital data related to sound extracted from video and image information, and typically includes information that is perceived by the human ear as sound.
[0830] "Text information" refers to information obtained by converting audio data into written characters, and is converted into a format that can be understood by humans as natural language.
[0831] "Summary information" is a condensed version of important information extracted from text, designed to help viewers quickly grasp the main content of a video.
[0832] A "chapter" refers to a specific segment within video information, indicating a particular part of the video generated based on summary information.
[0833] "Facial expression data" refers to digital data that indicates a user's emotional state, obtained from the user's facial movements and other factors.
[0834] "Means of recognizing emotions" refers to technologies that analyze a user's emotional state from their facial expressions and voice, and recognize it as digital information.
[0835] "Means of adjusting the display order" refers to the process of optimizing the order in which video chapters are presented to the user based on summary information and sentiment data.
[0836] To realize this application, the program builds a system that provides viewers with personalized video content tailored to their emotional state. The main hardware used is a server for processing video information, which communicates with terminals that collect user facial expressions and audio data.
[0837] The server acquires video and audio data, extracts audio data, and converts it into text. This uses the Google Cloud Speech-to-Text API as its speech recognition technology. Next, it uses the natural language processing library spaCy to generate a summary from the text, and then generates video segments based on that summary.
[0838] Simultaneously, the device uses its built-in camera and microphone to collect user facial expression data and voice, and analyzes this data in real time to recognize emotions. This utilizes image processing libraries such as OpenCV and specific algorithms useful for emotion analysis.
[0839] Based on user sentiment data, the server optimizes content and presents video segments in an order that suits the user. The user interface on the device displays video segments tailored to the user's sentiment, allowing the user to select content to watch according to their interests.
[0840] For example, if a user is feeling stressed, the server will prioritize displaying relaxing content. For instance, a prompt such as "When the user is deemed tired, prioritize displaying relaxing videos" will be generated, and content suggestions will be made to help the user relieve stress.
[0841] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0842] Step 1:
[0843] The server receives video information from the user and extracts audio data. At this stage, the input is the video file from the user, and the output is the audio data separated from the video. The process involves converting the video file to a digital media format and extracting the audio track.
[0844] Step 2:
[0845] The server uses speech recognition technology to convert audio data into text. The input here is the audio data obtained in step 1, and the output is the text information derived from it. The Google Cloud Speech-to-Text API is used to analyze the audio waveform and output the audio as text.
[0846] Step 3:
[0847] The server analyzes text information using natural language processing and generates summary information. In this step, text information is input, and information summarizing the main content of the video is output. The natural language processing library spaCy is used to extract important topics from the text and create a summary.
[0848] Step 4:
[0849] The server generates video chapters based on the summary information. The input is the summary information obtained in step 3, and the output is the chapter-structured video segments. The timeline is divided based on the summary information, and the information related to each segment is organized as chapters.
[0850] Step 5:
[0851] The device collects and analyzes the user's facial expression data and audio in real time. The input here is the user's camera video and audio, and the output is digital data representing the user's emotions. Camera video is captured using OpenCV, and emotion analysis is performed using a specific algorithm.
[0852] Step 6:
[0853] The server optimizes the display order of video segments based on sentiment data. The starting points are the chapter information generated in step 4 and the sentiment data obtained in step 5. The output is a list presenting chapters in an optimized viewing order for the user. The sentiment data is analyzed to prioritize videos based on the user's current state.
[0854] Step 7:
[0855] The device displays emotion-based video segment suggestions on the user interface and accepts user selections. The input is a server-prepared list of chapters, and the output is a list of video segment candidates displayed on the user interface. It visually displays a list of recommended content based on emotion and plays the video corresponding to the user's selection.
[0856] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0857] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0858] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the robot 414.
[0859] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[0860] Figure 9 shows an emotion map 400 in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.
[0861] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.
[0862] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.
[0863] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, motorcycles, etc., emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.
[0864] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."
[0865] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.
[0866] The above description primarily focuses on the functions of the data processing device 12 in relation to this disclosure. However, the system related to this disclosure is not necessarily implemented on a server. The system related to this disclosure may be implemented as a general information processing system. This disclosure may be implemented, for example, as a software program that runs on a personal computer or as an application that runs on a smartphone. The method related to this disclosure may be provided to users in SaaS (Software as a Service) format.
[0867] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing of the specific process may be performed by multiple computers, including computer 22. For example, a data generation model 58 may be provided in an external device of the data processing device 12, and the external device may generate data according to the input data.
[0868] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.
[0869] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[0870] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.
[0871] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.
[0872] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.
[0873] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.
[0874] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.
[0875] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and the like that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.
[0876] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted to be incorporated by reference.
[0877] The following is further disclosed regarding the embodiments described above.
[0878] (Claim 1)
[0879] A means for acquiring input video data,
[0880] A means for extracting audio information from the aforementioned video data and converting the audio information into text information,
[0881] A means for extracting important information and generating summary information based on the aforementioned text information,
[0882] A means for generating chapters in the video data based on the summary information,
[0883] A means for reverse-searching the video data based on the search terms entered by the user,
[0884] A system that includes this.
[0885] (Claim 2)
[0886] The system according to claim 1, wherein the chapter includes a timestamp for identifying the time corresponding to the important information.
[0887] (Claim 3)
[0888] The system according to claim 1, further comprising means for displaying the summary information and chapter information on a user interface and for playing back video data corresponding to the selected chapter when selected by the user.
[0889] "Example 1"
[0890] (Claim 1)
[0891] A means for acquiring input video information,
[0892] A means for extracting sound data from the aforementioned video information and converting the sound data into text information,
[0893] A means for identifying important information and generating summary data based on the aforementioned textual information,
[0894] A means for structuring the video information and generating chapters based on the summary data,
[0895] A means for reverse-searching the video information based on the search terms entered by the user,
[0896] The means for dividing the video into specific frames based on the identified topic during the chapter generation process,
[0897] A system that includes this.
[0898] (Claim 2)
[0899] The system according to claim 1, wherein the chapter includes a time marker indicating the time corresponding to the important information.
[0900] (Claim 3)
[0901] The system according to claim 1, further comprising means for presenting the summary data and chapter data on a user interface and for playing video information corresponding to the selected chapter when selected by the user.
[0902] "Application Example 1"
[0903] (Claim 1)
[0904] A means for acquiring input video information,
[0905] A means for extracting acoustic information from the aforementioned video information and converting the acoustic information into text information,
[0906] A means for extracting important information and generating summary information based on the aforementioned text information,
[0907] Means for generating segments in the video information based on the summary information,
[0908] A means for reverse-searching the video information based on search instructions entered by the user,
[0909] A means of providing users with summaries and segmented information based on specified topics, across multiple information platforms.
[0910] A system that includes this.
[0911] (Claim 2)
[0912] The system according to claim 1, wherein the segment includes a time marker for identifying the time corresponding to the important information.
[0913] (Claim 3)
[0914] The system according to claim 1, further comprising means for displaying the summary information and segment information on a user interface and for playing video information corresponding to the selected segment when selected by the user.
[0915] "Example 2 of combining an emotion engine"
[0916] (Claim 1)
[0917] A means for acquiring input video data,
[0918] A means for extracting audio data from the aforementioned video data and converting the audio data into text data,
[0919] A means for extracting important information and generating summary information based on the aforementioned text data,
[0920] Means for generating chapters in the video data based on the summary information,
[0921] A means of analyzing the user's emotions using a recognition engine,
[0922] Means for adjusting video data segments based on the aforementioned emotions,
[0923] A means of displaying customized content to the user,
[0924] A system that includes this.
[0925] (Claim 2)
[0926] The system according to claim 1, wherein the chapter includes a time stamp for identifying the time corresponding to the important information.
[0927] (Claim 3)
[0928] The system according to claim 1, further comprising means for displaying the summary information and chapter information on a user interface, and for playing video data corresponding to the selected chapter when selected by the user.
[0929] "Application example 2 when combining with an emotional engine"
[0930] (Claim 1)
[0931] A means for acquiring input video information,
[0932] A means for extracting audio data from the aforementioned video information and converting the audio data into text information,
[0933] A means for extracting important information and generating summary information based on the aforementioned text information,
[0934] Means for generating chapters in the video information based on the summary information,
[0935] A means for reverse-searching the video information based on the search terms entered by the user,
[0936] A means of analyzing user facial expression data and voice data to recognize emotions,
[0937] A means for adjusting the display order of the chapters in the video based on the aforementioned emotions,
[0938] A system that includes this.
[0939] (Claim 2)
[0940] The system according to claim 1, wherein the chapter includes a time marker for identifying the time corresponding to the important information, and is provided to the user interface in a state adjusted based on the emotion.
[0941] (Claim 3)
[0942] The system according to claim 1, further comprising means for displaying the summary information and chapter information on a user interface, suggesting chapters with adjusted priority according to the user's emotional state, and playing video information corresponding to the selected chapter when the user makes a selection. [Explanation of symbols]
[0943] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots< / url:> < / url:> < / url:> < / url:>
Claims
1. A means for acquiring input video data, A means for extracting audio information from the aforementioned video data and converting the audio information into text information, A means for extracting important information and generating summary information based on the aforementioned text information, A means for generating chapters in the video data based on the summary information, A means for reverse-searching the video data based on the search terms entered by the user, A system that includes this.
2. The system according to claim 1, wherein the chapter includes a timestamp for identifying the time corresponding to the important information.
3. The system according to claim 1, further comprising means for displaying the summary information and chapter information on a user interface, and for playing back video data corresponding to the selected chapter when selected by the user.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A