system
The system addresses the challenge of conveying cooking video content by preprocessing, identifying ingredients, and generating organized captions, improving viewer engagement through viewer feedback analysis.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- SOFTBANK GROUP CORP
- Filing Date
- 2024-10-16
- Publication Date
- 2026-04-28
AI Technical Summary
Cooking videos on video platforms face challenges in clearly conveying necessary materials and procedures, requiring significant effort from creators and lacking mechanisms to analyze viewer interests for content improvement.
A system that preprocesses video input, identifies food and seasonings through image analysis, converts audio to text, extracts cooking procedures, and generates visually organized captions, while collecting viewer reactions to suggest content enhancements.
Enhances viewer understanding and reduces creator effort by accurately conveying recipe information and optimizing content based on viewer feedback.
Smart Images

Figure 2026071023000001_ABST
Abstract
Description
Technical Field
[0001] The technology of the present disclosure relates to a system.
Background Art
[0002] Patent Document 1 discloses a method for controlling a persona chatbot, which is performed by at least one processor, and includes steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance that responds to the user utterance.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] In recent years, cooking videos on video platforms have become increasingly popular, but there is a problem that it is difficult for viewers to clearly grasp the necessary materials and procedures from the videos. Also, for creators, it takes time and effort to accurately convey detailed recipe information within the video, and it is difficult to do so within a limited time. Furthermore, there is still room for improvement in how to analyze the interests and concerns of viewers after watching the video and reflect them in the next content. Thus, a mechanism for improving the convenience of both viewers and creators is required.
Means for Solving the Problems
[0005] To solve this problem, the present invention provides a system that receives video input, performs preprocessing on it, and analyzes image frames for each material to identify food and seasonings. By using a technology that converts audio data to text and extracts cooking procedures, text information based on the extracted procedures and identified food can be generated, and visually organized captions can be created. Furthermore, by collecting viewer reactions, analyzing viewing trends, and providing suggestions for content improvement based on these, the system supports the creation of highly engaging content while reducing the burden on creators. In addition, it provides users with the generated captions and enables editing, providing a mechanism that allows users to easily optimize content.
[0006] "Video input" refers to video data received by the system through the platform.
[0007] "Preprocessing" refers to the process of performing data transformations and adjustments necessary for video analysis.
[0008] An "image frame" refers to each individual still image that makes up a video.
[0009] "Food" is a general term for the ingredients used in cooking videos.
[0010] "Seasonings" are substances used to adjust the taste of food.
[0011] "Audio data" refers to audio information recorded within a video.
[0012] "Text conversion" is the process of generating text data from audio data.
[0013] "Cooking instructions" are the instructions for each stage in completing a dish.
[0014] "Text information" refers to information in the form of characters extracted or generated from a video.
[0015] "Caption" refers to the subtitles and explanatory texts displayed within a video.
[0016] "Viewer reaction" refers to the actions and feedback that viewers perform on a video.
[0017] "Viewing tendency" refers to the pattern of data regarding the content that viewers are interested in.
[0018] "Content improvement plan" refers to the recommended proposals for improving the quality of a video.
[0019] "User" refers to a person who creates or edits cooking videos using the system.
Brief Explanation of Drawings
[0020] [Figure 1] It is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] It is a conceptual diagram showing an example of the main functions of a data processing device and a smart device according to the first embodiment. [Figure 3] It is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] It is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] It is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] It is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] It is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] It is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] It shows an emotion map to which multiple emotions are mapped. [Figure 10] Shows an emotion map to which a plurality of emotions are mapped. [Figure 11] It is a sequence diagram showing the processing flow of the data processing system in Embodiment 1. [Figure 12] It is a sequence diagram showing the processing flow of the data processing system in Application Example 1. [Figure 13] It is a sequence diagram showing the processing flow of the data processing system in Embodiment 2 when an emotion engine is combined. [Figure 14] It is a sequence diagram showing the processing flow of the data processing system in Application Example 2 when an emotion engine is combined.
Mode for Carrying Out the Invention
[0021] Hereinafter, an example of an embodiment of a system according to the technology of the present disclosure will be described with reference to the accompanying drawings.
[0022] First, the terms used in the following description will be described.
[0023] In the following embodiments, a numbered processor (hereinafter simply referred to as "processor") may be one arithmetic unit or a combination of a plurality of arithmetic units. Also, the processor may be one type of arithmetic unit or a combination of a plurality of types of arithmetic units. Examples of arithmetic units include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), an APU (Accelerated Processing Unit), and the like.
[0024] In the following embodiments, a numbered RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a work memory by the processor.
[0025] In the following embodiments, the signed storage is one or more non-volatile storage devices that store various programs and various parameters. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes.
[0026] In the following embodiments, the signed communication interface (I / F) is an interface that includes a communication processor and an antenna, etc. The communication interface manages communication between multiple computers. Examples of communication standards applicable to the communication interface include wireless communication standards such as 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark).
[0027] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."
[0028] [First Embodiment]
[0029] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.
[0030] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0031] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0032] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.
[0033] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0034] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0035] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.
[0036] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0037] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.
[0038] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0039] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0040] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".
[0041] In embodiments of the present invention, a process is described in detail for automatically extracting ingredients and cooking procedures from cooking videos using a video analysis system, thereby generating useful information for viewers and creators. This system operates in cooperation with a server, a terminal, and a user.
[0042] The server first receives video input uploaded from the terminal and performs preprocessing to analyze the video content. This preprocessing involves dividing the video into frames and converting the format as needed. Next, the server uses image recognition technology to identify information about food and seasonings from each frame and stores this information in a database. Audio data from the video is also extracted and converted into text using speech recognition technology. The server then uses natural language processing to extract cooking instructions from the transcribed data, organizes the instructions, and stores them in the database in chronological order.
[0043] The server then generates text information based on the food and procedure information stored in the database. This information is visually organized and used to create captions for use in the video being played. The generated captions and subtitles are provided to the user and can be edited by the user. This process enhances the accuracy of the information and helps creators effectively convey recipe information to viewers.
[0044] Furthermore, the server collects viewer reactions to published videos and analyzes viewing trends. Based on this analysis, it suggests improvements to future video content and notifies users, helping them create more engaging videos.
[0045] As a concrete example, when a user uploads a specific cooking video, the server recognizes the ingredients used, such as tomatoes and olive oil, and extracts how they were used based on the recipe. Next, the user can review the generated caption on their screen and edit details such as quantities and cooking time, making it easier for viewers to quickly replicate the recipe. This entire process provides convenient and useful information to both users and viewers, further enhancing the value of cooking videos.
[0046] The following describes the processing flow.
[0047] Step 1:
[0048] The user films a cooking video and opens the video selection screen to upload it to the platform. They click the upload button to send the video from their device to the server.
[0049] Step 2:
[0050] The server preprocesses the received video by checking and adjusting its format. Here, the video is broken down frame by frame, and the resolution and frame rate are adjusted to create a format suitable for analysis.
[0051] Step 3:
[0052] The server applies an image recognition algorithm to identify visible food items and condiments from each frame. It extracts the names and characteristics of the recognized objects and stores them in a database.
[0053] Step 4:
[0054] The server separates the audio from the video and converts it to text using speech recognition technology. The resulting text data is then subjected to natural language processing to analyze the cooking procedures and instructions, and the necessary information is extracted.
[0055] Step 5:
[0056] The server generates detailed recipe text based on ingredient and procedure information stored in the database, and creates visual captions and subtitle formats.
[0057] Step 6:
[0058] The device displays the generated captions and subtitles on the user's platform dashboard. The user can review and edit them, correcting the text information as needed.
[0059] Step 7:
[0060] The server collects viewer reaction data to published videos. It analyzes the viewing data and generates analysis results based on viewers' interests and preferences.
[0061] Step 8:
[0062] Based on the analysis of viewing data, the server sends users suggestions for improvement and ways to enhance engagement, which will be useful for future content creation.
[0063] (Example 1)
[0064] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0065] In modern dietary habits, visually presenting cooking procedures and ingredient information is crucial for viewers' understanding. However, manually extracting and organizing this information is time-consuming and laborious. Furthermore, there is no system in place to collect viewer feedback and use it to improve future content. This creates a challenge for creators, making it difficult to effectively communicate information and increase viewer engagement.
[0066] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0067] In this invention, the server includes a device that receives video data and performs preprocessing to facilitate analysis, a device that analyzes still images to identify ingredients and seasonings, and a device that converts sound data into text data and extracts cooking instructions. This allows viewers to intuitively understand the cooking instructions through subtitles, and enables improvements to future content based on viewer feedback.
[0068] "Video data" refers to digital or analog information media used to record and reproduce dynamic information that includes visual and auditory elements.
[0069] "Preprocessing" refers to the preliminary processing steps performed to facilitate data analysis, and includes data format conversion and quality improvement.
[0070] "Device" refers to a machine or system designed to perform a specific function.
[0071] A "still image" is a digital or analog representation of visual information that captures a specific moment in time.
[0072] "Ingredients" refers to materials classified according to their original cultural context and use in cooking.
[0073] "Seasoning" refers to ingredients or mixtures used to add flavor or taste to food.
[0074] "Audio data" refers to audio information recorded in digital or analog format and used for analysis and playback.
[0075] "Text data" refers to linguistic information represented in digital or analog format, and is in a readable and writable form.
[0076] "Cooking instructions" refer to a series of instructions or processes that should be followed when preparing a dish.
[0077] Subtitles are textual information displayed within visual media, and are typically used to provide transcripts or supplementary information for audio content.
[0078] This invention relates to a system that uses video analysis technology to automatically extract ingredients and cooking procedures from cooking videos and present them to viewers in an easy-to-understand format. This system operates through the coordinated efforts of a server, terminals, and users.
[0079] The server first receives video data sent from the terminal, recognizes the video format before analysis, and converts the format if necessary. Here, the server uses OpenCV to split the video into still images. These split still images are used as base data for identifying ingredients and seasonings. Machine learning libraries such as TENSORFLOW® and PyTorch are used for this image analysis, and the information is stored in a database.
[0080] Furthermore, the server extracts the audio data from the video and converts it into text data using the Google Cloud Speech-to-Text API. This text data is then analyzed and extracted using BERT or similar natural language processing models to describe the specific cooking steps. These steps are then organized into a sequence that is easy for viewers to understand.
[0081] The generated subtitle information and instructions are sent to the user's device, where the user can review and edit the captions. This editing function allows the user to supplement ingredient quantities and substitutions, ensuring viewers can accurately recreate the dish.
[0082] As a concrete example of its use, if a user uploads a cooking video of "lasagna" from their smartphone, the server analyzes it and identifies the main ingredients such as meat, tomato sauce, and cheese. Based on this information, it generates a caption, which the user can review and edit as needed to provide viewers with more accurate information, such as ingredient details and cooking time.
[0083] An example of a prompt message used in implementing this system is: "Analyze the cooking video, extract the ingredients used and cooking steps, and generate captions."
[0084] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0085] Step 1:
[0086] The terminal uploads cooking videos selected by the user to the server. The input is a cooking video. This video is sent to the server via data communication. The server temporarily stores the received video in its storage and checks the video format. The output is the saved video file.
[0087] Step 2:
[0088] The server begins preprocessing to make the video easier to analyze. The input is a saved video file. The server uses OpenCV to split the video into multiple still images. These split images are temporarily stored in a database and prepared for identification of ingredients and seasonings. The output is a set of individual still images.
[0089] Step 3:
[0090] The server analyzes each still image to identify ingredients and seasonings. The input is the set of still images obtained in step 2. Machine learning frameworks such as TensorFlow and PyTorch are used for this analysis. The server records the identified ingredient information in a database. The output is a dataset of the identified ingredient information.
[0091] Step 4:
[0092] The server extracts audio data from the video. The input is the original video file. This audio data is converted into text data using the Google Cloud Speech-to-Text API. The server then prepares the resulting text data for further analysis. The output is the text data obtained from the video's audio track.
[0093] Step 5:
[0094] The server analyzes the text data using a natural language processing model to extract cooking instructions. The input is the text data obtained in step 4. Using natural language processing techniques such as BERT, the cooking instructions are clearly analyzed, organized chronologically, and stored in a database. The output is a dataset of the organized cooking instructions.
[0095] Step 6:
[0096] The server generates captions for subtitles based on ingredient information and cooking procedure data. The inputs are the ingredient information from step 3 and the cooking procedure data from step 5. This caption information is visually organized and sent to the user's terminal. The output is the generated subtitle captions.
[0097] Step 7:
[0098] The terminal provides the user with the generated caption and displays an editable interface. The input is the caption sent from the server. The user uses this editing function to adjust ingredient quantities and supplementary information, and saves it as final information. The output is the final caption edited by the user.
[0099] (Application Example 1)
[0100] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0101] In video content, it is difficult to clearly convey cooking procedures and ingredients used to viewers. Furthermore, it is not easy for content creators to analyze viewer feedback and use it to improve future content. In this situation, there is a need for a system that effectively organizes information and provides useful information for both viewers and creators.
[0102] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0103] In this invention, the server includes means for receiving and processing video data, means for analyzing image data to identify ingredients and seasonings, and means for converting audio information into text information and extracting cooking procedures. This makes it possible to automatically extract ingredients and procedures from videos, organize them visually, and provide useful information to viewers and creators.
[0104] "Video data" refers to digital information that includes a sequence of images and sound in time.
[0105] "Data processing" is the process of performing calculations and operations to analyze, modify, or store received information.
[0106] "Image data" refers to digital data that contains still or moving visual information.
[0107] "Identifying ingredients and seasonings" is the process of identifying and classifying the substances used in cooking.
[0108] "Audio information" refers to data that records human speech and sounds in digital format.
[0109] "Text information" refers to data expressed in the form of characters or sentences.
[0110] "Extracting cooking steps" is the process of identifying and extracting the series of steps required to create a dish.
[0111] "Visual organization" means displaying information in a way that is easy to see and understand.
[0112] "Feedback" refers to information collected from viewers' reactions and opinions.
[0113] "Analyzing viewing trends" is the process of statistically investigating and analyzing viewers' behavior and preferences.
[0114] "Improvement suggestions" refer to providing specific ideas for making the existing situation better.
[0115] The system used to implement this application primarily relies on a server. The server first receives video data uploaded by users and processes it. This data processing often utilizes tools such as OpenCV or TensorFlow. The received video data is analyzed frame by frame as individual image data, identifying ingredients and seasonings. This process makes it possible to identify the elements used in a dish.
[0116] Next, the audio information contained in the video is converted into text using speech recognition technology such as the Google Cloud Speech-to-Text API. In this step, detailed cooking instructions can be extracted from the audio in cooking videos. The obtained information is then organized using Python's natural language processing libraries such as NLTK and spaCy. This makes it possible to generate subtitles in a visually easy-to-understand format.
[0117] Furthermore, the server uses Python's pandas and matplotlib to collect and analyze feedback from viewers. By analyzing viewing trends, such as which content viewers preferred, it is possible to make specific suggestions for improving future content.
[0118] For example, if a user uploads a video demonstrating a simple home cooking recipe, the server automatically extracts ingredients such as tomatoes and pasta, as well as cooking steps, and generates visually integrated captions. This allows viewers to immediately apply the recipe to their own cooking.
[0119] An example of a prompt for a generative AI model is, "Please tell me how to identify the ingredients used in the video, generate cooking instructions based on the recipe, and display them clearly for viewers." This helps the generative AI model provide specific steps and results.
[0120] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0121] Step 1:
[0122] The server receives video data from the user's terminal. This process begins with the user uploading a cooking video to the platform. The received data is a digital file containing chronologically sequential images and audio. This input video data is then subjected to the next analysis step.
[0123] Step 2:
[0124] The server uses the OpenCV library to split the received video into individual frames. TensorFlow is then used to identify ingredients and seasonings from this image data. The input is the video frames, and the output is a list of identified ingredient and seasoning names. In this step, the identified items are stored in a database for use in subsequent processes.
[0125] Step 3:
[0126] The server extracts the audio data from the video and converts it to text using the Google Cloud Speech-to-Text API. This process takes audio information as input and outputs it as transcribed text. The resulting text includes specific cooking instructions such as "chop the onions" and "add olive oil."
[0127] Step 4:
[0128] The extracted text information is processed using natural language processing with NLTK or spaCy to clarify the cooking instructions as commands. The input is transcribed audio data, and the output is an organized list of cooking instructions. This organized instruction information can be used to generate subtitles.
[0129] Step 5:
[0130] The server combines the obtained ingredient list and cooking instructions to generate visually easy-to-read subtitles and integrates them into the video. The input to this process is the identified ingredients and written procedure list, and the output is the captions displayed during video playback.
[0131] Step 6:
[0132] Viewer feedback is collected, and viewing trends are analyzed using Python's pandas and matplotlib. This data is input in the form of feedback, and the output is an analysis of which segments viewers preferred. Based on the analysis results, the server notifies the user with suggestions for improving the next content.
[0133] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0134] This invention incorporates an emotion engine into a video analysis system to individually optimize the production and viewing of cooking videos. This system operates through the cooperation of a server, terminal, and user, and provides a wide range of functions from video input to caption generation and viewer emotion recognition.
[0135] The server receives cooking videos uploaded by users and performs preprocessing for analysis. The video is divided frame by frame, converted to an appropriate format, and then image recognition technology is used to identify ingredients and seasonings from each frame, which are then stored in a database. During this process, audio data is converted to text, and cooking instructions are extracted using natural language processing technology.
[0136] Next, the server generates text information based on the identified ingredients and extracted procedures. The generated information is visually organized and used as captions and subtitles to be displayed on the video. Additionally, an emotion engine is incorporated, which analyzes the user's emotions from the audio and facial expressions in the video to suggest editing points and optimize the video.
[0137] The generated captions and sentiment analysis results are provided to the user via their device, allowing them to review and edit them to make adjustments for more effective video content delivery. For example, if the user makes expressions of tension or enjoyment in the video, the sentiment engine will recognize this and recommend editing to emphasize it.
[0138] Furthermore, the server continuously monitors viewer reactions after publication. It collects emotional responses viewers show to the video and uses this data to generate suggestions that individually optimize the viewing experience. For example, if many viewers smile at a particular part of the video, it becomes possible to create content that reflects viewer preferences, such as planning the next video using that part.
[0139] Systems that utilize emotion engines in this way provide video creators with the data and tools necessary to enhance engagement with viewers, making it easier to form deeper emotional connections. This can improve the quality of the viewing experience and increase viewer satisfaction.
[0140] The following describes the processing flow.
[0141] Step 1:
[0142] The user films a cooking video and operates the device to upload it to the platform. The device then transfers the video file to the server.
[0143] Step 2:
[0144] The server saves the received video to storage and performs preprocessing. This involves breaking down the video into frames and converting the format to make it easier to analyze.
[0145] Step 3:
[0146] The server uses image recognition technology to identify food items and condiments in each frame. It identifies the type and number of objects detected in each frame and stores this information in a database.
[0147] Step 4:
[0148] The server extracts audio from the video and converts it to text using a speech recognition engine. This allows the system to obtain the spoken words in the video as text information, and then analyze the content using natural language processing to extract the cooking instructions.
[0149] Step 5:
[0150] The server generates detailed recipe text based on the identified food items and extracted steps. This text information is visually organized and formatted as captions or subtitles for display in the video.
[0151] Step 6:
[0152] The emotion engine analyzes the user's facial expressions and voice tone in the video to understand their emotional state. This allows it to generate data that suggests areas for emphasis and improvement in editing.
[0153] Step 7:
[0154] The device displays generated captions and sentiment analysis results sent from the server on the user's dashboard. The user can then overwrite this information, edit the captions, and adjust the video based on sentiment.
[0155] Step 8:
[0156] After a video is published, the server collects viewer reactions into a database. It statistically analyzes the emotional responses viewers showed to specific parts of the video and identifies trends.
[0157] Step 9:
[0158] Based on the analysis results, the server notifies users with suggestions for improvements to the next video content and new content ideas in order to enhance viewer emotional engagement.
[0159] (Example 2)
[0160] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".
[0161] In modern society, video content is one of the primary means of information transmission, but its production and distribution require optimization to capture the viewer's interest. However, conventional systems lack editing suggestions based on viewer emotions and detailed analysis of viewing trends, making it difficult to create effective content.
[0162] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0163] In this invention, the server includes means for receiving video information and preparing it for analysis, means for analyzing image data to identify food substances and seasoning components, and means for converting audio signals into text data and obtaining cooking processes. This enables the analysis of viewers' emotions towards video content, and allows for specific editing suggestions and content optimization based on viewers' viewing behavior.
[0164] "Video information" refers to dynamic visual data transmitted through sight and hearing.
[0165] "Preparation for analysis" refers to the process of performing the necessary processing to analyze video information and converting the data into an appropriate format.
[0166] "Image data" refers to a collection of still images saved in a digital format.
[0167] "Food substances" refer to objects and ingredients that are consumed as part of a meal.
[0168] "Seasoning ingredients" refer to substances used to add taste and flavor to food.
[0169] "Audio signal" refers to the conversion of sound waves generated by the human voice into digital data.
[0170] "Text data" refers to a digital representation of character information.
[0171] "Cooking process" refers to a series of steps or processes necessary to complete a particular dish.
[0172] "Subtitles" refer to text information displayed on images or videos that serves to supplement or explain the content.
[0173] "Emotional state" refers to the emotional reactions of a person inferred from the audio and visuals within a video.
[0174] "Editing suggestions" refer to instructions or recommendations for editing that should be done to improve the quality and effectiveness of video content.
[0175] "Viewer reactions" refer to data about the emotions and behaviors exhibited by people who watched the video.
[0176] "Viewing behavior" refers to the methods and patterns in which viewers consume video content.
[0177] "Collecting data" refers to the process of gathering and storing information in a specific format.
[0178] "Media content" refers to the collection of information and expressions conveyed through the media.
[0179] An "improvement notice" refers to information that encourages specific corrections or additions aimed at improving the content of the media.
[0180] The embodiments for carrying out the invention are described below.
[0181] This system primarily consists of the cooperation of a server, terminals, and users. Specifically, the server receives video information and prepares it for analysis. The received video is first divided into frames and converted to image formats such as JPEG as needed. Video processing software is used for this process.
[0182] Next, the server uses image recognition libraries such as TensorFlow to analyze the image data. This allows it to identify food substances and seasoning components from each image frame. Meanwhile, the audio signal is converted into text data using speech recognition technology, and the cooking process is extracted from that text. This process is performed using NLP (Natural Language Processing) techniques.
[0183] Subsequently, the server generates captions and subtitles for the content based on the conversion and analysis processes described above. The generated subtitle information is managed and provided in a way that is easy for the user to understand. Furthermore, the server has a built-in emotion engine that can analyze the audio and visual information in the video to infer the user's emotional state. The results of this emotion analysis are provided to the user in the form of editing suggestions, etc.
[0184] For example, a user who has uploaded a cooking video can be recommended specific parts of the video as points that would be interesting to viewers. The emotion engine analyzes the user's facial expressions and tone of voice, and highlights scenes where the user is enjoying themselves and speaking, as points that should be emphasized more.
[0185] An example of a prompt might be: "Based on the analysis results of the cooking video, please suggest editing points to capture the viewer's interest. In particular, focus on scenes and expressions where the user is enjoying themselves, and edit in a way that evokes positive feelings in the viewer."
[0186] The device receives generated captions and editing suggestions and provides them to the user. Based on the provided information, the user can edit the video content and create a more engaging medium for viewers. This can provide a better viewing experience and improve viewer satisfaction.
[0187] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0188] Step 1:
[0189] The server receives video information from the user as input and converts the video into a format that can be analyzed. Specifically, it divides the video into frames and converts them to JPEG format. As a result of this conversion, still images of each frame are output.
[0190] Step 2:
[0191] The server processes the output frame images using image recognition software such as TensorFlow. Each frame image is used as input to identify food substances and seasoning components. Information on the recognized substances and components is stored in a database.
[0192] Step 3:
[0193] The server takes the audio signal from the video as input and converts it into text data using speech recognition technology. Based on the converted text data, it extracts the cooking steps using NLP technology. As a result of this extraction process, the text of the cooking procedure is output.
[0194] Step 4:
[0195] The server generates content captions and subtitles based on identified food substances and extracted cooking procedure information. This generated text information is then visually organized and output in a user-friendly format.
[0196] Step 5:
[0197] The server uses an emotion engine to analyze the user's emotional state, taking audio and visual information from the video as input. This analysis generates specific editing suggestions, which are then presented to the user. This analysis helps identify points that viewers might find interesting.
[0198] Step 6:
[0199] The device receives generated captions and editing suggestions as input and provides them to the user. The user then uses this information to edit the video and output it as more engaging content. Finally, the edited video content is ready to be published.
[0200] (Application Example 2)
[0201] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as a "server" and the smart device 14 as a "terminal".
[0202] When providing cooking videos to viewers, there is a challenge in delivering an optimal viewing experience that responds to viewers' emotions. Furthermore, there is a lack of concrete means for content creators to improve video content based on viewer reactions.
[0203] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0204] In this invention, the server includes means for receiving and pre-processing video input, means for analyzing image frames to identify ingredients and seasonings, means for converting audio data into text and extracting cooking procedures, and means for analyzing the viewer's emotions and suggesting video editing based on those emotions. This enables the provision of video content that takes the viewer's emotions into consideration and the individual optimization of the viewing experience.
[0205] The method of "receiving video input" refers to the process of acquiring video data provided by the user in a manipulateable format.
[0206] "Preprocessing" refers to the process of dividing video data into frames and converting them into the required format in order to make it analyzable.
[0207] "Analyzing image frames" is a technique that involves recognizing or identifying specific elements in each still image frame that makes up a video.
[0208] The method of "identifying ingredients and seasonings" involves detecting specific foods or seasonings from image frames within a video and classifying or naming them.
[0209] The method of "converting audio data to text" involves analyzing the audio content within a video and converting it into corresponding text.
[0210] The method of "extracting cooking procedures" is the process of clearly extracting the steps and methods of cooking from textual data and organizing that information.
[0211] "Generating text information" refers to a method of generating meaningful information, either visually or in written form, based on identified data.
[0212] "Creating captions" is the technique of constructing text or subtitles to visually present generated text information.
[0213] The method of "analyzing viewers' emotions" is a technique that estimates an emotional state by analyzing the nonverbal reactions (such as facial expressions and actions) that viewers exhibit.
[0214] The method of "providing video editing suggestions" is a technique that, based on analysis results, offers specific ideas for modifying or emphasizing parts of a video.
[0215] "Personalizing the viewing experience" is a method of adjusting video content to suit the individual preferences and needs of each viewer, taking into account their emotions and reactions.
[0216] This system is implemented by a series of programs running on a server. The server receives video from the user and first preprocesses it by dividing it into image frames. At this stage, it uses image processing libraries such as OpenCV to convert it into an analyzable representation. Next, it uses image recognition technology to identify ingredients and seasonings from each frame and stores them in a database. Audio data is converted into text using NLP technology and extracted as cooking instructions.
[0217] To analyze viewers' emotions, the server uses a model that implements emotion recognition technology. This model uses tools such as dlib to detect viewers' faces and estimate their emotions from their facial expressions. This enables the suggestion of video editing based on emotions.
[0218] The generated text information is visually organized by the server, and captions are created based on it. These captions and sentiment analysis results are provided to the user via the device. The user can then edit the video based on this information and make adjustments to enhance the viewing experience.
[0219] As a concrete example, the server sends a prompt message to an AI model in response to a cooking video filmed by a user: "What part of this cooking video made people smile the most? Identify that part and suggest editing to highlight it." This allows for the creation of videos that increase engagement and enables content improvement based on viewer reactions.
[0220] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0221] Step 1:
[0222] The server receives cooking videos from users. The input is a video file uploaded by the user. The server divides this video into frames and converts it into a format suitable for analysis. During this process, image processing libraries such as OpenCV are used to output image data for each frame.
[0223] Step 2:
[0224] The server analyzes the divided image frames and identifies ingredients and seasonings from each frame. The input is the image frames obtained in step 1. Using an image recognition algorithm, specific ingredients and seasonings are detected, and this information is recorded in the database. The output is a list of identified ingredients and seasonings.
[0225] Step 3:
[0226] The server processes audio data, converts it to text, and extracts cooking instructions. The input is an audio stream from a video. Natural language processing techniques are used to convert the audio to text, from which specific cooking steps are extracted. The output is text data showing the cooking instructions.
[0227] Step 4:
[0228] The server generates text information based on the identified ingredients and extracted procedures, and then visually organizes it. The inputs are the ingredient list from step 2 and the cooking procedure from step 3. Caption information is created based on these and linked to the user interface. The output is a visually organized caption.
[0229] Step 5:
[0230] The server analyzes the viewer's emotions and proposes emotion-based edits. The input is the currently playing video frame, including the viewer's facial expressions. Using an emotion recognition model such as dlib, it estimates emotions from facial expressions and generates editing suggestions based on the results. The output is a list of proposed editing points.
[0231] Step 6:
[0232] The terminal receives captions and editing suggestions generated from the server and presents them to the user. The input is data sent from the server. The user reviews these and uses them as an interface to edit and adjust the video. The output is the video content customized by the user.
[0233] Step 7:
[0234] The user continuously monitors viewer reactions and uses that data to individually optimize the viewing experience. The input is viewer reaction data. Viewer data is analyzed and used to create new content that improves the viewing experience. The output is a project proposal for the next project that reflects viewer reactions.
[0235] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0236] Data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include those described above. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions shown by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0237] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.
[0238] [Second Embodiment]
[0239] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.
[0240] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0241] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0242] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.
[0243] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0244] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0245] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0246] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0247] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0248] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0249] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0250] Next, the identification processing performed by the identification processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0251] In embodiments of the present invention, a process is described in detail for automatically extracting ingredients and cooking procedures from cooking videos using a video analysis system, thereby generating useful information for viewers and creators. This system operates in cooperation with a server, a terminal, and a user.
[0252] The server first receives video input uploaded from the terminal and performs preprocessing to analyze the video content. This preprocessing involves dividing the video into frames and converting the format as needed. Next, the server uses image recognition technology to identify information about food and seasonings from each frame and stores this information in a database. Audio data from the video is also extracted and converted into text using speech recognition technology. The server then uses natural language processing to extract cooking instructions from the transcribed data, organizes the instructions, and stores them in the database in chronological order.
[0253] The server then generates text information based on the food and procedure information stored in the database. This information is visually organized and used to create captions for use in the video being played. The generated captions and subtitles are provided to the user and can be edited by the user. This process enhances the accuracy of the information and helps creators effectively convey recipe information to viewers.
[0254] Furthermore, the server collects viewer reactions to published videos and analyzes viewing trends. Based on this analysis, it suggests improvements to future video content and notifies users, helping them create more engaging videos.
[0255] As a concrete example, when a user uploads a specific cooking video, the server recognizes the ingredients used, such as tomatoes and olive oil, and extracts how they were used based on the recipe. Next, the user can review the generated caption on their screen and edit details such as quantities and cooking time, making it easier for viewers to quickly replicate the recipe. This entire process provides convenient and useful information to both users and viewers, further enhancing the value of cooking videos.
[0256] The following describes the processing flow.
[0257] Step 1:
[0258] The user films a cooking video and opens the video selection screen to upload it to the platform. They click the upload button to send the video from their device to the server.
[0259] Step 2:
[0260] The server preprocesses the received video by checking and adjusting its format. Here, the video is broken down frame by frame, and the resolution and frame rate are adjusted to create a format suitable for analysis.
[0261] Step 3:
[0262] The server applies an image recognition algorithm to identify visible food items and condiments from each frame. It extracts the names and characteristics of the recognized objects and stores them in a database.
[0263] Step 4:
[0264] The server separates the audio from the video and converts it to text using speech recognition technology. The resulting text data is then subjected to natural language processing to analyze the cooking procedures and instructions, and the necessary information is extracted.
[0265] Step 5:
[0266] The server generates detailed recipe text based on ingredient and procedure information stored in the database, and creates visual captions and subtitle formats.
[0267] Step 6:
[0268] The device displays the generated captions and subtitles on the user's platform dashboard. The user can review and edit them, correcting the text information as needed.
[0269] Step 7:
[0270] The server collects viewer reaction data to published videos. It analyzes the viewing data and generates analysis results based on viewers' interests and preferences.
[0271] Step 8:
[0272] Based on the analysis of viewing data, the server sends users suggestions for improvement and ways to enhance engagement, which will be useful for future content creation.
[0273] (Example 1)
[0274] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0275] In modern dietary habits, visually presenting cooking procedures and ingredient information is crucial for viewers' understanding. However, manually extracting and organizing this information is time-consuming and laborious. Furthermore, there is no system in place to collect viewer feedback and use it to improve future content. This creates a challenge for creators, making it difficult to effectively communicate information and increase viewer engagement.
[0276] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0277] In this invention, the server includes a device that receives video data and performs preprocessing to facilitate analysis, a device that analyzes still images to identify ingredients and seasonings, and a device that converts sound data into text data and extracts cooking instructions. This allows viewers to intuitively understand the cooking instructions through subtitles, and enables improvements to future content based on viewer feedback.
[0278] "Video data" refers to digital or analog information media used to record and reproduce dynamic information that includes visual and auditory elements.
[0279] "Preprocessing" refers to the preliminary processing procedures carried out to facilitate the analysis of data, including data format conversion and quality improvement.
[0280] "Device" refers to a machine or system designed to perform a specific function.
[0281] "Still image" refers to the digital or analog representation of visual information that records a specific moment.
[0282] "Ingredient" refers to the materials classified in their original cultural and usage contexts when cooking.
[0283] "Seasoning" refers to the materials or mixtures used to add flavor and taste to dishes.
[0284] "Audio data" refers to the recording of audio information in digital or analog form and is used for analysis and playback.
[0285] "Character data" refers to the representation of language information in digital or analog form and is in a form that can be read and written.
[0286] "Cooking procedure" refers to a series of instructions or processes to be followed when cooking.
[0287] "Subtitle" refers to the character information displayed within visual media and is usually used for the purpose of providing a transcription of the audio content or supplementary information.
[0288] The present invention relates to a system that uses video analysis technology to automatically extract ingredients and cooking procedures from cooking videos and provide them in an easily understandable form to viewers. This system operates with the cooperation of each entity of the server, terminal, and user.
[0289] The server first receives video data sent from the terminal, recognizes the video format before analysis, and converts the format if necessary. Here, the server uses OpenCV to split the video into still images. These split still images are used as base data for identifying ingredients and seasonings. Machine learning libraries such as TensorFlow and PyTorch are used for this image analysis, and the information is stored in a database.
[0290] Furthermore, the server extracts the audio data from the video and converts it into text data using the Google Cloud Speech-to-Text API. This text data is then analyzed and extracted using BERT or similar natural language processing models to describe the specific cooking steps. These steps are then organized into a sequence that is easy for viewers to understand.
[0291] The generated subtitle information and instructions are sent to the user's device, where the user can review and edit the captions. This editing function allows the user to supplement ingredient quantities and substitutions, ensuring viewers can accurately recreate the dish.
[0292] As a concrete example of its use, if a user uploads a cooking video of "lasagna" from their smartphone, the server analyzes it and identifies the main ingredients such as meat, tomato sauce, and cheese. Based on this information, it generates a caption, which the user can review and edit as needed to provide viewers with more accurate information, such as ingredient details and cooking time.
[0293] An example of a prompt message used in implementing this system is: "Analyze the cooking video, extract the ingredients used and cooking steps, and generate captions."
[0294] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0295] Step 1:
[0296] The terminal uploads cooking videos selected by the user to the server. The input is a cooking video. This video is sent to the server via data communication. The server temporarily stores the received video in its storage and checks the video format. The output is the saved video file.
[0297] Step 2:
[0298] The server begins preprocessing to make the video easier to analyze. The input is a saved video file. The server uses OpenCV to split the video into multiple still images. These split images are temporarily stored in a database and prepared for identification of ingredients and seasonings. The output is a set of individual still images.
[0299] Step 3:
[0300] The server analyzes each still image to identify ingredients and seasonings. The input is the set of still images obtained in step 2. Machine learning frameworks such as TensorFlow and PyTorch are used for this analysis. The server records the identified ingredient information in a database. The output is a dataset of the identified ingredient information.
[0301] Step 4:
[0302] The server extracts audio data from the video. The input is the original video file. This audio data is converted into text data using the Google Cloud Speech-to-Text API. The server then prepares the resulting text data for further analysis. The output is the text data obtained from the video's audio track.
[0303] Step 5:
[0304] The server analyzes the character data with a natural language processing model and extracts cooking procedures. The input is the character data obtained in step 4. Using natural language processing technologies such as BERT, the cooking procedures are clearly analyzed, and the procedures are sorted by time series and stored in a database. The output is a dataset of sorted cooking procedures.
[0305] Step 6:
[0306] Based on the food ingredient information and the data of cooking procedures, the server generates captions for subtitles. The input is the food ingredient information in step 3 and the cooking procedure data in step 5. This caption information is visually sorted and transmitted to the user's terminal. The output is the generated subtitle captions.
[0307] Step 7:
[0308] The terminal provides the generated captions to the user and displays an editable interface. The input is the captions transmitted from the server. The user uses this editing function to adjust the quantity of ingredients and supplementary information and saves them as confirmed information. The output is the final captions edited by the user.
[0309] (Application Example 1)
[0310] Next, Application Example 1 will be described. In the following description, the data processing device 12 is referred to as the "server", and the smart glasses 214 are referred to as the "terminal".
[0311] <Q In video content, it is difficult to convey cooking procedures and used food ingredients to viewers in an easy-to-understand manner. Also, for content producers, it is not easy to analyze the reactions of viewers and improve the next content. In such a situation, there is a need for a system that effectively organizes information and provides information useful for both viewers and producers.
[0312] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0313] In this invention, the server includes means for receiving and processing video data, means for analyzing image data to identify ingredients and seasonings, and means for converting audio information into text information and extracting cooking procedures. This makes it possible to automatically extract ingredients and procedures from videos, organize them visually, and provide useful information to viewers and creators.
[0314] "Video data" refers to digital information that includes a sequence of images and sound in time.
[0315] "Data processing" is the process of performing calculations and operations to analyze, modify, or store received information.
[0316] "Image data" refers to digital data that contains still or moving visual information.
[0317] "Identifying ingredients and seasonings" is the process of identifying and classifying the substances used in cooking.
[0318] "Audio information" refers to data that records human speech and sounds in digital format.
[0319] "Text information" refers to data expressed in the form of characters or sentences.
[0320] "Extracting cooking steps" is the process of identifying and extracting the series of steps required to create a dish.
[0321] "Visual organization" means displaying information in a way that is easy to see and understand.
[0322] "Feedback" refers to information collected from viewers' reactions and opinions.
[0323] "Analyzing viewing trends" is the process of statistically investigating and analyzing viewers' behavior and preferences.
[0324] "Improvement suggestions" refer to providing specific ideas for making the existing situation better.
[0325] The system used to implement this application primarily relies on a server. The server first receives video data uploaded by users and processes it. This data processing often utilizes tools such as OpenCV or TensorFlow. The received video data is analyzed frame by frame as individual image data, identifying ingredients and seasonings. This process makes it possible to identify the elements used in a dish.
[0326] Next, the audio information contained in the video is converted into text using speech recognition technology such as the Google Cloud Speech-to-Text API. In this step, detailed cooking instructions can be extracted from the audio in cooking videos. The obtained information is then organized using Python's natural language processing libraries such as NLTK and spaCy. This makes it possible to generate subtitles in a visually easy-to-understand format.
[0327] Furthermore, the server uses Python's pandas and matplotlib to collect and analyze feedback from viewers. By analyzing viewing trends, such as which content viewers preferred, it is possible to make specific suggestions for improving future content.
[0328] For example, if a user uploads a video demonstrating a simple home cooking recipe, the server automatically extracts ingredients such as tomatoes and pasta, as well as cooking steps, and generates visually integrated captions. This allows viewers to immediately apply the recipe to their own cooking.
[0329] An example of a prompt for a generative AI model is, "Please tell me how to identify the ingredients used in the video, generate cooking instructions based on the recipe, and display them clearly for viewers." This helps the generative AI model provide specific steps and results.
[0330] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0331] Step 1:
[0332] The server receives video data from the user's terminal. This process begins with the user uploading a cooking video to the platform. The received data is a digital file containing chronologically sequential images and audio. This input video data is then subjected to the next analysis step.
[0333] Step 2:
[0334] The server uses the OpenCV library to split the received video into individual frames. TensorFlow is then used to identify ingredients and seasonings from this image data. The input is the video frames, and the output is a list of identified ingredient and seasoning names. In this step, the identified items are stored in a database for use in subsequent processes.
[0335] Step 3:
[0336] The server extracts the audio data from the video and converts it to text using the Google Cloud Speech-to-Text API. This process takes audio information as input and outputs it as transcribed text. The resulting text includes specific cooking instructions such as "chop the onions" and "add olive oil."
[0337] Step 4:
[0338] The extracted text information is processed using natural language processing with NLTK or spaCy to clarify the cooking instructions as commands. The input is transcribed audio data, and the output is an organized list of cooking instructions. This organized instruction information can be used to generate subtitles.
[0339] Step 5:
[0340] The server combines the obtained ingredient list and cooking instructions to generate visually easy-to-read subtitles and integrates them into the video. The input to this process is the identified ingredients and written procedure list, and the output is the captions displayed during video playback.
[0341] Step 6:
[0342] Viewer feedback is collected, and viewing trends are analyzed using Python's pandas and matplotlib. This data is input in the form of feedback, and the output is an analysis of which segments viewers preferred. Based on the analysis results, the server notifies the user with suggestions for improving the next content.
[0343] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0344] This invention incorporates an emotion engine into a video analysis system to individually optimize the production and viewing of cooking videos. This system operates through the cooperation of a server, terminal, and user, and provides a wide range of functions from video input to caption generation and viewer emotion recognition.
[0345] The server receives cooking videos uploaded by users and performs preprocessing for analysis. The video is divided frame by frame, converted to an appropriate format, and then image recognition technology is used to identify ingredients and seasonings from each frame, which are then stored in a database. During this process, audio data is converted to text, and cooking instructions are extracted using natural language processing technology.
[0346] Next, the server generates text information based on the identified ingredients and extracted procedures. The generated information is visually organized and used as captions and subtitles to be displayed on the video. Additionally, an emotion engine is incorporated, which analyzes the user's emotions from the audio and facial expressions in the video to suggest editing points and optimize the video.
[0347] The generated captions and sentiment analysis results are provided to the user via their device, allowing them to review and edit them to make adjustments for more effective video content delivery. For example, if the user makes expressions of tension or enjoyment in the video, the sentiment engine will recognize this and recommend editing to emphasize it.
[0348] Furthermore, the server continuously monitors viewer reactions after publication. It collects emotional responses viewers show to the video and uses this data to generate suggestions that individually optimize the viewing experience. For example, if many viewers smile at a particular part of the video, it becomes possible to create content that reflects viewer preferences, such as planning the next video using that part.
[0349] Systems that utilize emotion engines in this way provide video creators with the data and tools necessary to enhance engagement with viewers, making it easier to form deeper emotional connections. This can improve the quality of the viewing experience and increase viewer satisfaction.
[0350] The following describes the processing flow.
[0351] Step 1:
[0352] The user films a cooking video and operates the device to upload it to the platform. The device then transfers the video file to the server.
[0353] Step 2:
[0354] The server saves the received video to storage and performs preprocessing. This involves breaking down the video into frames and converting the format to make it easier to analyze.
[0355] Step 3:
[0356] The server uses image recognition technology to identify food items and condiments in each frame. It identifies the type and number of objects detected in each frame and stores this information in a database.
[0357] Step 4:
[0358] The server extracts audio from the video and converts it to text using a speech recognition engine. This allows the system to obtain the spoken words in the video as text information, and then analyze the content using natural language processing to extract the cooking instructions.
[0359] Step 5:
[0360] The server generates detailed recipe text based on the identified food items and extracted steps. This text information is visually organized and formatted as captions or subtitles for display in the video.
[0361] Step 6:
[0362] The emotion engine analyzes the user's facial expressions and voice tone in the video to understand their emotional state. This allows it to generate data that suggests areas for emphasis and improvement in editing.
[0363] Step 7:
[0364] The device displays generated captions and sentiment analysis results sent from the server on the user's dashboard. The user can then overwrite this information, edit the captions, and adjust the video based on sentiment.
[0365] Step 8:
[0366] After a video is published, the server collects viewer reactions into a database. It statistically analyzes the emotional responses viewers showed to specific parts of the video and identifies trends.
[0367] Step 9:
[0368] Based on the analysis results, the server notifies users with suggestions for improvements to the next video content and new content ideas in order to enhance viewer emotional engagement.
[0369] (Example 2)
[0370] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0371] In modern society, video content is one of the primary means of information transmission, but its production and distribution require optimization to capture the viewer's interest. However, conventional systems lack editing suggestions based on viewer emotions and detailed analysis of viewing trends, making it difficult to create effective content.
[0372] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0373] In this invention, the server includes means for receiving video information and preparing it for analysis, means for analyzing image data to identify food substances and seasoning components, and means for converting audio signals into text data and obtaining cooking processes. This enables the analysis of viewers' emotions towards video content, and allows for specific editing suggestions and content optimization based on viewers' viewing behavior.
[0374] "Video information" refers to dynamic visual data transmitted through sight and hearing.
[0375] "Preparation for analysis" refers to the process of performing the necessary processing to analyze video information and converting the data into an appropriate format.
[0376] "Image data" refers to a collection of still images saved in a digital format.
[0377] "Food substances" refer to objects and ingredients that are consumed as part of a meal.
[0378] "Seasoning ingredients" refer to substances used to add taste and flavor to food.
[0379] "Audio signal" refers to the conversion of sound waves generated by the human voice into digital data.
[0380] "Text data" refers to a digital representation of character information.
[0381] "Cooking process" refers to a series of steps or processes necessary to complete a particular dish.
[0382] "Subtitles" refer to text information displayed on images or videos that serves to supplement or explain the content.
[0383] "Emotional state" refers to the emotional reactions of a person inferred from the audio and visuals within a video.
[0384] "Editing suggestions" refer to instructions or recommendations for editing that should be done to improve the quality and effectiveness of video content.
[0385] "Viewer reactions" refer to data about the emotions and behaviors exhibited by people who watched the video.
[0386] "Viewing behavior" refers to the methods and patterns in which viewers consume video content.
[0387] "Collecting data" refers to the process of gathering and storing information in a specific format.
[0388] "Media content" refers to the collection of information and expressions conveyed through the media.
[0389] An "improvement notice" refers to information that encourages specific corrections or additions aimed at improving the content of the media.
[0390] The embodiments for carrying out the invention are described below.
[0391] This system primarily consists of the cooperation of a server, terminals, and users. Specifically, the server receives video information and prepares it for analysis. The received video is first divided into frames and converted to image formats such as JPEG as needed. Video processing software is used for this process.
[0392] Next, the server uses image recognition libraries such as TensorFlow to analyze the image data. This allows it to identify food substances and seasoning components from each image frame. Meanwhile, the audio signal is converted into text data using speech recognition technology, and the cooking process is extracted from that text. This process is performed using NLP (Natural Language Processing) techniques.
[0393] Subsequently, the server generates captions and subtitles for the content based on the conversion and analysis processes described above. The generated subtitle information is managed and provided in a way that is easy for the user to understand. Furthermore, the server has a built-in emotion engine that can analyze the audio and visual information in the video to infer the user's emotional state. The results of this emotion analysis are provided to the user in the form of editing suggestions, etc.
[0394] For example, a user who has uploaded a cooking video can be recommended specific parts of the video as points that would be interesting to viewers. The emotion engine analyzes the user's facial expressions and tone of voice, and highlights scenes where the user is enjoying themselves and speaking, as points that should be emphasized more.
[0395] An example of a prompt might be: "Based on the analysis results of the cooking video, please suggest editing points to capture the viewer's interest. In particular, focus on scenes and expressions where the user is enjoying themselves, and edit in a way that evokes positive feelings in the viewer."
[0396] The device receives generated captions and editing suggestions and provides them to the user. Based on the provided information, the user can edit the video content and create a more engaging medium for viewers. This can provide a better viewing experience and improve viewer satisfaction.
[0397] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0398] Step 1:
[0399] The server receives video information from the user as input and converts the video into a format that can be analyzed. Specifically, it divides the video into frames and converts them to JPEG format. As a result of this conversion, still images of each frame are output.
[0400] Step 2:
[0401] The server processes the output frame images using image recognition software such as TensorFlow. Each frame image is used as input to identify food substances and seasoning components. Information on the recognized substances and components is stored in a database.
[0402] Step 3:
[0403] The server takes the audio signal from the video as input and converts it into text data using speech recognition technology. Based on the converted text data, it extracts the cooking steps using NLP technology. As a result of this extraction process, the text of the cooking procedure is output.
[0404] Step 4:
[0405] The server generates content captions and subtitles based on identified food substances and extracted cooking procedure information. This generated text information is then visually organized and output in a user-friendly format.
[0406] Step 5:
[0407] The server uses an emotion engine to analyze the user's emotional state, taking audio and visual information from the video as input. This analysis generates specific editing suggestions, which are then presented to the user. This analysis helps identify points that viewers might find interesting.
[0408] Step 6:
[0409] The device receives generated captions and editing suggestions as input and provides them to the user. The user then uses this information to edit the video and output it as more engaging content. Finally, the edited video content is ready to be published.
[0410] (Application Example 2)
[0411] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0412] When providing cooking videos to viewers, there is a challenge in delivering an optimal viewing experience that responds to viewers' emotions. Furthermore, there is a lack of concrete means for content creators to improve video content based on viewer reactions.
[0413] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0414] In this invention, the server includes means for receiving and pre-processing video input, means for analyzing image frames to identify ingredients and seasonings, means for converting audio data into text and extracting cooking procedures, and means for analyzing the viewer's emotions and suggesting video editing based on those emotions. This enables the provision of video content that takes the viewer's emotions into consideration and the individual optimization of the viewing experience.
[0415] The method of "receiving video input" refers to the process of acquiring video data provided by the user in a manipulateable format.
[0416] "Preprocessing" refers to the process of dividing video data into frames and converting them into the required format in order to make it analyzable.
[0417] "Analyzing image frames" is a technique that involves recognizing or identifying specific elements in each still image frame that makes up a video.
[0418] The method of "identifying ingredients and seasonings" involves detecting specific foods or seasonings from image frames within a video and classifying or naming them.
[0419] The method of "converting audio data to text" involves analyzing the audio content within a video and converting it into corresponding text.
[0420] The method of "extracting cooking procedures" is the process of clearly extracting the steps and methods of cooking from textual data and organizing that information.
[0421] "Generating text information" refers to a method of generating meaningful information, either visually or in written form, based on identified data.
[0422] "Creating captions" is the technique of constructing text or subtitles to visually present generated text information.
[0423] The method of "analyzing viewers' emotions" is a technique that estimates an emotional state by analyzing the nonverbal reactions (such as facial expressions and actions) that viewers exhibit.
[0424] The method of "providing video editing suggestions" is a technique that, based on analysis results, offers specific ideas for modifying or emphasizing parts of a video.
[0425] "Personalizing the viewing experience" is a method of adjusting video content to suit the individual preferences and needs of each viewer, taking into account their emotions and reactions.
[0426] This system is implemented by a series of programs running on a server. The server receives video from the user and first preprocesses it by dividing it into image frames. At this stage, it uses image processing libraries such as OpenCV to convert it into an analyzable representation. Next, it uses image recognition technology to identify ingredients and seasonings from each frame and stores them in a database. Audio data is converted into text using NLP technology and extracted as cooking instructions.
[0427] To analyze viewers' emotions, the server uses a model that implements emotion recognition technology. This model uses tools such as dlib to detect viewers' faces and estimate their emotions from their facial expressions. This enables the suggestion of video editing based on emotions.
[0428] The generated text information is visually organized by the server, and captions are created based on it. These captions and sentiment analysis results are provided to the user via the device. The user can then edit the video based on this information and make adjustments to enhance the viewing experience.
[0429] As a concrete example, the server sends a prompt message to an AI model in response to a cooking video filmed by a user: "What part of this cooking video made people smile the most? Identify that part and suggest editing to highlight it." This allows for the creation of videos that increase engagement and enables content improvement based on viewer reactions.
[0430] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0431] Step 1:
[0432] The server receives cooking videos from users. The input is a video file uploaded by the user. The server divides this video into frames and converts it into a format suitable for analysis. During this process, image processing libraries such as OpenCV are used to output image data for each frame.
[0433] Step 2:
[0434] The server analyzes the divided image frames and identifies ingredients and seasonings from each frame. The input is the image frames obtained in step 1. Using an image recognition algorithm, specific ingredients and seasonings are detected, and this information is recorded in the database. The output is a list of identified ingredients and seasonings.
[0435] Step 3:
[0436] The server processes audio data, converts it to text, and extracts cooking instructions. The input is an audio stream from a video. Natural language processing techniques are used to convert the audio to text, from which specific cooking steps are extracted. The output is text data showing the cooking instructions.
[0437] Step 4:
[0438] The server generates text information based on the identified ingredients and extracted procedures, and then visually organizes it. The inputs are the ingredient list from step 2 and the cooking procedure from step 3. Caption information is created based on these and linked to the user interface. The output is a visually organized caption.
[0439] Step 5:
[0440] The server analyzes the viewer's emotions and proposes emotion-based edits. The input is the currently playing video frame, including the viewer's facial expressions. Using an emotion recognition model such as dlib, it estimates emotions from facial expressions and generates editing suggestions based on the results. The output is a list of proposed editing points.
[0441] Step 6:
[0442] The terminal receives captions and editing suggestions generated from the server and presents them to the user. The input is data sent from the server. The user reviews these and uses them as an interface to edit and adjust the video. The output is the video content customized by the user.
[0443] Step 7:
[0444] The user continuously monitors viewer reactions and uses that data to individually optimize the viewing experience. The input is viewer reaction data. Viewer data is analyzed and used to create new content that improves the viewing experience. The output is a project proposal for the next project that reflects viewer reactions.
[0445] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0446] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include those described above. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions shown by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0447] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.
[0448] [Third Embodiment]
[0449] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.
[0450] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.
[0451] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0452] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.
[0453] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0454] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0455] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0456] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0457] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0458] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0459] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0460] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the headset terminal 314 will be referred to as the "terminal".
[0461] In embodiments of the present invention, a process is described in detail for automatically extracting ingredients and cooking procedures from cooking videos using a video analysis system, thereby generating useful information for viewers and creators. This system operates in cooperation with a server, a terminal, and a user.
[0462] The server first receives video input uploaded from the terminal and performs preprocessing to analyze the video content. This preprocessing involves dividing the video into frames and converting the format as needed. Next, the server uses image recognition technology to identify information about food and seasonings from each frame and stores this information in a database. Audio data from the video is also extracted and converted into text using speech recognition technology. The server then uses natural language processing to extract cooking instructions from the transcribed data, organizes the instructions, and stores them in the database in chronological order.
[0463] The server then generates text information based on the food and procedure information stored in the database. This information is visually organized and used to create captions for use in the video being played. The generated captions and subtitles are provided to the user and can be edited by the user. This process enhances the accuracy of the information and helps creators effectively convey recipe information to viewers.
[0464] Furthermore, the server collects viewer reactions to published videos and analyzes viewing trends. Based on this analysis, it suggests improvements to future video content and notifies users, helping them create more engaging videos.
[0465] As a concrete example, when a user uploads a specific cooking video, the server recognizes the ingredients used, such as tomatoes and olive oil, and extracts how they were used based on the recipe. Next, the user can review the generated caption on their screen and edit details such as quantities and cooking time, making it easier for viewers to quickly replicate the recipe. This entire process provides convenient and useful information to both users and viewers, further enhancing the value of cooking videos.
[0466] The following describes the processing flow.
[0467] Step 1:
[0468] The user films a cooking video and opens the video selection screen to upload it to the platform. They click the upload button to send the video from their device to the server.
[0469] Step 2:
[0470] The server preprocesses the received video by checking and adjusting its format. Here, the video is broken down frame by frame, and the resolution and frame rate are adjusted to create a format suitable for analysis.
[0471] Step 3:
[0472] The server applies an image recognition algorithm to identify visible food items and condiments from each frame. It extracts the names and characteristics of the recognized objects and stores them in a database.
[0473] Step 4:
[0474] The server separates the audio from the video and converts it to text using speech recognition technology. The resulting text data is then subjected to natural language processing to analyze the cooking procedures and instructions, and the necessary information is extracted.
[0475] Step 5:
[0476] The server generates detailed recipe text based on ingredient and procedure information stored in the database, and creates visual captions and subtitle formats.
[0477] Step 6:
[0478] The device displays the generated captions and subtitles on the user's platform dashboard. The user can review and edit them, correcting the text information as needed.
[0479] Step 7:
[0480] The server collects viewer reaction data to published videos. It analyzes the viewing data and generates analysis results based on viewers' interests and preferences.
[0481] Step 8:
[0482] Based on the analysis of viewing data, the server sends users suggestions for improvement and ways to enhance engagement, which will be useful for future content creation.
[0483] (Example 1)
[0484] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0485] In modern dietary habits, visually presenting cooking procedures and ingredient information is crucial for viewers' understanding. However, manually extracting and organizing this information is time-consuming and laborious. Furthermore, there is no system in place to collect viewer feedback and use it to improve future content. This creates a challenge for creators, making it difficult to effectively communicate information and increase viewer engagement.
[0486] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0487] In this invention, the server includes a device that receives video data and performs preprocessing to facilitate analysis, a device that analyzes still images to identify ingredients and seasonings, and a device that converts sound data into text data and extracts cooking instructions. This allows viewers to intuitively understand the cooking instructions through subtitles, and enables improvements to future content based on viewer feedback.
[0488] "Video data" refers to digital or analog information media used to record and reproduce dynamic information that includes visual and auditory elements.
[0489] "Preprocessing" refers to the preliminary processing steps performed to facilitate data analysis, and includes data format conversion and quality improvement.
[0490] "Device" refers to a machine or system designed to perform a specific function.
[0491] A "still image" is a digital or analog representation of visual information that captures a specific moment in time.
[0492] "Ingredients" refers to materials classified according to their original cultural context and use in cooking.
[0493] "Seasoning" refers to ingredients or mixtures used to add flavor or taste to food.
[0494] "Audio data" refers to audio information recorded in digital or analog format and used for analysis and playback.
[0495] "Text data" refers to linguistic information represented in digital or analog format, and is in a readable and writable form.
[0496] "Cooking instructions" refer to a series of instructions or processes that should be followed when preparing a dish.
[0497] Subtitles are textual information displayed within visual media, and are typically used to provide transcripts or supplementary information for audio content.
[0498] This invention relates to a system that uses video analysis technology to automatically extract ingredients and cooking procedures from cooking videos and present them to viewers in an easy-to-understand format. This system operates through the coordinated efforts of a server, terminals, and users.
[0499] The server first receives video data sent from the terminal, recognizes the video format before analysis, and converts the format if necessary. Here, the server uses OpenCV to split the video into still images. These split still images are used as base data for identifying ingredients and seasonings. Machine learning libraries such as TensorFlow and PyTorch are used for this image analysis, and the information is stored in a database.
[0500] Furthermore, the server extracts the audio data from the video and converts it into text data using the Google Cloud Speech-to-Text API. This text data is then analyzed and extracted using BERT or similar natural language processing models to describe the specific cooking steps. These steps are then organized into a sequence that is easy for viewers to understand.
[0501] The generated subtitle information and instructions are sent to the user's device, where the user can review and edit the captions. This editing function allows the user to supplement ingredient quantities and substitutions, ensuring viewers can accurately recreate the dish.
[0502] As a concrete example of its use, if a user uploads a cooking video of "lasagna" from their smartphone, the server analyzes it and identifies the main ingredients such as meat, tomato sauce, and cheese. Based on this information, it generates a caption, which the user can review and edit as needed to provide viewers with more accurate information, such as ingredient details and cooking time.
[0503] An example of a prompt message used in implementing this system is: "Analyze the cooking video, extract the ingredients used and cooking steps, and generate captions."
[0504] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0505] Step 1:
[0506] The terminal uploads cooking videos selected by the user to the server. The input is a cooking video. This video is sent to the server via data communication. The server temporarily stores the received video in its storage and checks the video format. The output is the saved video file.
[0507] Step 2:
[0508] The server begins preprocessing to make the video easier to analyze. The input is a saved video file. The server uses OpenCV to split the video into multiple still images. These split images are temporarily stored in a database and prepared for identification of ingredients and seasonings. The output is a set of individual still images.
[0509] Step 3:
[0510] The server analyzes each still image to identify ingredients and seasonings. The input is the set of still images obtained in step 2. Machine learning frameworks such as TensorFlow and PyTorch are used for this analysis. The server records the identified ingredient information in a database. The output is a dataset of the identified ingredient information.
[0511] Step 4:
[0512] The server extracts audio data from the video. The input is the original video file. This audio data is converted into text data using the Google Cloud Speech-to-Text API. The server then prepares the resulting text data for further analysis. The output is the text data obtained from the video's audio track.
[0513] Step 5:
[0514] The server analyzes the text data using a natural language processing model to extract cooking instructions. The input is the text data obtained in step 4. Using natural language processing techniques such as BERT, the cooking instructions are clearly analyzed, organized chronologically, and stored in a database. The output is a dataset of the organized cooking instructions.
[0515] Step 6:
[0516] The server generates captions for subtitles based on ingredient information and cooking procedure data. The inputs are the ingredient information from step 3 and the cooking procedure data from step 5. This caption information is visually organized and sent to the user's terminal. The output is the generated subtitle captions.
[0517] Step 7:
[0518] The terminal provides the user with the generated caption and displays an editable interface. The input is the caption sent from the server. The user uses this editing function to adjust ingredient quantities and supplementary information, and saves it as final information. The output is the final caption edited by the user.
[0519] (Application Example 1)
[0520] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0521] In video content, it is difficult to clearly convey cooking procedures and ingredients used to viewers. Furthermore, it is not easy for content creators to analyze viewer feedback and use it to improve future content. In this situation, there is a need for a system that effectively organizes information and provides useful information for both viewers and creators.
[0522] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0523] In this invention, the server includes means for receiving and processing video data, means for analyzing image data to identify ingredients and seasonings, and means for converting audio information into text information and extracting cooking procedures. This makes it possible to automatically extract ingredients and procedures from videos, organize them visually, and provide useful information to viewers and creators.
[0524] "Video data" refers to digital information that includes a sequence of images and sound in time.
[0525] "Data processing" is the process of performing calculations and operations to analyze, modify, or store received information.
[0526] "Image data" refers to digital data that contains still or moving visual information.
[0527] "Identifying ingredients and seasonings" is the process of identifying and classifying the substances used in cooking.
[0528] "Audio information" refers to data that records human speech and sounds in digital format.
[0529] "Text information" refers to data expressed in the form of characters or sentences.
[0530] "Extracting cooking steps" is the process of identifying and extracting the series of steps required to create a dish.
[0531] "Visual organization" means displaying information in a way that is easy to see and understand.
[0532] "Feedback" refers to information collected from viewers' reactions and opinions.
[0533] "Analyzing viewing trends" is the process of statistically investigating and analyzing viewers' behavior and preferences.
[0534] "Improvement suggestions" refer to providing specific ideas for making the existing situation better.
[0535] The system used to implement this application primarily relies on a server. The server first receives video data uploaded by users and processes it. This data processing often utilizes tools such as OpenCV or TensorFlow. The received video data is analyzed frame by frame as individual image data, identifying ingredients and seasonings. This process makes it possible to identify the elements used in a dish.
[0536] Next, the audio information contained in the video is converted into text using speech recognition technology such as the Google Cloud Speech-to-Text API. In this step, detailed cooking instructions can be extracted from the audio in cooking videos. The obtained information is then organized using Python's natural language processing libraries such as NLTK and spaCy. This makes it possible to generate subtitles in a visually easy-to-understand format.
[0537] Furthermore, the server uses Python's pandas and matplotlib to collect and analyze feedback from viewers. By analyzing viewing trends, such as which content viewers preferred, it is possible to make specific suggestions for improving future content.
[0538] For example, if a user uploads a video demonstrating a simple home cooking recipe, the server automatically extracts ingredients such as tomatoes and pasta, as well as cooking steps, and generates visually integrated captions. This allows viewers to immediately apply the recipe to their own cooking.
[0539] An example of a prompt for a generative AI model is, "Please tell me how to identify the ingredients used in the video, generate cooking instructions based on the recipe, and display them clearly for viewers." This helps the generative AI model provide specific steps and results.
[0540] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0541] Step 1:
[0542] The server receives video data from the user's terminal. This process begins with the user uploading a cooking video to the platform. The received data is a digital file containing chronologically sequential images and audio. This input video data is then subjected to the next analysis step.
[0543] Step 2:
[0544] The server uses the OpenCV library to split the received video into individual frames. TensorFlow is then used to identify ingredients and seasonings from this image data. The input is the video frames, and the output is a list of identified ingredient and seasoning names. In this step, the identified items are stored in a database for use in subsequent processes.
[0545] Step 3:
[0546] The server extracts the audio data from the video and converts it to text using the Google Cloud Speech-to-Text API. This process takes audio information as input and outputs it as transcribed text. The resulting text includes specific cooking instructions such as "chop the onions" and "add olive oil."
[0547] Step 4:
[0548] The extracted text information is processed using natural language processing with NLTK or spaCy to clarify the cooking instructions as commands. The input is transcribed audio data, and the output is an organized list of cooking instructions. This organized instruction information can be used to generate subtitles.
[0549] Step 5:
[0550] The server combines the obtained ingredient list and cooking instructions to generate visually easy-to-read subtitles and integrates them into the video. The input to this process is the identified ingredients and written procedure list, and the output is the captions displayed during video playback.
[0551] Step 6:
[0552] Viewer feedback is collected, and viewing trends are analyzed using Python's pandas and matplotlib. This data is input in the form of feedback, and the output is an analysis of which segments viewers preferred. Based on the analysis results, the server notifies the user with suggestions for improving the next content.
[0553] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0554] This invention incorporates an emotion engine into a video analysis system to individually optimize the production and viewing of cooking videos. This system operates through the cooperation of a server, terminal, and user, and provides a wide range of functions from video input to caption generation and viewer emotion recognition.
[0555] The server receives cooking videos uploaded by users and performs preprocessing for analysis. The video is divided frame by frame, converted to an appropriate format, and then image recognition technology is used to identify ingredients and seasonings from each frame, which are then stored in a database. During this process, audio data is converted to text, and cooking instructions are extracted using natural language processing technology.
[0556] Next, the server generates text information based on the identified ingredients and extracted procedures. The generated information is visually organized and used as captions and subtitles to be displayed on the video. Additionally, an emotion engine is incorporated, which analyzes the user's emotions from the audio and facial expressions in the video to suggest editing points and optimize the video.
[0557] The generated captions and sentiment analysis results are provided to the user via their device, allowing them to review and edit them to make adjustments for more effective video content delivery. For example, if the user makes expressions of tension or enjoyment in the video, the sentiment engine will recognize this and recommend editing to emphasize it.
[0558] Furthermore, the server continuously monitors viewer reactions after publication. It collects emotional responses viewers show to the video and uses this data to generate suggestions that individually optimize the viewing experience. For example, if many viewers smile at a particular part of the video, it becomes possible to create content that reflects viewer preferences, such as planning the next video using that part.
[0559] Systems that utilize emotion engines in this way provide video creators with the data and tools necessary to enhance engagement with viewers, making it easier to form deeper emotional connections. This can improve the quality of the viewing experience and increase viewer satisfaction.
[0560] The following describes the processing flow.
[0561] Step 1:
[0562] The user films a cooking video and operates the device to upload it to the platform. The device then transfers the video file to the server.
[0563] Step 2:
[0564] The server saves the received video to storage and performs preprocessing. This involves breaking down the video into frames and converting the format to make it easier to analyze.
[0565] Step 3:
[0566] The server uses image recognition technology to identify food items and condiments in each frame. It identifies the type and number of objects detected in each frame and stores this information in a database.
[0567] Step 4:
[0568] The server extracts audio from the video and converts it to text using a speech recognition engine. This allows the system to obtain the spoken words in the video as text information, and then analyze the content using natural language processing to extract the cooking instructions.
[0569] Step 5:
[0570] The server generates detailed recipe text based on the identified food items and extracted steps. This text information is visually organized and formatted as captions or subtitles for display in the video.
[0571] Step 6:
[0572] The emotion engine analyzes the user's facial expressions and voice tone in the video to understand their emotional state. This allows it to generate data that suggests areas for emphasis and improvement in editing.
[0573] Step 7:
[0574] The device displays generated captions and sentiment analysis results sent from the server on the user's dashboard. The user can then overwrite this information, edit the captions, and adjust the video based on sentiment.
[0575] Step 8:
[0576] After a video is published, the server collects viewer reactions into a database. It statistically analyzes the emotional responses viewers showed to specific parts of the video and identifies trends.
[0577] Step 9:
[0578] Based on the analysis results, the server notifies users with suggestions for improvements to the next video content and new content ideas in order to enhance viewer emotional engagement.
[0579] (Example 2)
[0580] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0581] In modern society, video content is one of the primary means of information transmission, but its production and distribution require optimization to capture the viewer's interest. However, conventional systems lack editing suggestions based on viewer emotions and detailed analysis of viewing trends, making it difficult to create effective content.
[0582] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0583] In this invention, the server includes means for receiving video information and preparing it for analysis, means for analyzing image data to identify food substances and seasoning components, and means for converting audio signals into text data and obtaining cooking processes. This enables the analysis of viewers' emotions towards video content, and allows for specific editing suggestions and content optimization based on viewers' viewing behavior.
[0584] "Video information" refers to dynamic visual data transmitted through sight and hearing.
[0585] "Preparation for analysis" refers to the process of performing the necessary processing to analyze video information and converting the data into an appropriate format.
[0586] "Image data" refers to a collection of still images saved in a digital format.
[0587] "Food substances" refer to objects and ingredients that are consumed as part of a meal.
[0588] "Seasoning ingredients" refer to substances used to add taste and flavor to food.
[0589] "Audio signal" refers to the conversion of sound waves generated by the human voice into digital data.
[0590] "Text data" refers to a digital representation of character information.
[0591] "Cooking process" refers to a series of steps or processes necessary to complete a particular dish.
[0592] "Subtitles" refer to text information displayed on images or videos that serves to supplement or explain the content.
[0593] "Emotional state" refers to the emotional reactions of a person inferred from the audio and visuals within a video.
[0594] "Editing suggestions" refer to instructions or recommendations for editing that should be done to improve the quality and effectiveness of video content.
[0595] "Viewer reactions" refer to data about the emotions and behaviors exhibited by people who watched the video.
[0596] "Viewing behavior" refers to the methods and patterns in which viewers consume video content.
[0597] "Collecting data" refers to the process of gathering and storing information in a specific format.
[0598] "Media content" refers to the collection of information and expressions conveyed through the media.
[0599] An "improvement notice" refers to information that encourages specific corrections or additions aimed at improving the content of the media.
[0600] The embodiments for carrying out the invention are described below.
[0601] This system primarily consists of the cooperation of a server, terminals, and users. Specifically, the server receives video information and prepares it for analysis. The received video is first divided into frames and converted to image formats such as JPEG as needed. Video processing software is used for this process.
[0602] Next, the server uses image recognition libraries such as TensorFlow to analyze the image data. This allows it to identify food substances and seasoning components from each image frame. Meanwhile, the audio signal is converted into text data using speech recognition technology, and the cooking process is extracted from that text. This process is performed using NLP (Natural Language Processing) techniques.
[0603] Subsequently, the server generates captions and subtitles for the content based on the conversion and analysis processes described above. The generated subtitle information is managed and provided in a way that is easy for the user to understand. Furthermore, the server has a built-in emotion engine that can analyze the audio and visual information in the video to infer the user's emotional state. The results of this emotion analysis are provided to the user in the form of editing suggestions, etc.
[0604] For example, a user who has uploaded a cooking video can be recommended specific parts of the video as points that would be interesting to viewers. The emotion engine analyzes the user's facial expressions and tone of voice, and highlights scenes where the user is enjoying themselves and speaking, as points that should be emphasized more.
[0605] An example of a prompt might be: "Based on the analysis results of the cooking video, please suggest editing points to capture the viewer's interest. In particular, focus on scenes and expressions where the user is enjoying themselves, and edit in a way that evokes positive feelings in the viewer."
[0606] The device receives generated captions and editing suggestions and provides them to the user. Based on the provided information, the user can edit the video content and create a more engaging medium for viewers. This can provide a better viewing experience and improve viewer satisfaction.
[0607] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0608] Step 1:
[0609] The server receives video information from the user as input and converts the video into a format that can be analyzed. Specifically, it divides the video into frames and converts them to JPEG format. As a result of this conversion, still images of each frame are output.
[0610] Step 2:
[0611] The server processes the output frame images using image recognition software such as TensorFlow. Each frame image is used as input to identify food substances and seasoning components. Information on the recognized substances and components is stored in a database.
[0612] Step 3:
[0613] The server takes the audio signal from the video as input and converts it into text data using speech recognition technology. Based on the converted text data, it extracts the cooking steps using NLP technology. As a result of this extraction process, the text of the cooking procedure is output.
[0614] Step 4:
[0615] The server generates content captions and subtitles based on identified food substances and extracted cooking procedure information. This generated text information is then visually organized and output in a user-friendly format.
[0616] Step 5:
[0617] The server uses an emotion engine to analyze the user's emotional state, taking audio and visual information from the video as input. This analysis generates specific editing suggestions, which are then presented to the user. This analysis helps identify points that viewers might find interesting.
[0618] Step 6:
[0619] The device receives generated captions and editing suggestions as input and provides them to the user. The user then uses this information to edit the video and output it as more engaging content. Finally, the edited video content is ready to be published.
[0620] (Application Example 2)
[0621] Next, we will explain Application Example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0622] When providing cooking videos to viewers, there is a challenge in delivering an optimal viewing experience that responds to viewers' emotions. Furthermore, there is a lack of concrete means for content creators to improve video content based on viewer reactions.
[0623] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0624] In this invention, the server includes means for receiving and pre-processing video input, means for analyzing image frames to identify ingredients and seasonings, means for converting audio data into text and extracting cooking procedures, and means for analyzing the viewer's emotions and suggesting video editing based on those emotions. This enables the provision of video content that takes the viewer's emotions into consideration and the individual optimization of the viewing experience.
[0625] The method of "receiving video input" refers to the process of acquiring video data provided by the user in a manipulateable format.
[0626] "Preprocessing" refers to the process of dividing video data into frames and converting them into the required format in order to make it analyzable.
[0627] "Analyzing image frames" is a technique that involves recognizing or identifying specific elements in each still image frame that makes up a video.
[0628] The method of "identifying ingredients and seasonings" involves detecting specific foods or seasonings from image frames within a video and classifying or naming them.
[0629] The method of "converting audio data to text" involves analyzing the audio content within a video and converting it into corresponding text.
[0630] The method of "extracting cooking procedures" is the process of clearly extracting the steps and methods of cooking from textual data and organizing that information.
[0631] "Generating text information" refers to a method of generating meaningful information, either visually or in written form, based on identified data.
[0632] "Creating captions" is the technique of constructing text or subtitles to visually present generated text information.
[0633] The method of "analyzing viewers' emotions" is a technique that estimates an emotional state by analyzing the nonverbal reactions (such as facial expressions and actions) that viewers exhibit.
[0634] The method of "providing video editing suggestions" is a technique that, based on analysis results, offers specific ideas for modifying or emphasizing parts of a video.
[0635] "Personalizing the viewing experience" is a method of adjusting video content to suit the individual preferences and needs of each viewer, taking into account their emotions and reactions.
[0636] This system is implemented by a series of programs running on a server. The server receives video from the user and first preprocesses it by dividing it into image frames. At this stage, it uses image processing libraries such as OpenCV to convert it into an analyzable representation. Next, it uses image recognition technology to identify ingredients and seasonings from each frame and stores them in a database. Audio data is converted into text using NLP technology and extracted as cooking instructions.
[0637] To analyze viewers' emotions, the server uses a model that implements emotion recognition technology. This model uses tools such as dlib to detect viewers' faces and estimate their emotions from their facial expressions. This enables the suggestion of video editing based on emotions.
[0638] The generated text information is visually organized by the server, and captions are created based on it. These captions and sentiment analysis results are provided to the user via the device. The user can then edit the video based on this information and make adjustments to enhance the viewing experience.
[0639] As a concrete example, the server sends a prompt message to an AI model in response to a cooking video filmed by a user: "What part of this cooking video made people smile the most? Identify that part and suggest editing to highlight it." This allows for the creation of videos that increase engagement and enables content improvement based on viewer reactions.
[0640] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0641] Step 1:
[0642] The server receives cooking videos from users. The input is a video file uploaded by the user. The server divides this video into frames and converts it into a format suitable for analysis. During this process, image processing libraries such as OpenCV are used to output image data for each frame.
[0643] Step 2:
[0644] The server analyzes the divided image frames and identifies ingredients and seasonings from each frame. The input is the image frames obtained in step 1. Using an image recognition algorithm, specific ingredients and seasonings are detected, and this information is recorded in the database. The output is a list of identified ingredients and seasonings.
[0645] Step 3:
[0646] The server processes audio data, converts it to text, and extracts cooking instructions. The input is an audio stream from a video. Natural language processing techniques are used to convert the audio to text, from which specific cooking steps are extracted. The output is text data showing the cooking instructions.
[0647] Step 4:
[0648] The server generates text information based on the identified ingredients and extracted procedures, and then visually organizes it. The inputs are the ingredient list from step 2 and the cooking procedure from step 3. Caption information is created based on these and linked to the user interface. The output is a visually organized caption.
[0649] Step 5:
[0650] The server analyzes the viewer's emotions and proposes emotion-based edits. The input is the currently playing video frame, including the viewer's facial expressions. Using an emotion recognition model such as dlib, it estimates emotions from facial expressions and generates editing suggestions based on the results. The output is a list of proposed editing points.
[0651] Step 6:
[0652] The terminal receives captions and editing suggestions generated from the server and presents them to the user. The input is data sent from the server. The user reviews these and uses them as an interface to edit and adjust the video. The output is the video content customized by the user.
[0653] Step 7:
[0654] The user continuously monitors viewer reactions and uses that data to individually optimize the viewing experience. The input is viewer reaction data. Viewer data is analyzed and used to create new content that improves the viewing experience. The output is a project proposal for the next project that reflects viewer reactions.
[0655] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0656] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include those described above. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions shown by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0657] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.
[0658] [Fourth Embodiment]
[0659] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.
[0660] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[0661] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0662] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.
[0663] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0664] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0665] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0666] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.
[0667] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0668] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0669] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0670] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0671] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0672] In embodiments of the present invention, a process is described in detail for automatically extracting ingredients and cooking procedures from cooking videos using a video analysis system, thereby generating useful information for viewers and creators. This system operates in cooperation with a server, a terminal, and a user.
[0673] The server first receives video input uploaded from the terminal and performs preprocessing to analyze the video content. This preprocessing involves dividing the video into frames and converting the format as needed. Next, the server uses image recognition technology to identify information about food and seasonings from each frame and stores this information in a database. Audio data from the video is also extracted and converted into text using speech recognition technology. The server then uses natural language processing to extract cooking instructions from the transcribed data, organizes the instructions, and stores them in the database in chronological order.
[0674] The server then generates text information based on the food and procedure information stored in the database. This information is visually organized and used to create captions for use in the video being played. The generated captions and subtitles are provided to the user and can be edited by the user. This process enhances the accuracy of the information and helps creators effectively convey recipe information to viewers.
[0675] Furthermore, the server collects viewer reactions to published videos and analyzes viewing trends. Based on this analysis, it suggests improvements to future video content and notifies users, helping them create more engaging videos.
[0676] As a concrete example, when a user uploads a specific cooking video, the server recognizes the ingredients used, such as tomatoes and olive oil, and extracts how they were used based on the recipe. Next, the user can review the generated caption on their screen and edit details such as quantities and cooking time, making it easier for viewers to quickly replicate the recipe. This entire process provides convenient and useful information to both users and viewers, further enhancing the value of cooking videos.
[0677] The following describes the processing flow.
[0678] Step 1:
[0679] The user films a cooking video and opens the video selection screen to upload it to the platform. They click the upload button to send the video from their device to the server.
[0680] Step 2:
[0681] The server preprocesses the received video by checking and adjusting its format. Here, the video is broken down frame by frame, and the resolution and frame rate are adjusted to create a format suitable for analysis.
[0682] Step 3:
[0683] The server applies an image recognition algorithm to identify visible food items and condiments from each frame. It extracts the names and characteristics of the recognized objects and stores them in a database.
[0684] Step 4:
[0685] The server separates the audio from the video and converts it to text using speech recognition technology. The resulting text data is then subjected to natural language processing to analyze the cooking procedures and instructions, and the necessary information is extracted.
[0686] Step 5:
[0687] The server generates detailed recipe text based on ingredient and procedure information stored in the database, and creates visual captions and subtitle formats.
[0688] Step 6:
[0689] The device displays the generated captions and subtitles on the user's platform dashboard. The user can review and edit them, correcting the text information as needed.
[0690] Step 7:
[0691] The server collects viewer reaction data to published videos. It analyzes the viewing data and generates analysis results based on viewers' interests and preferences.
[0692] Step 8:
[0693] Based on the analysis of viewing data, the server sends users suggestions for improvement and ways to enhance engagement, which will be useful for future content creation.
[0694] (Example 1)
[0695] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0696] In modern dietary habits, visually presenting cooking procedures and ingredient information is crucial for viewers' understanding. However, manually extracting and organizing this information is time-consuming and laborious. Furthermore, there is no system in place to collect viewer feedback and use it to improve future content. This creates a challenge for creators, making it difficult to effectively communicate information and increase viewer engagement.
[0697] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0698] In this invention, the server includes a device that receives video data and performs preprocessing to facilitate analysis, a device that analyzes still images to identify ingredients and seasonings, and a device that converts sound data into text data and extracts cooking instructions. This allows viewers to intuitively understand the cooking instructions through subtitles, and enables improvements to future content based on viewer feedback.
[0699] "Video data" refers to digital or analog information media used to record and reproduce dynamic information that includes visual and auditory elements.
[0700] "Preprocessing" refers to the preliminary processing steps performed to facilitate data analysis, and includes data format conversion and quality improvement.
[0701] "Device" refers to a machine or system designed to perform a specific function.
[0702] A "still image" is a digital or analog representation of visual information that captures a specific moment in time.
[0703] "Ingredients" refers to materials classified according to their original cultural context and use in cooking.
[0704] "Seasoning" refers to ingredients or mixtures used to add flavor or taste to food.
[0705] "Audio data" refers to audio information recorded in digital or analog format and used for analysis and playback.
[0706] "Text data" refers to linguistic information represented in digital or analog format, and is in a readable and writable form.
[0707] "Cooking instructions" refer to a series of instructions or processes that should be followed when preparing a dish.
[0708] Subtitles are textual information displayed within visual media, and are typically used to provide transcripts or supplementary information for audio content.
[0709] This invention relates to a system that uses video analysis technology to automatically extract ingredients and cooking procedures from cooking videos and present them to viewers in an easy-to-understand format. This system operates through the coordinated efforts of a server, terminals, and users.
[0710] The server first receives video data sent from the terminal, recognizes the video format before analysis, and converts the format if necessary. Here, the server uses OpenCV to split the video into still images. These split still images are used as base data for identifying ingredients and seasonings. Machine learning libraries such as TensorFlow and PyTorch are used for this image analysis, and the information is stored in a database.
[0711] Furthermore, the server extracts the audio data from the video and converts it into text data using the Google Cloud Speech-to-Text API. This text data is then analyzed and extracted using BERT or similar natural language processing models to describe the specific cooking steps. These steps are then organized into a sequence that is easy for viewers to understand.
[0712] The generated subtitle information and instructions are sent to the user's device, where the user can review and edit the captions. This editing function allows the user to supplement ingredient quantities and substitutions, ensuring viewers can accurately recreate the dish.
[0713] As a concrete example of its use, if a user uploads a cooking video of "lasagna" from their smartphone, the server analyzes it and identifies the main ingredients such as meat, tomato sauce, and cheese. Based on this information, it generates a caption, which the user can review and edit as needed to provide viewers with more accurate information, such as ingredient details and cooking time.
[0714] An example of a prompt message used in implementing this system is: "Analyze the cooking video, extract the ingredients used and cooking steps, and generate captions."
[0715] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0716] Step 1:
[0717] The terminal uploads cooking videos selected by the user to the server. The input is a cooking video. This video is sent to the server via data communication. The server temporarily stores the received video in its storage and checks the video format. The output is the saved video file.
[0718] Step 2:
[0719] The server begins preprocessing to make the video easier to analyze. The input is a saved video file. The server uses OpenCV to split the video into multiple still images. These split images are temporarily stored in a database and prepared for identification of ingredients and seasonings. The output is a set of individual still images.
[0720] Step 3:
[0721] The server analyzes each still image to identify ingredients and seasonings. The input is the set of still images obtained in step 2. Machine learning frameworks such as TensorFlow and PyTorch are used for this analysis. The server records the identified ingredient information in a database. The output is a dataset of the identified ingredient information.
[0722] Step 4:
[0723] The server extracts audio data from the video. The input is the original video file. This audio data is converted into text data using the Google Cloud Speech-to-Text API. The server then prepares the resulting text data for further analysis. The output is the text data obtained from the video's audio track.
[0724] Step 5:
[0725] The server analyzes the text data using a natural language processing model to extract cooking instructions. The input is the text data obtained in step 4. Using natural language processing techniques such as BERT, the cooking instructions are clearly analyzed, organized chronologically, and stored in a database. The output is a dataset of the organized cooking instructions.
[0726] Step 6:
[0727] The server generates captions for subtitles based on ingredient information and cooking procedure data. The inputs are the ingredient information from step 3 and the cooking procedure data from step 5. This caption information is visually organized and sent to the user's terminal. The output is the generated subtitle captions.
[0728] Step 7:
[0729] The terminal provides the user with the generated caption and displays an editable interface. The input is the caption sent from the server. The user uses this editing function to adjust ingredient quantities and supplementary information, and saves it as final information. The output is the final caption edited by the user.
[0730] (Application Example 1)
[0731] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0732] In video content, it is difficult to clearly convey cooking procedures and ingredients used to viewers. Furthermore, it is not easy for content creators to analyze viewer feedback and use it to improve future content. In this situation, there is a need for a system that effectively organizes information and provides useful information for both viewers and creators.
[0733] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0734] In this invention, the server includes means for receiving and processing video data, means for analyzing image data to identify ingredients and seasonings, and means for converting audio information into text information and extracting cooking procedures. This makes it possible to automatically extract ingredients and procedures from videos, organize them visually, and provide useful information to viewers and creators.
[0735] "Video data" refers to digital information that includes a sequence of images and sound in time.
[0736] "Data processing" is the process of performing calculations and operations to analyze, modify, or store received information.
[0737] "Image data" refers to digital data that contains still or moving visual information.
[0738] "Identifying ingredients and seasonings" is the process of identifying and classifying the substances used in cooking.
[0739] "Audio information" refers to data that records human speech and sounds in digital format.
[0740] "Text information" refers to data expressed in the form of characters or sentences.
[0741] "Extracting cooking steps" is the process of identifying and extracting the series of steps required to create a dish.
[0742] "Visual organization" means displaying information in a way that is easy to see and understand.
[0743] "Feedback" refers to information collected from viewers' reactions and opinions.
[0744] "Analyzing viewing trends" is the process of statistically investigating and analyzing viewers' behavior and preferences.
[0745] "Improvement suggestions" refer to providing specific ideas for making the existing situation better.
[0746] The system used to implement this application primarily relies on a server. The server first receives video data uploaded by users and processes it. This data processing often utilizes tools such as OpenCV or TensorFlow. The received video data is analyzed frame by frame as individual image data, identifying ingredients and seasonings. This process makes it possible to identify the elements used in a dish.
[0747] Next, the audio information contained in the video is converted into text using speech recognition technology such as the Google Cloud Speech-to-Text API. In this step, detailed cooking instructions can be extracted from the audio in cooking videos. The obtained information is then organized using Python's natural language processing libraries such as NLTK and spaCy. This makes it possible to generate subtitles in a visually easy-to-understand format.
[0748] Furthermore, the server uses Python's pandas and matplotlib to collect and analyze feedback from viewers. By analyzing viewing trends, such as which content viewers preferred, it is possible to make specific suggestions for improving future content.
[0749] For example, if a user uploads a video demonstrating a simple home cooking recipe, the server automatically extracts ingredients such as tomatoes and pasta, as well as cooking steps, and generates visually integrated captions. This allows viewers to immediately apply the recipe to their own cooking.
[0750] An example of a prompt for a generative AI model is, "Please tell me how to identify the ingredients used in the video, generate cooking instructions based on the recipe, and display them clearly for viewers." This helps the generative AI model provide specific steps and results.
[0751] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0752] Step 1:
[0753] The server receives video data from the user's terminal. This process begins with the user uploading a cooking video to the platform. The received data is a digital file containing chronologically sequential images and audio. This input video data is then subjected to the next analysis step.
[0754] Step 2:
[0755] The server uses the OpenCV library to split the received video into individual frames. TensorFlow is then used to identify ingredients and seasonings from this image data. The input is the video frames, and the output is a list of identified ingredient and seasoning names. In this step, the identified items are stored in a database for use in subsequent processes.
[0756] Step 3:
[0757] The server extracts the audio data from the video and converts it to text using the Google Cloud Speech-to-Text API. This process takes audio information as input and outputs it as transcribed text. The resulting text includes specific cooking instructions such as "chop the onions" and "add olive oil."
[0758] Step 4:
[0759] The extracted text information is processed using natural language processing with NLTK or spaCy to clarify the cooking instructions as commands. The input is transcribed audio data, and the output is an organized list of cooking instructions. This organized instruction information can be used to generate subtitles.
[0760] Step 5:
[0761] The server combines the obtained ingredient list and cooking instructions to generate visually easy-to-read subtitles and integrates them into the video. The input to this process is the identified ingredients and written procedure list, and the output is the captions displayed during video playback.
[0762] Step 6:
[0763] Viewer feedback is collected, and viewing trends are analyzed using Python's pandas and matplotlib. This data is input in the form of feedback, and the output is an analysis of which segments viewers preferred. Based on the analysis results, the server notifies the user with suggestions for improving the next content.
[0764] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0765] This invention incorporates an emotion engine into a video analysis system to individually optimize the production and viewing of cooking videos. This system operates through the cooperation of a server, terminal, and user, and provides a wide range of functions from video input to caption generation and viewer emotion recognition.
[0766] The server receives cooking videos uploaded by users and performs preprocessing for analysis. The video is divided frame by frame, converted to an appropriate format, and then image recognition technology is used to identify ingredients and seasonings from each frame, which are then stored in a database. During this process, audio data is converted to text, and cooking instructions are extracted using natural language processing technology.
[0767] Next, the server generates text information based on the identified ingredients and extracted procedures. The generated information is visually organized and used as captions and subtitles to be displayed on the video. Additionally, an emotion engine is incorporated, which analyzes the user's emotions from the audio and facial expressions in the video to suggest editing points and optimize the video.
[0768] The generated captions and sentiment analysis results are provided to the user via their device, allowing them to review and edit them to make adjustments for more effective video content delivery. For example, if the user makes expressions of tension or enjoyment in the video, the sentiment engine will recognize this and recommend editing to emphasize it.
[0769] Furthermore, the server continuously monitors viewer reactions after publication. It collects emotional responses viewers show to the video and uses this data to generate suggestions that individually optimize the viewing experience. For example, if many viewers smile at a particular part of the video, it becomes possible to create content that reflects viewer preferences, such as planning the next video using that part.
[0770] Systems that utilize emotion engines in this way provide video creators with the data and tools necessary to enhance engagement with viewers, making it easier to form deeper emotional connections. This can improve the quality of the viewing experience and increase viewer satisfaction.
[0771] The following describes the processing flow.
[0772] Step 1:
[0773] The user films a cooking video and operates the device to upload it to the platform. The device then transfers the video file to the server.
[0774] Step 2:
[0775] The server saves the received video to storage and performs preprocessing. This involves breaking down the video into frames and converting the format to make it easier to analyze.
[0776] Step 3:
[0777] The server uses image recognition technology to identify food items and condiments in each frame. It identifies the type and number of objects detected in each frame and stores this information in a database.
[0778] Step 4:
[0779] The server extracts audio from the video and converts it to text using a speech recognition engine. This allows the system to obtain the spoken words in the video as text information, and then analyze the content using natural language processing to extract the cooking instructions.
[0780] Step 5:
[0781] The server generates detailed recipe text based on the identified food items and extracted steps. This text information is visually organized and formatted as captions or subtitles for display in the video.
[0782] Step 6:
[0783] The emotion engine analyzes the user's facial expressions and voice tone in the video to understand their emotional state. This allows it to generate data that suggests areas for emphasis and improvement in editing.
[0784] Step 7:
[0785] The device displays generated captions and sentiment analysis results sent from the server on the user's dashboard. The user can then overwrite this information, edit the captions, and adjust the video based on sentiment.
[0786] Step 8:
[0787] After a video is published, the server collects viewer reactions into a database. It statistically analyzes the emotional responses viewers showed to specific parts of the video and identifies trends.
[0788] Step 9:
[0789] Based on the analysis results, the server notifies users with suggestions for improvements to the next video content and new content ideas in order to enhance viewer emotional engagement.
[0790] (Example 2)
[0791] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0792] In modern society, video content is one of the primary means of information transmission, but its production and distribution require optimization to capture the viewer's interest. However, conventional systems lack editing suggestions based on viewer emotions and detailed analysis of viewing trends, making it difficult to create effective content.
[0793] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0794] In this invention, the server includes means for receiving video information and preparing it for analysis, means for analyzing image data to identify food substances and seasoning components, and means for converting audio signals into text data and obtaining cooking processes. This enables the analysis of viewers' emotions towards video content, and allows for specific editing suggestions and content optimization based on viewers' viewing behavior.
[0795] "Video information" refers to dynamic visual data transmitted through sight and hearing.
[0796] "Preparation for analysis" refers to the process of performing the necessary processing to analyze video information and converting the data into an appropriate format.
[0797] "Image data" refers to a collection of still images saved in a digital format.
[0798] "Food substances" refer to objects and ingredients that are consumed as part of a meal.
[0799] "Seasoning ingredients" refer to substances used to add taste and flavor to food.
[0800] "Audio signal" refers to the conversion of sound waves generated by the human voice into digital data.
[0801] "Text data" refers to a digital representation of character information.
[0802] "Cooking process" refers to a series of steps or processes necessary to complete a particular dish.
[0803] "Subtitles" refer to text information displayed on images or videos that serves to supplement or explain the content.
[0804] "Emotional state" refers to the emotional reactions of a person inferred from the audio and visuals within a video.
[0805] "Editing suggestions" refer to instructions or recommendations for editing that should be done to improve the quality and effectiveness of video content.
[0806] "Viewer reactions" refer to data about the emotions and behaviors exhibited by people who watched the video.
[0807] "Viewing behavior" refers to the methods and patterns in which viewers consume video content.
[0808] "Collecting data" refers to the process of gathering and storing information in a specific format.
[0809] "Media content" refers to the collection of information and expressions conveyed through the media.
[0810] An "improvement notice" refers to information that encourages specific corrections or additions aimed at improving the content of the media.
[0811] The embodiments for carrying out the invention are described below.
[0812] This system primarily consists of the cooperation of a server, terminals, and users. Specifically, the server receives video information and prepares it for analysis. The received video is first divided into frames and converted to image formats such as JPEG as needed. Video processing software is used for this process.
[0813] Next, the server uses image recognition libraries such as TensorFlow to analyze the image data. This allows it to identify food substances and seasoning components from each image frame. Meanwhile, the audio signal is converted into text data using speech recognition technology, and the cooking process is extracted from that text. This process is performed using NLP (Natural Language Processing) techniques.
[0814] Subsequently, the server generates captions and subtitles for the content based on the conversion and analysis processes described above. The generated subtitle information is managed and provided in a way that is easy for the user to understand. Furthermore, the server has a built-in emotion engine that can analyze the audio and visual information in the video to infer the user's emotional state. The results of this emotion analysis are provided to the user in the form of editing suggestions, etc.
[0815] For example, a user who has uploaded a cooking video can be recommended specific parts of the video as points that would be interesting to viewers. The emotion engine analyzes the user's facial expressions and tone of voice, and highlights scenes where the user is enjoying themselves and speaking, as points that should be emphasized more.
[0816] An example of a prompt might be: "Based on the analysis results of the cooking video, please suggest editing points to capture the viewer's interest. In particular, focus on scenes and expressions where the user is enjoying themselves, and edit in a way that evokes positive feelings in the viewer."
[0817] The device receives generated captions and editing suggestions and provides them to the user. Based on the provided information, the user can edit the video content and create a more engaging medium for viewers. This can provide a better viewing experience and improve viewer satisfaction.
[0818] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0819] Step 1:
[0820] The server receives video information from the user as input and converts the video into a format that can be analyzed. Specifically, it divides the video into frames and converts them to JPEG format. As a result of this conversion, still images of each frame are output.
[0821] Step 2:
[0822] The server processes the output frame images using image recognition software such as TensorFlow. Each frame image is used as input to identify food substances and seasoning components. Information on the recognized substances and components is stored in a database.
[0823] Step 3:
[0824] The server takes the audio signal from the video as input and converts it into text data using speech recognition technology. Based on the converted text data, it extracts the cooking steps using NLP technology. As a result of this extraction process, the text of the cooking procedure is output.
[0825] Step 4:
[0826] The server generates content captions and subtitles based on identified food substances and extracted cooking procedure information. This generated text information is then visually organized and output in a user-friendly format.
[0827] Step 5:
[0828] The server uses an emotion engine to analyze the user's emotional state, taking audio and visual information from the video as input. This analysis generates specific editing suggestions, which are then presented to the user. This analysis helps identify points that viewers might find interesting.
[0829] Step 6:
[0830] The device receives generated captions and editing suggestions as input and provides them to the user. The user then uses this information to edit the video and output it as more engaging content. Finally, the edited video content is ready to be published.
[0831] (Application Example 2)
[0832] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0833] When providing cooking videos to viewers, there is a challenge in delivering an optimal viewing experience that responds to viewers' emotions. Furthermore, there is a lack of concrete means for content creators to improve video content based on viewer reactions.
[0834] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0835] In this invention, the server includes means for receiving and pre-processing video input, means for analyzing image frames to identify ingredients and seasonings, means for converting audio data into text and extracting cooking procedures, and means for analyzing the viewer's emotions and suggesting video editing based on those emotions. This enables the provision of video content that takes the viewer's emotions into consideration and the individual optimization of the viewing experience.
[0836] The method of "receiving video input" refers to the process of acquiring video data provided by the user in a manipulateable format.
[0837] "Preprocessing" refers to the process of dividing video data into frames and converting them into the required format in order to make it analyzable.
[0838] "Analyzing image frames" is a technique that involves recognizing or identifying specific elements in each still image frame that makes up a video.
[0839] The method of "identifying ingredients and seasonings" involves detecting specific foods or seasonings from image frames within a video and classifying or naming them.
[0840] The method of "converting audio data to text" involves analyzing the audio content within a video and converting it into corresponding text.
[0841] The method of "extracting cooking procedures" is the process of clearly extracting the steps and methods of cooking from textual data and organizing that information.
[0842] "Generating text information" refers to a method of generating meaningful information, either visually or in written form, based on identified data.
[0843] "Creating captions" is the technique of constructing text or subtitles to visually present generated text information.
[0844] The method of "analyzing viewers' emotions" is a technique that estimates an emotional state by analyzing the nonverbal reactions (such as facial expressions and actions) that viewers exhibit.
[0845] The method of "providing video editing suggestions" is a technique that, based on analysis results, offers specific ideas for modifying or emphasizing parts of a video.
[0846] "Personalizing the viewing experience" is a method of adjusting video content to suit the individual preferences and needs of each viewer, taking into account their emotions and reactions.
[0847] This system is implemented by a series of programs running on a server. The server receives video from the user and first preprocesses it by dividing it into image frames. At this stage, it uses image processing libraries such as OpenCV to convert it into an analyzable representation. Next, it uses image recognition technology to identify ingredients and seasonings from each frame and stores them in a database. Audio data is converted into text using NLP technology and extracted as cooking instructions.
[0848] To analyze viewers' emotions, the server uses a model that implements emotion recognition technology. This model uses tools such as dlib to detect viewers' faces and estimate their emotions from their facial expressions. This enables the suggestion of video editing based on emotions.
[0849] The generated text information is visually organized by the server, and captions are created based on it. These captions and sentiment analysis results are provided to the user via the device. The user can then edit the video based on this information and make adjustments to enhance the viewing experience.
[0850] As a concrete example, the server sends a prompt message to an AI model in response to a cooking video filmed by a user: "What part of this cooking video made people smile the most? Identify that part and suggest editing to highlight it." This allows for the creation of videos that increase engagement and enables content improvement based on viewer reactions.
[0851] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0852] Step 1:
[0853] The server receives cooking videos from users. The input is a video file uploaded by the user. The server divides this video into frames and converts it into a format suitable for analysis. During this process, image processing libraries such as OpenCV are used to output image data for each frame.
[0854] Step 2:
[0855] The server analyzes the divided image frames and identifies ingredients and seasonings from each frame. The input is the image frames obtained in step 1. Using an image recognition algorithm, specific ingredients and seasonings are detected, and this information is recorded in the database. The output is a list of identified ingredients and seasonings.
[0856] Step 3:
[0857] The server processes audio data, converts it to text, and extracts cooking instructions. The input is an audio stream from a video. Natural language processing techniques are used to convert the audio to text, from which specific cooking steps are extracted. The output is text data showing the cooking instructions.
[0858] Step 4:
[0859] The server generates text information based on the identified ingredients and extracted procedures, and then visually organizes it. The inputs are the ingredient list from step 2 and the cooking procedure from step 3. Caption information is created based on these and linked to the user interface. The output is a visually organized caption.
[0860] Step 5:
[0861] The server analyzes the viewer's emotions and proposes emotion-based edits. The input is the currently playing video frame, including the viewer's facial expressions. Using an emotion recognition model such as dlib, it estimates emotions from facial expressions and generates editing suggestions based on the results. The output is a list of proposed editing points.
[0862] Step 6:
[0863] The terminal receives captions and editing suggestions generated from the server and presents them to the user. The input is data sent from the server. The user reviews these and uses them as an interface to edit and adjust the video. The output is the video content customized by the user.
[0864] Step 7:
[0865] The user continuously monitors viewer reactions and uses that data to individually optimize the viewing experience. The input is viewer reaction data. Viewer data is analyzed and used to create new content that improves the viewing experience. The output is a project proposal for the next project that reflects viewer reactions.
[0866] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0867] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include those described above. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions shown by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0868] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the robot 414.
[0869] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[0870] Figure 9 shows an emotion map 400 in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.
[0871] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.
[0872] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.
[0873] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, motorcycles, etc., emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.
[0874] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."
[0875] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.
[0876] The above description primarily focuses on the functions of the data processing device 12 in relation to this disclosure. However, the system related to this disclosure is not necessarily implemented on a server. The system related to this disclosure may be implemented as a general information processing system. This disclosure may be implemented, for example, as a software program that runs on a personal computer or as an application that runs on a smartphone. The method related to this disclosure may be provided to users in SaaS (Software as a Service) format.
[0877] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing of the specific process may be performed by multiple computers, including computer 22. For example, a data generation model 58 may be provided in an external device of the data processing device 12, and the external device may generate data according to the input data.
[0878] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.
[0879] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[0880] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.
[0881] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.
[0882] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.
[0883] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.
[0884] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.
[0885] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and the like that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.
[0886] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted to be incorporated by reference.
[0887] The following is further disclosed regarding the embodiments described above.
[0888] (Claim 1)
[0889] A means for receiving video input and performing preprocessing,
[0890] A means for analyzing image frames to identify food and seasonings,
[0891] A method for converting audio data into text and extracting cooking instructions,
[0892] Means for generating text information based on identified food and extracted procedures,
[0893] A means of visually organizing the generated text information to create captions,
[0894] A means of collecting viewer reactions and analyzing viewing trends,
[0895] A system that includes this.
[0896] (Claim 2)
[0897] The system according to claim 1, further comprising means for providing the generated caption to the user and enabling editing.
[0898] (Claim 3)
[0899] The system according to claim 1, further comprising means for notifying viewers of suggestions for improving content based on their viewing habits.
[0900] "Example 1"
[0901] (Claim 1)
[0902] A device that receives video data and performs preprocessing to facilitate analysis,
[0903] A device that analyzes still images to identify ingredients and seasonings,
[0904] A device that converts sound data into text data and extracts cooking instructions,
[0905] A device that generates textual information based on identified ingredients and extracted procedures,
[0906] A device that visually organizes generated text information and creates subtitles for viewers,
[0907] A device that acquires viewer reactions and analyzes viewing trends,
[0908] A device that creates a plan for improvement to the next content based on the results of data analysis,
[0909] A system that includes this.
[0910] (Claim 2)
[0911] The system according to claim 1, comprising a device for manipulating and providing generated subtitles in an editable format.
[0912] (Claim 3)
[0913] The system according to claim 1, comprising a device for notifying viewers of the next content improvement suggestions based on the results of viewer analysis.
[0914] "Application Example 1"
[0915] (Claim 1)
[0916] A means for receiving video data and processing the data,
[0917] A means of identifying ingredients and seasonings by analyzing image data,
[0918] A means for converting audio information into text information and extracting cooking instructions,
[0919] Means for generating textual information based on identified ingredients and extracted procedures,
[0920] A method for organizing the generated text information to create subtitles,
[0921] A means of collecting viewer feedback and analyzing viewing trends,
[0922] A means of proposing improvements to the next content based on the analysis results,
[0923] A system that includes this.
[0924] (Claim 2)
[0925] The system according to claim 1, which provides generated subtitles to the user and enables editing.
[0926] (Claim 3)
[0927] The system according to claim 1 for notifying improvement suggestions.
[0928] "Example 2 of combining an emotion engine"
[0929] (Claim 1)
[0930] A means of receiving video information and preparing it for analysis,
[0931] A means of analyzing image data to identify food substances and seasoning components,
[0932] A means for converting audio signals into text data and obtaining cooking procedures,
[0933] Means for generating document information based on identified food substances and acquired processes,
[0934] A means of visually organizing the generated document information to generate subtitles,
[0935] A method for analyzing emotional states based on audio and video and making editing suggestions,
[0936] A means of collecting viewer reactions as data and analyzing viewing behavior,
[0937] A system that includes this.
[0938] (Claim 2)
[0939] The system according to claim 1, further comprising means for making the generated subtitles operable and providing them to the user.
[0940] (Claim 3)
[0941] The system according to claim 1, comprising means for notifying viewers of improvements to the media content based on their viewing behavior.
[0942] "Application example 2 when combining with an emotional engine"
[0943] (Claim 1)
[0944] A means for receiving video input and performing preprocessing,
[0945] A means for analyzing image frames to identify ingredients and seasonings,
[0946] A method for converting audio data into text and extracting cooking instructions,
[0947] Means for generating text information based on identified ingredients and extracted procedures,
[0948] A means of visually organizing the generated text information to create captions,
[0949] A means of analyzing viewers' emotions and suggesting video editing based on those emotions,
[0950] A means of collecting viewer feedback and individually optimizing the viewing experience,
[0951] A system that includes this.
[0952] (Claim 2)
[0953] The system according to claim 1, further comprising means for providing the generated caption to the user and enabling editing.
[0954] (Claim 3)
[0955] The system according to claim 1, further comprising means for notifying viewers of suggestions for improving content based on the results of sentiment analysis. [Explanation of Symbols]
[0956] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots< / url:> < / url:> < / url:> < / url:>
Claims
1. A means for receiving video input and performing preprocessing, A means for analyzing image frames to identify food and seasonings, A method for converting audio data into text and extracting cooking instructions, Means for generating text information based on identified food and extracted procedures, A means of visually organizing the generated text information to create captions, A means of collecting viewer reactions and analyzing viewing trends, A system that includes this.
2. The system according to claim 1, further comprising means for providing the generated caption to the user and enabling editing.
3. The system according to claim 1, further comprising means for notifying viewers of suggestions for improving content based on their viewing habits.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A