System

The system addresses the limitations of generative AI by using a video analysis model and natural language generation model to generate accurate and emotion-sensitive natural language summaries from video content.

JP2026014288APending Publication Date: 2026-01-29SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024115285
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-07-18
Publication Date
2026-01-29

AI Technical Summary

Technical Problem

Conventional generative AI systems struggle to provide accurate and efficient natural language responses to non-verbal communication, such as hobby music or dance, as they are limited to immature verbalization and lack the ability to analyze and summarize non-verbal content like videos.

Method used

A system that includes a video analysis model and a natural language generation model to extract features from video frames, normalize and pad them to a fixed length, enabling accurate natural language generation.

Benefits of technology

Enables efficient and accurate analysis of non-verbal information from videos, generating natural language responses that summarize and provide actionable advice based on the content, considering user emotions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026014288000001_ABST
    Figure 2026014288000001_ABST
Patent Text Reader

Abstract

To provide a system for generating a natural language text from a moving image which is a non-linguistic content.SOLUTION: The data processing system includes means for reading a video file provided by a user, means for extracting each frame of the video file, resizing the frame to a predetermined size, and normalizing the frame, means for extracting the normalized frame as a feature, means for inputting the feature to a machine learning model to generate a natural language text, and means for returning the generated natural language text.SELECTED DRAWING: Figure 11
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The technology of the present disclosure relates to a system. [Background technology]

[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]

[0004] Conventional generative AI demonstrates excellent performance in fields that require verbalization and systematization, but has the problem of only being able to return immature answers in fields that require non-verbal communication, such as hobby music or dance. Many modern users often obtain information through non-verbal content such as videos, and there is a demand for technology to analyze and verbalize this information. The objective of this invention is to provide a system that can analyze such non-verbal information and provide accurate answers in natural language. [Means for solving the problem]

[0005] The present invention provides a system that includes a means for reading a video file provided by a user, a means for extracting each frame of the video file, resizing the frame to a predetermined size, and normalizing the frame, a means for extracting features from the normalized frames, a means for inputting the features into a machine learning model to generate natural language text, and a means for returning the generated natural language text. This realizes a technology that analyzes non-verbal information and enables responses in natural language. In particular, by using a pre-trained video analysis model and a natural language generation model and including a means for padding features to a fixed length, more accurate responses can be generated.

[0006] "User" refers to a person who uses the system.

[0007] "Video file" refers to an electronic file that records video data.

[0008] A "frame" refers to each still image that makes up a video file.

[0009] "Resizing" refers to the process of changing the size of an image.

[0010] "Normalization" refers to the process of constraining data values ​​to a specific range.

[0011] A "feature" refers to the information contained in data expressed as a numerical value.

[0012] A "machine learning model" refers to an algorithm that learns from data and performs a specific task.

[0013] "Natural language text" refers to text written in a language commonly used by humans.

[0014] "Means for returning" refers to a method for sending the generated text to the user.

[0015] A "system" refers to a set of devices and software designed to perform a specific function.

[0016] A "pre-trained model" refers to a machine learning model that has been trained in advance using a large amount of data.

[0017] "Padding to fixed length" refers to the process of adding missing data to make the length of the data constant.

[0018] "Video analysis model" refers to a machine learning model designed to extract features from videos.

[0019] "Natural language generation model" refers to a machine learning model that generates natural language text based on input data. [Brief explanation of the drawings]

[0020] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9]1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION

[0021] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.

[0022] First, the terms used in the following description will be explained.

[0023] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).

[0024] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.

[0025] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.

[0026] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.

[0027] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."

[0028] [First embodiment]

[0029] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.

[0030] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0031] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0032] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.

[0033] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0034] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0035] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.

[0036] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0037] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0038] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0039] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0040] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0041] The present invention is a system that analyzes a video file provided by a user and generates natural language text based on the video file. The present invention will be described with specific examples.

[0042] First, a user uses a terminal to provide a path to a video file to the server. The video file is any video data, such as a dance practice video.

[0043] The server then reads the provided video file. Once the video file is read, the server extracts each frame to analyze its content. Each frame is a separate still image, which is then resized to a specific size, for example, 224x224 pixels.

[0044] Each resized frame is then normalized, a process that scales the frame's data to the range 0 to 1, allowing machine learning models to process the data more efficiently.

[0045] Next, the normalized frames are extracted as features. These features are numerical representations of specific information contained in the data. The extracted features are used as input for machine learning models.

[0046] The server uses these features to generate natural language text using a pre-trained video analysis model and natural language generation model. First, the video analysis model analyzes the features and extracts the necessary information. Next, that information is input into the natural language generation model, which generates natural language text based on the content of the video.

[0047] For example, based on a dance video provided by the user, a response such as, "This dance consists of basic hip-hop steps, and it's important to step left and right in time with the rhythm" is generated.

[0048] Finally, the server returns the generated natural language text to the user, who can then check and refer to the generated text through their terminal.

[0049] Through this series of processes, the present invention realizes a technology that analyzes non-verbal information and enables responses in natural language. In particular, the use of a pre-trained model enables more accurate analysis and responses.

[0050] The processing flow will be explained below.

[0051] Step 1:

[0052] The user uses the terminal to provide the server with the path of the video file to be analyzed, for example, the path of a dance practice video file.

[0053] Step 2:

[0054] The server receives the path to the video file provided by the user and opens the file using the cv2.VideoCapture function.

[0055] Step 3:

[0056] The server loads the video file frame by frame. For each frame, it does the following until it reaches the end of the video:

[0057] Step 4:

[0058] Use the cv2.resize function to resize the frame loaded by the server to a given size, for example 224x224 pixels.

[0059] Step 5:

[0060] The server normalizes the resized frame by dividing each pixel's value by 255, scaling it to the range 0 to 1.

[0061] Step 6:

[0062] The server adds the normalized frame to the list as a feature. This operation is performed for all frames.

[0063] Step 7:

[0064] Once the server has finished processing all frames, it converts the feature list into a NumPy array.

[0065] Step 8:

[0066] The pad_sequences function is used to pad the feature data acquired by the server to a fixed length, thereby preparing the input format for the machine learning model.

[0067] Step 9:

[0068] The server inputs the feature data into a pre-trained video analysis model and extracts the necessary information.

[0069] Step 10:

[0070] The server generates natural language text based on the extracted information using a natural language generation model.

[0071] Step 11:

[0072] The server returns the generated natural language text to the user, for example, providing advice or explanation based on the dance content.

[0073] Step 12:

[0074] The user can view the natural language text generated through the device and take action based on it, such as improving their activities.

[0075] Example 1

[0076] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0077] Conventional systems have had difficulty extracting useful information from video files and providing it to users in natural language. Furthermore, due to the low accuracy of video analysis, the generated natural language text was often inaccurate. Furthermore, processing efficiency was low, and real-time performance was lacking.

[0078] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0079] In this invention, the server includes means for reading a video file provided by a user, means for extracting each frame of the video file, resizing it to a predetermined size, and normalizing it, means for extracting the normalized frames as features, means for inputting the features into a machine learning model to generate natural language text, and means for returning the generated natural language text. This makes it possible to efficiently and accurately extract necessary information from a video and generate natural language text based on that information.

[0080] "User" refers to a person who provides a video file to the system.

[0081] "Video file" refers to a file containing various video data that a user provides to a server.

[0082] A "frame" refers to the individual still images that make up a video file.

[0083] "Resizing" refers to the process of changing the size of a frame to a specified dimension.

[0084] "Normalization" refers to the process of scaling data to the range 0 to 1.

[0085] "Features" refer to data that expresses specific information contained in a frame as numerical data.

[0086] A "machine learning model" refers to an algorithm or system designed to perform a specific task by analyzing data and learning.

[0087] "Video analysis model" refers to a machine learning model trained to analyze video data and understand its content.

[0088] "Natural Language Generation Model" refers to a machine learning model trained to generate natural language text from data.

[0089] "Return" refers to the process by which the server provides the generated natural language text to the user.

[0090] The present invention is a system for analyzing a video file provided by a user and generating natural language text based on the video file. Specific embodiments of the present invention will be described below.

[0091] First, the user uses their device to provide the server with the path to a video file. The video file can be any video data, such as a dance practice video. The user enters the path to the video file and sends it to the server. This operation is performed through a form on a web browser or a dedicated application.

[0092] Next, the server retrieves the provided video file and prepares its contents for analysis. Specifically, it uses a video processing library such as FFmpeg to extract each frame of the video file. The extracted frames are individual still images. These frames are then resized to a predetermined size, for example, 224x224 pixels.

[0093] Each resized frame is then normalized, which is the process of scaling a frame's data to the range 0 to 1, allowing the machine learning model to process the data more efficiently.

[0094] The server then extracts features from the normalized frames. These features are numerical representations of specific information contained in the frames. Typically, deep learning models (e.g., ResNet) are used to extract the features.

[0095] The server then inputs the extracted features into a video analysis model and a natural language generation model. The video analysis model is a machine learning model trained to analyze video data and extract specific information, while the natural language generation model is a machine learning model trained to generate natural language text based on that information. For example, generative AI models such as GPT-3 are often used.

[0096] An example of a prompt to be input to a generative AI model is, "Please analyze the following dance practice video and generate natural language text based on its content," followed by information extracted by the video analysis model.

[0097] Finally, the server returns the generated natural language text to the user. The user can check the generated text through their device and use it as reference. For example, based on a dance video provided by the user, a natural language text such as "This dance consists of basic hip-hop steps, and it is important to step left and right in time with the rhythm" may be generated.

[0098] Through this series of processes, the present invention realizes a technology that can analyze non-verbal information such as video and generate natural language responses. In particular, the use of a pre-trained model enables highly accurate analysis and natural language generation.

[0099] The above is a specific embodiment for carrying out the present invention.

[0100] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0101] Step 1:

[0102] The user uses a terminal to provide the server with the path to the video file. The user enters the path to the video file into the input form on the terminal and presses the send button. This operation sends the path to the video file to the server. The input is the video file path specified by the user, and the output is the path sent to the server.

[0103] Step 2:

[0104] The server retrieves the video file and extracts each frame. The server reads the video file using the received video file path and splits the video into frames using a video processing library such as FFmpeg. The input is the video file path, and the output is the image data of each extracted frame.

[0105] Step 3:

[0106] The server resizes and normalizes each frame to a specified size. The server resizes the extracted frames to a specified size (e.g., 224x224 pixels) and normalizes the value of each pixel to a range of 0 to 1. The input is the image data of the extracted frames, and the output is the resized and normalized frame data. This process allows the machine learning model to process the data efficiently.

[0107] Step 4:

[0108] The server extracts features from the normalized frames. The server uses a deep learning model (e.g., ResNet) to extract features from the normalized frames. The input is the normalized frame data, and the output is feature data for each frame. Features are numerical representations of important information contained in the frame.

[0109] Step 5:

[0110] The server inputs the features into a video analysis model and extracts the necessary information. The server inputs the features into a pre-trained video analysis model and extracts information based on the content of the video. The input is feature data, and the output is the information extracted by the video analysis model.

[0111] Step 6:

[0112] The server inputs information from the video analysis model into a natural language generation model to generate natural language text. The server inputs information obtained from the video analysis model into a natural language generation model (e.g., GPT-3) and generates natural language text based on that information. The input is information from the video analysis model, and the output is the generated natural language text.

[0113] Step 7:

[0114] The server returns the generated natural language text to the user. The server sends the generated natural language text to the user's terminal so that the user can check it. The input is the generated natural language text, and the output is the text displayed on the user's terminal. The user can refer to this text.

[0115] (Application example 1)

[0116] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0117] In recent years, the increase in video content has created a demand for users to quickly understand the content of videos. However, systems that automatically summarize video content and provide it as text are not widely available, meaning users are unable to save time watching videos. Therefore, there is a need for a system that automatically generates video summaries and provides them to users.

[0118] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0119] In this invention, the server includes means for reading a video file provided by a user, means for extracting frames from the video file, resizing the frames to a predetermined size, and normalizing the frames, means for extracting features from the normalized frames, means for inputting the features into a machine learning model to generate a natural language text summary, and means for returning the generated natural language text summary, thereby enabling the content of the video provided by the user to be automatically analyzed and presented as a summary.

[0120] "User" means an individual or legal entity that uses the service or application.

[0121] A "video file" is a digital file that contains video data, and may also contain audio and metadata.

[0122] "Each frame" refers to an individual still image extracted from a video file.

[0123] "Resizing" is the process of changing the width and height of an image (frame) to fit a new size.

[0124] "Normalization" is the process of scaling data to a particular range (typically between 0 and 1).

[0125] "Features" are important numerical data or patterns extracted from data and input into machine learning models.

[0126] A "machine learning model" is an algorithm or model that is trained to analyze data and perform a specific task.

[0127] A "video analysis model" is a machine learning model trained to perform a specific task to extract and analyze features from video data.

[0128] A "natural language generation model" is a machine learning model trained to generate coherent natural language text based on input data.

[0129] A "natural language text summary" is a document that condenses the content of the original video and presents it in an easy-to-understand format.

[0130] "Returning means" refers to the process or mechanism for providing the generated data or information to the user.

[0131] "Padding" is an operation to fill in missing parts to make data a fixed length.

[0132] The present invention is a system that analyzes video files provided by users and generates natural language text summaries based on the files. This system is realized through the interaction of a server, a terminal, and a user.

[0133] First, a user uses a terminal to provide a video file to the server. The provided video file can be any video data, such as a dance practice video or a cooking recipe video.

[0134] The server reads the received video file. Once the video file is read, the server extracts each frame to analyze its content. Each frame is an individual still image, which is then resized to a predetermined size (e.g., 224x224 pixels). Each resized frame is then normalized, scaling the data range from 0 to 1.

[0135] Next, the normalized frames are extracted as features. Features are numerical representations of important information contained in the video frames. Based on these features, the server generates data to be input into the machine learning model. This machine learning model consists of a pre-trained video analysis model and a natural language generation model.

[0136] The video analysis model analyzes the feature data and extracts important information. This information is input into the natural language generation model, which generates a natural language text summary based on the content of the video. For example, based on a dance practice video, the summary generated might be, "This dance consists of basic hip-hop steps. It is important to step left and right in time with the rhythm."

[0137] The server then returns the generated natural language text summary to the user, who can then review the generated text and use it as a reference. Through this process, the present invention realizes a technology for analyzing non-verbal information and providing a summary in natural language.

[0138] The present invention uses the following hardware and software:

[0139] OpenCV is used to read the video file and extract frames.

[0140] A pre-trained model using TensorFlow is used to extract features.

[0141] For natural language generation, generative AI models such as GPT-2 are used.

[0142] For example, the prompt for the dance practice video is as follows:

[0143] "This video teaches basic dance steps."

[0144] For recipe videos:

[0145] This video shows the cooking steps.

[0146] By using the above method, the present invention can efficiently summarize the contents of a video and provide it to the user.

[0147] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0148] Step 1:

[0149] A user uses a terminal to provide a video file to the server. The video file is any video data, and the user specifies a file path of their choice. The input here is the path of the video file provided by the user. As an output, the server obtains the specified video file.

[0150] Step 2:

[0151] The server reads the video file and extracts each frame. Each extracted frame is an individual still image, and it uses a library such as OpenCV to extract frames from the video file. The input to this process is the video file obtained in step 1, and the output is the individual extracted frames.

[0152] Step 3:

[0153] The server resizes and normalizes each frame to a predetermined size. Specifically, it resizes each frame to 224x224 pixels and then scales the pixel data to the range of 0 to 1. Data processing involves resizing and normalization. The input is the frame extracted in step 2, and the output is the resized and normalized frame.

[0154] Step 4:

[0155] The server extracts features from the resized and normalized frames. A pre-trained model using TensorFlow is used for feature extraction. Specifically, each frame is input to a specific neural network, which outputs a feature vector. The input is the frame processed in step 3, and the output is a feature vector.

[0156] Step 5:

[0157] The server inputs the feature vectors into a machine learning model to generate a natural language text summary. A video analysis model analyzes the feature vectors and extracts important information. A natural language generation model (e.g., GPT-2) then generates a text summary based on that information. Data calculations involve converting features to text. The input is the feature vector, and the output is the natural language text summary.

[0158] Step 6:

[0159] The server returns the generated natural language text summary to the user. The generated text summary is provided in a format that the user can view through their terminal. The input is the natural language text summary generated in step 5, and the output is the text summary returned to the user's terminal. This allows the user to refer to the provided summary.

[0160] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.

[0161] The present invention combines a system that analyzes video files provided by users and generates natural language text from them with an emotion engine that recognizes the user's emotions. The present invention will be described below with specific examples.

[0162] First, the user uses the terminal to provide the server with the path to the video file to be analyzed. The video file can be any video data, such as a dance practice video or a presentation practice video.

[0163] The server receives the video file path provided by the user and opens the file. Once the video file is read, the server extracts each frame to analyze its content. Each frame is resized, for example to 224x224 pixels. Each resized frame is then normalized. Normalization is the process of scaling the frame's data to a range of 0 to 1, allowing the machine learning model to process the data more efficiently.

[0164] The normalized frames are extracted as features. These features represent specific information contained in the data as numerical data. The extracted features are used as input to a machine learning model. The server uses these features to generate natural language text using a pre-trained video analysis model and natural language generation model.

[0165] A distinctive feature of the present invention is the emotion engine. The emotion engine analyzes emotions from the user's facial expressions and voice. This analyzed emotion data is fed back to the machine learning model and reflected in the generated natural language text. For example, if the user is nervous, the generated advice is "Relax and try again."

[0166] For example, based on a dance video provided by the user, a response such as "This dance consists of basic hip-hop steps, and it is important to step left and right in time with the rhythm" is generated. Furthermore, if the user shows a happy emotion, positive feedback such as "You're doing great!" is added.

[0167] Finally, the server returns the generated natural language text to the user, who can then review the text and take action, such as improving their activities, based on the text.

[0168] Through this series of processes, the present invention realizes a technology that analyzes non-verbal information and enables responses in natural language that take the user's emotions into account. In particular, by combining a pre-trained model with an emotion engine, more accurate analysis and responses that take emotions into account become possible.

[0169] The processing flow will be explained below.

[0170] Step 1:

[0171] The user uses the terminal to provide the server with the path of the video file to be analyzed, for example, the path of a dance practice video file.

[0172] Step 2:

[0173] The server receives the path to the video file provided by the user and opens the file using the cv2.VideoCapture function.

[0174] Step 3:

[0175] The server loads the video file frame by frame. For each frame, it does the following until it reaches the end of the video:

[0176] Step 4:

[0177] Use the cv2.resize function to resize the frame loaded by the server to a given size, for example 224x224 pixels.

[0178] Step 5:

[0179] The server normalizes the resized frame by dividing each pixel's value by 255, scaling it to the range 0 to 1.

[0180] Step 6:

[0181] The server adds the normalized frame to the list as a feature. This operation is performed for all frames.

[0182] Step 7:

[0183] Once the server has finished processing all frames, it converts the feature list into a NumPy array.

[0184] Step 8:

[0185] The pad_sequences function is used to pad the feature data acquired by the server to a fixed length, thereby preparing the input format for the machine learning model.

[0186] Step 9:

[0187] The server inputs the feature data into a pre-trained video analysis model and extracts the necessary information.

[0188] Step 10:

[0189] The server runs an emotion engine using the user's real-time video and audio data to analyze the user's emotions, for example, by using facial recognition technology or voice analysis technology to identify the user's emotions.

[0190] Step 11:

[0191] The server then feeds the analyzed emotion data back to the machine learning model, which then uses that data to generate natural language text. For example, if the user is nervous, the system generates advice like, "Relax and try again."

[0192] Step 12:

[0193] The server returns the generated natural language text to the user, providing advice and explanations based on the dance content, as well as emotionally sensitive feedback.

[0194] Step 13:

[0195] The user can review the generated natural language text through the device and take action based on it to improve their activity, for example, practicing a dance step again in accordance with the new advice.

[0196] Example 2

[0197] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0198] Conventional video analysis systems were unable to consider the user's emotions when generating text from video content. As a result, the generated text could not provide advice or feedback appropriate to the user's situation, resulting in poor usability. A particular issue was the inability to properly reflect the user's emotions, such as tension or joy.

[0199] The identification process by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes: means for reading a video file provided by a user; means for extracting each frame of the video file, resizing it to a predetermined size, and normalizing it; means for extracting the normalized frames as features; means for inputting the features into a machine learning model and generating natural language text; means for returning the generated natural language text; means including an emotion engine for analyzing the user's facial expressions and voice and recognizing their emotions; and means for feeding back the emotion data to the machine learning model and reflecting it in the natural language text. This makes it possible to generate natural language text that reflects the user's emotions in the analysis results of the video content.

[0200] "User" means an individual or organization that uses the system to provide video files and receive analysis results.

[0201] A "video file" is a digital file in a video data format that a user provides to the system.

[0202] A "frame" refers to one of the still images that make up a video file, and a video is made up of a series of many of these frames.

[0203] "Resizing" is the process of changing the resolution of each frame of a video file to a specific size.

[0204] "Normalization" is the process of scaling the range of data to a certain range, typically to a range from 0 to 1.

[0205] A "feature" is a specific pattern or piece of information in data that has been quantified and is used as input into a machine learning model.

[0206] A "machine learning model" is a pre-trained algorithm or network that analyzes data and makes predictions or classifications.

[0207] "Natural language text" is text in a language format that humans can understand, generated based on the results of video analysis.

[0208] An "emotion engine" is an algorithm or system that analyzes and recognizes emotions from a user's facial expressions, voice, etc.

[0209] "Feedback" is the process of inputting analysis results and data back into the system and reflecting the results.

[0210] MODE FOR CARRYING OUT THE INVENTION

[0211] The present invention combines a system that analyzes video files provided by users and generates natural language text from them with an emotion engine that recognizes the user's emotions. The present invention will be described below with specific examples.

[0212] First, the user uses their device to provide the server with the path to the video file to be analyzed. The video file can be any video data, such as a dance practice video or a presentation practice video. The server receives the path to the video file provided by the user and opens the file. Once the video file is read, the server extracts each frame to analyze its content. Each frame is resized, for example, to 224x224 pixels. Each resized frame is then normalized. Normalization is the process of scaling the frame data to a range of 0 to 1, which allows the machine learning model to process the data more efficiently.

[0213] The normalized frames are extracted as features. These features represent specific information contained in the data as numerical data. The server then inputs these features into a pre-trained video analysis model to analyze the content of the video. Based on the results of the video analysis, the server then generates natural language text using a pre-trained natural language generation model.

[0214] A distinctive feature of this system is the emotion engine, which analyzes emotions from the user's facial expressions and voice. This analyzed emotion data is fed back to the machine learning model and reflected in the generated natural language text. As a specific example, based on a dance video provided by the user, a response such as "This dance consists of basic hip-hop steps, and it is important to step left and right in time with the rhythm" is generated. Furthermore, if the user shows a happy emotion, positive feedback such as "You're doing great!" is added.

[0215] Finally, the server returns the generated natural language text to the user, who can review the text through their device and take action based on it, such as improving their activities.

[0216] Hardware and software used

[0217] This system consists of a server, a terminal, and related software components. Specifically, the following hardware and software are used:

[0218] Hardware:

[0219] Server: A high-performance computer that analyzes video files and generates natural language text.

[0220] Device: The device (PC, tablet, smartphone, etc.) where the user uploads the video file and reviews the generated text.

[0221] software:

[0222] Video analysis libraries: Libraries such as OpenCV for extracting, resizing, and normalizing frames from video files.

[0223] Machine learning frameworks, such as TensorFlow and PyTorch, are used to implement and run video analysis and natural language generation models.

[0224] Emotion analysis engine: A dedicated library or API for analyzing a user's facial expressions and voice.

[0225] Prompt Sentence Examples

[0226] "What happens if a user uploads a dance practice video?"

[0227] In response to this prompt, the system will generate an explanation like this:

[0228] "When a user uploads a dance practice video, the server analyzes the video, extracts each frame, and resizes and normalizes it. It then extracts features and generates natural language text based on them using a pre-trained model. If the user's sentiment is positive, it will add feedback like, 'You're doing great!'"

[0229] Through this series of processes, the present invention realizes a technology that analyzes non-verbal information and generates a response in natural language that takes into account the user's emotions.

[0230] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0231] Step 1:

[0232] The user inputs the path of the video file to be analyzed using the device. The input video file path is, for example, "C:\Users\SampleUser\Videos\dance_practice.mp4". This is the input.

[0233] Step 2:

[0234] The device sends the path of the input video file to the server. Specifically, it sends the path information using an HTTP request or other communication method. The output is the path information received by the server.

[0235] Step 3:

[0236] Opens the specified video file based on the video file path received by the server. Here, the input is the video file path and the output is the video file in its open state.

[0237] Step 4:

[0238] The server extracts each frame of the video file. Specifically, it uses a library such as OpenCV to obtain image data for each frame. The input is the video file, and the output is the extracted frames.

[0239] Step 5:

[0240] The server resizes each extracted frame to 224x224 pixels, for example using the OpenCV cv2.resize function. The input is the image data of the frame, and the output is the resized frames.

[0241] Step 6:

[0242] The server normalizes each resized frame, scaling each pixel value to the range 0 to 1 by dividing it by 255. The input is a set of resized frames, and the output is a set of normalized frames.

[0243] Step 7:

[0244] The server extracts features from the normalized frames. For example, it uses a convolutional neural network (CNN) to extract specific features from the frames as numerical data. The input is the normalized frames, and the output is the feature data.

[0245] Step 8:

[0246] The server analyzes the features using a pre-trained video analysis model. To obtain the analysis results, the feature data is input into the model. The input is the feature data, and the output is the analysis results.

[0247] Step 9:

[0248] The server generates text based on the analysis results using a pre-trained natural language generation model. The input is the analysis results, and the output is the generated natural language text.

[0249] Step 10:

[0250] The server uses an emotion engine to analyze the user's facial expressions and voice to obtain emotion data. The input is the user's facial expression image and voice data, and the output is emotion data.

[0251] Step 11:

[0252] The server feeds the emotion data back to the machine learning model and reflects it in the natural language text. Specifically, it adds advice and feedback according to the emotion to the text. The input is the emotion data and the initial natural language text, and the output is the final natural language text that takes emotion into account.

[0253] Step 12:

[0254] The server returns the generated natural language text to the user. Specifically, it returns the text using an HTTP response, etc. The input is the final natural language text, and the output is the text displayed on the user's terminal.

[0255] Step 13:

[0256] The user views the generated natural language text through a terminal. The input is the returned text, and the output is the user's understanding and action based on it.

[0257] The above are the processing steps of the present invention.

[0258] (Application example 2)

[0259] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0260] Conventional video analysis systems simply analyze video data and generate text without considering the user's emotions. As a result, in fields such as customer service and product recommendations, they are unable to generate responses that appropriately reflect the customer's emotions and reactions, making it difficult to improve customer satisfaction and communicate effectively. The objective of this invention is to provide more effective and personalized responses by analyzing the user's emotions and generating natural language text that reflects them.

[0261] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[0262] In this invention, the server includes means for reading a video file provided by a user, means for extracting each frame of the video file, resizing it to a predetermined size, and normalizing it, means for extracting the normalized frames as features, means for inputting the features into a machine learning model to generate natural language text, means for returning the generated natural language text, means for creating a context for generating appropriate responses and suggestions for customer service, and means for analyzing the user's emotions and reflecting them in the context. This makes it possible to generate natural language text that reflects the user's emotions, thereby enabling more effective and personalized customer service.

[0263] A "user" is an entity that uses this system to provide videos and receive analysis results.

[0264] "Video file" refers to video data provided by the user, and is the media data to be analyzed.

[0265] The term "means" refers to a method or technique for realizing the functions or processes of the present invention.

[0266] A "frame" is an individual still image that makes up a video file.

[0267] "Resizing" is a process of changing the size of an image or frame to a predetermined size.

[0268] "Normalization" is the process of scaling data to a certain range, primarily used to make data easier to handle uniformly in machine learning.

[0269] A "feature" is specific information extracted from data expressed as numerical data.

[0270] A "machine learning model" refers to an algorithm or mathematical model that learns patterns from data and makes predictions and classifications.

[0271] "Natural language text" refers to human-understandable sentences generated from analyzed video data.

[0272] "Context" refers to the background information and circumstances for generating customer service responses and suggestions, and also includes the results of user sentiment analysis.

[0273] "Means for analyzing emotions" refers to a technology or method that detects the user's emotional state from facial expressions, voice, etc., and uses that information within the system.

[0274] "Means for returning" refers to the method or technique for presenting the generated natural language text to the user.

[0275] The present invention is a system that analyzes a video file provided by a user, generates natural language text from the video file, and further combines it with an emotion engine that recognizes the user's emotions. Specific embodiments of the present invention will be described in detail below.

[0276] First, the user uses their device to provide the server with the path to the video file to be analyzed. The video file can be any video data, such as a recorded video of a customer service session. The server receives the path to the video file provided by the user and opens the file. Once the video file is read, the server extracts each frame to analyze its content. Each frame is resized, for example, to 224x224 pixels. Each resized frame is then normalized. Normalization is the process of scaling the frame data to a range of 0 to 1, which allows the machine learning model to process the data more efficiently.

[0277] The normalized frames are extracted as features. These features represent specific information contained in the data as numerical data. The extracted features are used as input to a machine learning model. The server uses these features to generate natural language text using a pre-trained video analysis model and natural language generation model.

[0278] A distinctive feature of the present invention is the emotion engine. The emotion engine analyzes emotions from the user's facial expressions and voice. This analyzed emotion data is fed back to the machine learning model and reflected in the generated natural language text. For example, if the user shows interest, a response such as "This product is new this season and is particularly popular" is generated. If the user loses interest, the response switches to "Can you tell me what other designs you like?"

[0279] Specifically, the server uses OpenCV to extract frames from the video (assuming a video capture camera), resizes the frames to 224x224 pixels, and normalizes them. The normalized data is then input into machine learning models: EmotionEngine (emotion recognition) and TextGenerationModel (natural language generation). The data obtained from sentiment analysis influences the generated text, resulting in natural language text that reflects the user's emotions.

[0280] Examples:

[0281] This application can be used in a scenario where a salesperson in a fashion store is wearing smart glasses and serving customers. When a customer sees an item of clothing that interests them, the smart glasses respond to their emotions and provide appropriate guidance, such as, "This item is new this season and is particularly popular. Please let us know if you need more information." If the customer appears to be losing interest, the smart glasses can switch to a suggestion such as, "Can you tell us what other designs you like?"

[0282] Example prompt:

[0283] "Customer interested" prompt:

[0284] Context: Your customer is interested.

[0285] Frame Data: [frame 1, frame 2, ...]

[0286] Generated text type: Product description

[0287] "Customer is losing interest" prompt:

[0288] Context: Your customer is losing interest.

[0289] Frame Data: [frame 1, frame 2, ...]

[0290] Type of generated text: Questions and Answers

[0291] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0292] Step 1:

[0293] The user uses the device to provide the server with the path to the video file to be analyzed. The path to the video file is specified as input, and the data is sent to the server as output. In this step, the specific actions are to select and send the video file.

[0294] Step 2:

[0295] The server receives the video file path provided by the user and opens the file. The input is the video file path, and the output is the video file itself. This step involves reading the target file from the file system.

[0296] Step 3:

[0297] The server extracts each frame to analyze the video file. The input is the video file, and the output is a list of extracted frames. Specifically, the server uses OpenCV to split the video into frames.

[0298] Step 4:

[0299] Each frame is resized and normalized to 224x224 pixels. The input is the extracted frame, and the output is the resized and normalized frame. Specifically, the resizing and normalization processes are performed.

[0300] Step 5:

[0301] The normalized frames are extracted as features. The input is the normalized frames and the output is the feature data. In this step, a feature extraction algorithm is applied.

[0302] Step 6:

[0303] The server inputs the feature data into a machine learning model and generates natural language text. The input is the feature data, and the output is the generated natural language text. This includes operations that utilize a pre-trained video analysis model and a natural language generation model.

[0304] Step 7:

[0305] The emotion engine analyzes emotions from the user's facial expressions and voice. The input is normalized frame and voice data, and the output is analyzed emotion data. In this step, emotion recognition algorithms are applied.

[0306] Step 8:

[0307] The emotion data is fed back into the machine learning model and reflected in the generated natural language text. The input is emotion data and normal frame data for the generated text, and the output is the final natural language text. Here, the text generation process integrates emotion data.

[0308] Step 9:

[0309] The server returns the generated natural language text to the user. The input is the final natural language text and the output is the text that is displayed on the terminal. This step includes displaying the text in a user interface.

[0310] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0311] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0312] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.

[0313] [Second embodiment]

[0314] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.

[0315] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0316] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0317] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.

[0318] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0319] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0320] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0321] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0322] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0323] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0324] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0325] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."

[0326] The present invention is a system that analyzes a video file provided by a user and generates natural language text based on the video file. The present invention will be described with specific examples.

[0327] First, a user uses a terminal to provide a path to a video file to the server. The video file is any video data, such as a dance practice video.

[0328] The server then reads the provided video file. Once the video file is read, the server extracts each frame to analyze its content. Each frame is a separate still image, which is then resized to a specific size, for example, 224x224 pixels.

[0329] Each resized frame is then normalized, a process that scales the frame's data to the range 0 to 1, allowing machine learning models to process the data more efficiently.

[0330] Next, the normalized frames are extracted as features. These features are numerical representations of specific information contained in the data. The extracted features are used as input for machine learning models.

[0331] The server uses these features to generate natural language text using a pre-trained video analysis model and natural language generation model. First, the video analysis model analyzes the features and extracts the necessary information. Next, that information is input into the natural language generation model, which generates natural language text based on the content of the video.

[0332] For example, based on a dance video provided by the user, a response such as, "This dance consists of basic hip-hop steps, and it's important to step left and right in time with the rhythm" is generated.

[0333] Finally, the server returns the generated natural language text to the user, who can then check and refer to the generated text through their terminal.

[0334] Through this series of processes, the present invention realizes a technology that analyzes non-verbal information and enables responses in natural language. In particular, the use of a pre-trained model enables more accurate analysis and responses.

[0335] The processing flow will be explained below.

[0336] Step 1:

[0337] The user uses the terminal to provide the server with the path of the video file to be analyzed, for example, the path of a dance practice video file.

[0338] Step 2:

[0339] The server receives the path to the video file provided by the user and opens the file using the cv2.VideoCapture function.

[0340] Step 3:

[0341] The server loads the video file frame by frame. For each frame, it does the following until it reaches the end of the video:

[0342] Step 4:

[0343] Use the cv2.resize function to resize the frame loaded by the server to a given size, for example 224x224 pixels.

[0344] Step 5:

[0345] The server normalizes the resized frame by dividing each pixel's value by 255, scaling it to the range 0 to 1.

[0346] Step 6:

[0347] The server adds the normalized frame to the list as a feature. This operation is performed for all frames.

[0348] Step 7:

[0349] Once the server has finished processing all frames, it converts the feature list into a NumPy array.

[0350] Step 8:

[0351] The pad_sequences function is used to pad the feature data acquired by the server to a fixed length, thereby preparing the input format for the machine learning model.

[0352] Step 9:

[0353] The server inputs the feature data into a pre-trained video analysis model and extracts the necessary information.

[0354] Step 10:

[0355] The server generates natural language text based on the extracted information using a natural language generation model.

[0356] Step 11:

[0357] The server returns the generated natural language text to the user, for example, providing advice or explanation based on the dance content.

[0358] Step 12:

[0359] The user can view the natural language text generated through the device and take action based on it, such as improving their activities.

[0360] Example 1

[0361] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0362] Conventional systems have had difficulty extracting useful information from video files and providing it to users in natural language. Furthermore, due to the low accuracy of video analysis, the generated natural language text was often inaccurate. Furthermore, processing efficiency was low, and real-time performance was lacking.

[0363] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0364] In this invention, the server includes means for reading a video file provided by a user, means for extracting each frame of the video file, resizing it to a predetermined size, and normalizing it, means for extracting the normalized frames as features, means for inputting the features into a machine learning model to generate natural language text, and means for returning the generated natural language text. This makes it possible to efficiently and accurately extract necessary information from a video and generate natural language text based on that information.

[0365] "User" refers to a person who provides a video file to the system.

[0366] "Video file" refers to a file containing various video data that a user provides to a server.

[0367] A "frame" refers to the individual still images that make up a video file.

[0368] "Resizing" refers to the process of changing the size of a frame to a specified dimension.

[0369] "Normalization" refers to the process of scaling data to the range 0 to 1.

[0370] "Features" refer to data that expresses specific information contained in a frame as numerical data.

[0371] A "machine learning model" refers to an algorithm or system designed to perform a specific task by analyzing data and learning.

[0372] "Video analysis model" refers to a machine learning model trained to analyze video data and understand its content.

[0373] "Natural Language Generation Model" refers to a machine learning model trained to generate natural language text from data.

[0374] "Return" refers to the process by which the server provides the generated natural language text to the user.

[0375] The present invention is a system for analyzing a video file provided by a user and generating natural language text based on the video file. Specific embodiments of the present invention will be described below.

[0376] First, the user uses their device to provide the server with the path to a video file. The video file can be any video data, such as a dance practice video. The user enters the path to the video file and sends it to the server. This operation is performed through a form on a web browser or a dedicated application.

[0377] Next, the server retrieves the provided video file and prepares its contents for analysis. Specifically, it uses a video processing library such as FFmpeg to extract each frame of the video file. The extracted frames are individual still images. These frames are then resized to a predetermined size, for example, 224x224 pixels.

[0378] Each resized frame is then normalized, which is the process of scaling a frame's data to the range 0 to 1, allowing the machine learning model to process the data more efficiently.

[0379] The server then extracts features from the normalized frames. These features are numerical representations of specific information contained in the frames. Typically, deep learning models (e.g., ResNet) are used to extract the features.

[0380] The server then inputs the extracted features into a video analysis model and a natural language generation model. The video analysis model is a machine learning model trained to analyze video data and extract specific information, while the natural language generation model is a machine learning model trained to generate natural language text based on that information. For example, generative AI models such as GPT-3 are often used.

[0381] An example of a prompt to be input to a generative AI model is, "Please analyze the following dance practice video and generate natural language text based on its content," followed by information extracted by the video analysis model.

[0382] Finally, the server returns the generated natural language text to the user. The user can check the generated text through their device and use it as reference. For example, based on a dance video provided by the user, a natural language text such as "This dance consists of basic hip-hop steps, and it is important to step left and right in time with the rhythm" may be generated.

[0383] Through this series of processes, the present invention realizes a technology that can analyze non-verbal information such as video and generate natural language responses. In particular, the use of a pre-trained model enables highly accurate analysis and natural language generation.

[0384] The above is a specific embodiment for carrying out the present invention.

[0385] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0386] Step 1:

[0387] The user uses a terminal to provide the server with the path to the video file. The user enters the path to the video file into the input form on the terminal and presses the send button. This operation sends the path to the video file to the server. The input is the video file path specified by the user, and the output is the path sent to the server.

[0388] Step 2:

[0389] The server retrieves the video file and extracts each frame. The server reads the video file using the received video file path and splits the video into frames using a video processing library such as FFmpeg. The input is the video file path, and the output is the image data of each extracted frame.

[0390] Step 3:

[0391] The server resizes and normalizes each frame to a specified size. The server resizes the extracted frames to a specified size (e.g., 224x224 pixels) and normalizes the value of each pixel to a range of 0 to 1. The input is the image data of the extracted frames, and the output is the resized and normalized frame data. This process allows the machine learning model to process the data efficiently.

[0392] Step 4:

[0393] The server extracts features from the normalized frames. The server uses a deep learning model (e.g., ResNet) to extract features from the normalized frames. The input is the normalized frame data, and the output is feature data for each frame. Features are numerical representations of important information contained in the frame.

[0394] Step 5:

[0395] The server inputs the features into a video analysis model and extracts the necessary information. The server inputs the features into a pre-trained video analysis model and extracts information based on the content of the video. The input is feature data, and the output is the information extracted by the video analysis model.

[0396] Step 6:

[0397] The server inputs information from the video analysis model into a natural language generation model to generate natural language text. The server inputs information obtained from the video analysis model into a natural language generation model (e.g., GPT-3) and generates natural language text based on that information. The input is information from the video analysis model, and the output is the generated natural language text.

[0398] Step 7:

[0399] The server returns the generated natural language text to the user. The server sends the generated natural language text to the user's terminal so that the user can check it. The input is the generated natural language text, and the output is the text displayed on the user's terminal. The user can refer to this text.

[0400] (Application example 1)

[0401] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0402] In recent years, the increase in video content has created a demand for users to quickly understand the content of videos. However, systems that automatically summarize video content and provide it as text are not widely available, meaning users are unable to save time watching videos. Therefore, there is a need for a system that automatically generates video summaries and provides them to users.

[0403] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0404] In this invention, the server includes means for reading a video file provided by a user, means for extracting frames from the video file, resizing the frames to a predetermined size, and normalizing the frames, means for extracting features from the normalized frames, means for inputting the features into a machine learning model to generate a natural language text summary, and means for returning the generated natural language text summary, thereby enabling the content of the video provided by the user to be automatically analyzed and presented as a summary.

[0405] "User" means an individual or legal entity that uses the service or application.

[0406] A "video file" is a digital file that contains video data, and may also contain audio and metadata.

[0407] "Each frame" refers to an individual still image extracted from a video file.

[0408] "Resizing" is the process of changing the width and height of an image (frame) to fit a new size.

[0409] "Normalization" is the process of scaling data to a particular range (typically between 0 and 1).

[0410] "Features" are important numerical data or patterns extracted from data and input into machine learning models.

[0411] A "machine learning model" is an algorithm or model that is trained to analyze data and perform a specific task.

[0412] A "video analysis model" is a machine learning model trained to perform a specific task to extract and analyze features from video data.

[0413] A "natural language generation model" is a machine learning model trained to generate coherent natural language text based on input data.

[0414] A "natural language text summary" is a document that condenses the content of the original video and presents it in an easy-to-understand format.

[0415] "Returning means" refers to the process or mechanism for providing the generated data or information to the user.

[0416] "Padding" is an operation to fill in missing parts to make data a fixed length.

[0417] The present invention is a system that analyzes video files provided by users and generates natural language text summaries based on the files. This system is realized through the interaction of a server, a terminal, and a user.

[0418] First, a user uses a terminal to provide a video file to the server. The provided video file can be any video data, such as a dance practice video or a cooking recipe video.

[0419] The server reads the received video file. Once the video file is read, the server extracts each frame to analyze its content. Each frame is an individual still image, which is then resized to a predetermined size (e.g., 224x224 pixels). Each resized frame is then normalized, scaling the data range from 0 to 1.

[0420] Next, the normalized frames are extracted as features. Features are numerical representations of important information contained in the video frames. Based on these features, the server generates data to be input into the machine learning model. This machine learning model consists of a pre-trained video analysis model and a natural language generation model.

[0421] The video analysis model analyzes the feature data and extracts important information. This information is input into the natural language generation model, which generates a natural language text summary based on the content of the video. For example, based on a dance practice video, the summary generated might be, "This dance consists of basic hip-hop steps. It is important to step left and right in time with the rhythm."

[0422] The server then returns the generated natural language text summary to the user, who can then review the generated text and use it as a reference. Through this process, the present invention realizes a technology for analyzing non-verbal information and providing a summary in natural language.

[0423] The present invention uses the following hardware and software:

[0424] OpenCV is used to read the video file and extract frames.

[0425] A pre-trained model using TensorFlow is used to extract features.

[0426] For natural language generation, generative AI models such as GPT-2 are used.

[0427] For example, the prompt for the dance practice video is as follows:

[0428] "This video teaches basic dance steps."

[0429] For recipe videos:

[0430] This video shows the cooking steps.

[0431] By using the above method, the present invention can efficiently summarize the contents of a video and provide it to the user.

[0432] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0433] Step 1:

[0434] A user uses a terminal to provide a video file to the server. The video file is any video data, and the user specifies a file path of their choice. The input here is the path of the video file provided by the user. As an output, the server obtains the specified video file.

[0435] Step 2:

[0436] The server reads the video file and extracts each frame. Each extracted frame is an individual still image, and it uses a library such as OpenCV to extract frames from the video file. The input to this process is the video file obtained in step 1, and the output is the individual extracted frames.

[0437] Step 3:

[0438] The server resizes and normalizes each frame to a predetermined size. Specifically, it resizes each frame to 224x224 pixels and then scales the pixel data to the range of 0 to 1. Data processing involves resizing and normalization. The input is the frame extracted in step 2, and the output is the resized and normalized frame.

[0439] Step 4:

[0440] The server extracts features from the resized and normalized frames. A pre-trained model using TensorFlow is used for feature extraction. Specifically, each frame is input to a specific neural network, which outputs a feature vector. The input is the frame processed in step 3, and the output is a feature vector.

[0441] Step 5:

[0442] The server inputs the feature vectors into a machine learning model to generate a natural language text summary. A video analysis model analyzes the feature vectors and extracts important information. A natural language generation model (e.g., GPT-2) then generates a text summary based on that information. Data calculations involve converting features to text. The input is the feature vector, and the output is the natural language text summary.

[0443] Step 6:

[0444] The server returns the generated natural language text summary to the user. The generated text summary is provided in a format that the user can view through their terminal. The input is the natural language text summary generated in step 5, and the output is the text summary returned to the user's terminal. This allows the user to refer to the provided summary.

[0445] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[0446] The present invention combines a system that analyzes video files provided by users and generates natural language text from them with an emotion engine that recognizes the user's emotions. The present invention will be described below with specific examples.

[0447] First, the user uses the terminal to provide the server with the path to the video file to be analyzed. The video file can be any video data, such as a dance practice video or a presentation practice video.

[0448] The server receives the video file path provided by the user and opens the file. Once the video file is read, the server extracts each frame to analyze its content. Each frame is resized, for example to 224x224 pixels. Each resized frame is then normalized. Normalization is the process of scaling the frame's data to a range of 0 to 1, allowing the machine learning model to process the data more efficiently.

[0449] The normalized frames are extracted as features. These features represent specific information contained in the data as numerical data. The extracted features are used as input to a machine learning model. The server uses these features to generate natural language text using a pre-trained video analysis model and natural language generation model.

[0450] A distinctive feature of the present invention is the emotion engine. The emotion engine analyzes emotions from the user's facial expressions and voice. This analyzed emotion data is fed back to the machine learning model and reflected in the generated natural language text. For example, if the user is nervous, the generated advice is "Relax and try again."

[0451] For example, based on a dance video provided by the user, a response such as "This dance consists of basic hip-hop steps, and it is important to step left and right in time with the rhythm" is generated. Furthermore, if the user shows a happy emotion, positive feedback such as "You're doing great!" is added.

[0452] Finally, the server returns the generated natural language text to the user, who can then review the text and take action, such as improving their activities, based on the text.

[0453] Through this series of processes, the present invention realizes a technology that analyzes non-verbal information and enables responses in natural language that take the user's emotions into account. In particular, by combining a pre-trained model with an emotion engine, more accurate analysis and responses that take emotions into account become possible.

[0454] The processing flow will be explained below.

[0455] Step 1:

[0456] The user uses the terminal to provide the server with the path of the video file to be analyzed, for example, the path of a dance practice video file.

[0457] Step 2:

[0458] The server receives the path to the video file provided by the user and opens the file using the cv2.VideoCapture function.

[0459] Step 3:

[0460] The server loads the video file frame by frame. For each frame, it does the following until it reaches the end of the video:

[0461] Step 4:

[0462] Use the cv2.resize function to resize the frame loaded by the server to a given size, for example 224x224 pixels.

[0463] Step 5:

[0464] The server normalizes the resized frame by dividing each pixel's value by 255, scaling it to the range 0 to 1.

[0465] Step 6:

[0466] The server adds the normalized frame to the list as a feature. This operation is performed for all frames.

[0467] Step 7:

[0468] Once the server has finished processing all frames, it converts the feature list into a NumPy array.

[0469] Step 8:

[0470] The pad_sequences function is used to pad the feature data acquired by the server to a fixed length, thereby preparing the input format for the machine learning model.

[0471] Step 9:

[0472] The server inputs the feature data into a pre-trained video analysis model and extracts the necessary information.

[0473] Step 10:

[0474] The server runs an emotion engine using the user's real-time video and audio data to analyze the user's emotions, for example, by using facial recognition technology or voice analysis technology to identify the user's emotions.

[0475] Step 11:

[0476] The server then feeds the analyzed emotion data back to the machine learning model, which then uses that data to generate natural language text. For example, if the user is nervous, the system generates advice like, "Relax and try again."

[0477] Step 12:

[0478] The server returns the generated natural language text to the user, providing advice and explanations based on the dance content, as well as emotionally sensitive feedback.

[0479] Step 13:

[0480] The user can review the generated natural language text through the device and take action based on it to improve their activity, for example, practicing a dance step again in accordance with the new advice.

[0481] Example 2

[0482] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0483] Conventional video analysis systems were unable to consider the user's emotions when generating text from video content. As a result, the generated text could not provide advice or feedback appropriate to the user's situation, resulting in poor usability. A particular issue was the inability to properly reflect the user's emotions, such as tension or joy.

[0484] The identification process by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes: means for reading a video file provided by a user; means for extracting each frame of the video file, resizing it to a predetermined size, and normalizing it; means for extracting the normalized frames as features; means for inputting the features into a machine learning model and generating natural language text; means for returning the generated natural language text; means including an emotion engine for analyzing the user's facial expressions and voice and recognizing their emotions; and means for feeding back the emotion data to the machine learning model and reflecting it in the natural language text. This makes it possible to generate natural language text that reflects the user's emotions in the analysis results of the video content.

[0485] "User" means an individual or organization that uses the system to provide video files and receive analysis results.

[0486] A "video file" is a digital file in a video data format that a user provides to the system.

[0487] A "frame" refers to one of the still images that make up a video file, and a video is made up of a series of many of these frames.

[0488] "Resizing" is the process of changing the resolution of each frame of a video file to a specific size.

[0489] "Normalization" is the process of scaling the range of data to a certain range, typically to a range from 0 to 1.

[0490] A "feature" is a specific pattern or piece of information in data that has been quantified and is used as input into a machine learning model.

[0491] A "machine learning model" is a pre-trained algorithm or network that analyzes data and makes predictions or classifications.

[0492] "Natural language text" is text in a language format that humans can understand, generated based on the results of video analysis.

[0493] An "emotion engine" is an algorithm or system that analyzes and recognizes emotions from a user's facial expressions, voice, etc.

[0494] "Feedback" is the process of inputting analysis results and data back into the system and reflecting the results.

[0495] MODE FOR CARRYING OUT THE INVENTION

[0496] The present invention combines a system that analyzes video files provided by users and generates natural language text from them with an emotion engine that recognizes the user's emotions. The present invention will be described below with specific examples.

[0497] First, the user uses their device to provide the server with the path to the video file to be analyzed. The video file can be any video data, such as a dance practice video or a presentation practice video. The server receives the path to the video file provided by the user and opens the file. Once the video file is read, the server extracts each frame to analyze its content. Each frame is resized, for example, to 224x224 pixels. Each resized frame is then normalized. Normalization is the process of scaling the frame data to a range of 0 to 1, which allows the machine learning model to process the data more efficiently.

[0498] The normalized frames are extracted as features. These features represent specific information contained in the data as numerical data. The server then inputs these features into a pre-trained video analysis model to analyze the content of the video. Based on the results of the video analysis, the server then generates natural language text using a pre-trained natural language generation model.

[0499] A distinctive feature of this system is the emotion engine, which analyzes emotions from the user's facial expressions and voice. This analyzed emotion data is fed back to the machine learning model and reflected in the generated natural language text. As a specific example, based on a dance video provided by the user, a response such as "This dance consists of basic hip-hop steps, and it is important to step left and right in time with the rhythm" is generated. Furthermore, if the user shows a happy emotion, positive feedback such as "You're doing great!" is added.

[0500] Finally, the server returns the generated natural language text to the user, who can review the text through their device and take action based on it, such as improving their activities.

[0501] Hardware and software used

[0502] This system consists of a server, a terminal, and related software components. Specifically, the following hardware and software are used:

[0503] Hardware:

[0504] Server: A high-performance computer that analyzes video files and generates natural language text.

[0505] Device: The device (PC, tablet, smartphone, etc.) where the user uploads the video file and reviews the generated text.

[0506] software:

[0507] Video analysis libraries: Libraries such as OpenCV for extracting, resizing, and normalizing frames from video files.

[0508] Machine learning frameworks, such as TensorFlow and PyTorch, are used to implement and run video analysis and natural language generation models.

[0509] Emotion analysis engine: A dedicated library or API for analyzing a user's facial expressions and voice.

[0510] Prompt Sentence Examples

[0511] "What happens if a user uploads a dance practice video?"

[0512] In response to this prompt, the system will generate an explanation like this:

[0513] "When a user uploads a dance practice video, the server analyzes the video, extracts each frame, and resizes and normalizes it. It then extracts features and generates natural language text based on them using a pre-trained model. If the user's sentiment is positive, it will add feedback like, 'You're doing great!'"

[0514] Through this series of processes, the present invention realizes a technology that analyzes non-verbal information and generates a response in natural language that takes into account the user's emotions.

[0515] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0516] Step 1:

[0517] The user inputs the path of the video file to be analyzed using the device. The input video file path is, for example, "C:\Users\SampleUser\Videos\dance_practice.mp4". This is the input.

[0518] Step 2:

[0519] The device sends the path of the input video file to the server. Specifically, it sends the path information using an HTTP request or other communication method. The output is the path information received by the server.

[0520] Step 3:

[0521] Opens the specified video file based on the video file path received by the server. Here, the input is the video file path and the output is the video file in its open state.

[0522] Step 4:

[0523] The server extracts each frame of the video file. Specifically, it uses a library such as OpenCV to obtain image data for each frame. The input is the video file, and the output is the extracted frames.

[0524] Step 5:

[0525] The server resizes each extracted frame to 224x224 pixels, for example using the OpenCV cv2.resize function. The input is the image data of the frame, and the output is the resized frames.

[0526] Step 6:

[0527] The server normalizes each resized frame, scaling each pixel value to the range 0 to 1 by dividing it by 255. The input is a set of resized frames, and the output is a set of normalized frames.

[0528] Step 7:

[0529] The server extracts features from the normalized frames. For example, it uses a convolutional neural network (CNN) to extract specific features from the frames as numerical data. The input is the normalized frames, and the output is the feature data.

[0530] Step 8:

[0531] The server analyzes the features using a pre-trained video analysis model. To obtain the analysis results, the feature data is input into the model. The input is the feature data, and the output is the analysis results.

[0532] Step 9:

[0533] The server generates text based on the analysis results using a pre-trained natural language generation model. The input is the analysis results, and the output is the generated natural language text.

[0534] Step 10:

[0535] The server uses an emotion engine to analyze the user's facial expressions and voice to obtain emotion data. The input is the user's facial expression image and voice data, and the output is emotion data.

[0536] Step 11:

[0537] The server feeds the emotion data back to the machine learning model and reflects it in the natural language text. Specifically, it adds advice and feedback according to the emotion to the text. The input is the emotion data and the initial natural language text, and the output is the final natural language text that takes emotion into account.

[0538] Step 12:

[0539] The server returns the generated natural language text to the user. Specifically, it returns the text using an HTTP response, etc. The input is the final natural language text, and the output is the text displayed on the user's terminal.

[0540] Step 13:

[0541] The user views the generated natural language text through a terminal. The input is the returned text, and the output is the user's understanding and action based on it.

[0542] The above are the processing steps of the present invention.

[0543] (Application example 2)

[0544] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0545] Conventional video analysis systems simply analyze video data and generate text without considering the user's emotions. As a result, in fields such as customer service and product recommendations, they are unable to generate responses that appropriately reflect the customer's emotions and reactions, making it difficult to improve customer satisfaction and communicate effectively. The objective of this invention is to provide more effective and personalized responses by analyzing the user's emotions and generating natural language text that reflects them.

[0546] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[0547] In this invention, the server includes means for reading a video file provided by a user, means for extracting each frame of the video file, resizing it to a predetermined size, and normalizing it, means for extracting the normalized frames as features, means for inputting the features into a machine learning model to generate natural language text, means for returning the generated natural language text, means for creating a context for generating appropriate responses and suggestions for customer service, and means for analyzing the user's emotions and reflecting them in the context. This makes it possible to generate natural language text that reflects the user's emotions, thereby enabling more effective and personalized customer service.

[0548] A "user" is an entity that uses this system to provide videos and receive analysis results.

[0549] "Video file" refers to video data provided by the user, and is the media data to be analyzed.

[0550] The term "means" refers to a method or technique for realizing the functions or processes of the present invention.

[0551] A "frame" is an individual still image that makes up a video file.

[0552] "Resizing" is a process of changing the size of an image or frame to a predetermined size.

[0553] "Normalization" is the process of scaling data to a certain range, primarily used to make data easier to handle uniformly in machine learning.

[0554] A "feature" is specific information extracted from data expressed as numerical data.

[0555] A "machine learning model" refers to an algorithm or mathematical model that learns patterns from data and makes predictions and classifications.

[0556] "Natural language text" refers to human-understandable sentences generated from analyzed video data.

[0557] "Context" refers to the background information and circumstances for generating customer service responses and suggestions, and also includes the results of user sentiment analysis.

[0558] "Means for analyzing emotions" refers to a technology or method that detects the user's emotional state from facial expressions, voice, etc., and uses that information within the system.

[0559] "Means for returning" refers to the method or technique for presenting the generated natural language text to the user.

[0560] The present invention is a system that analyzes a video file provided by a user, generates natural language text from the video file, and further combines it with an emotion engine that recognizes the user's emotions. Specific embodiments of the present invention will be described in detail below.

[0561] First, the user uses their device to provide the server with the path to the video file to be analyzed. The video file can be any video data, such as a recorded video of a customer service session. The server receives the path to the video file provided by the user and opens the file. Once the video file is read, the server extracts each frame to analyze its content. Each frame is resized, for example, to 224x224 pixels. Each resized frame is then normalized. Normalization is the process of scaling the frame data to a range of 0 to 1, which allows the machine learning model to process the data more efficiently.

[0562] The normalized frames are extracted as features. These features represent specific information contained in the data as numerical data. The extracted features are used as input to a machine learning model. The server uses these features to generate natural language text using a pre-trained video analysis model and natural language generation model.

[0563] A distinctive feature of the present invention is the emotion engine. The emotion engine analyzes emotions from the user's facial expressions and voice. This analyzed emotion data is fed back to the machine learning model and reflected in the generated natural language text. For example, if the user shows interest, a response such as "This product is new this season and is particularly popular" is generated. If the user loses interest, the response switches to "Can you tell me what other designs you like?"

[0564] Specifically, the server uses OpenCV to extract frames from the video (assuming a video capture camera), resizes the frames to 224x224 pixels, and normalizes them. The normalized data is then input into machine learning models: EmotionEngine (emotion recognition) and TextGenerationModel (natural language generation). The data obtained from sentiment analysis influences the generated text, resulting in natural language text that reflects the user's emotions.

[0565] Examples:

[0566] This application can be used in a scenario where a salesperson in a fashion store is wearing smart glasses and serving customers. When a customer sees an item of clothing that interests them, the smart glasses respond to their emotions and provide appropriate guidance, such as, "This item is new this season and is particularly popular. Please let us know if you need more information." If the customer appears to be losing interest, the smart glasses can switch to a suggestion such as, "Can you tell us what other designs you like?"

[0567] Example prompt:

[0568] "Customer interested" prompt:

[0569] Context: Your customer is interested.

[0570] Frame Data: [frame 1, frame 2, ...]

[0571] Generated text type: Product description

[0572] "Customer is losing interest" prompt:

[0573] Context: Your customer is losing interest.

[0574] Frame Data: [frame 1, frame 2, ...]

[0575] Type of generated text: Questions and Answers

[0576] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0577] Step 1:

[0578] The user uses the device to provide the server with the path to the video file to be analyzed. The path to the video file is specified as input, and the data is sent to the server as output. In this step, the specific actions are to select and send the video file.

[0579] Step 2:

[0580] The server receives the video file path provided by the user and opens the file. The input is the video file path, and the output is the video file itself. This step involves reading the target file from the file system.

[0581] Step 3:

[0582] The server extracts each frame to analyze the video file. The input is the video file, and the output is a list of extracted frames. Specifically, the server uses OpenCV to split the video into frames.

[0583] Step 4:

[0584] Each frame is resized and normalized to 224x224 pixels. The input is the extracted frame, and the output is the resized and normalized frame. Specifically, the resizing and normalization processes are performed.

[0585] Step 5:

[0586] The normalized frames are extracted as features. The input is the normalized frames and the output is the feature data. In this step, a feature extraction algorithm is applied.

[0587] Step 6:

[0588] The server inputs the feature data into a machine learning model and generates natural language text. The input is the feature data, and the output is the generated natural language text. This includes operations that utilize a pre-trained video analysis model and a natural language generation model.

[0589] Step 7:

[0590] The emotion engine analyzes emotions from the user's facial expressions and voice. The input is normalized frame and voice data, and the output is analyzed emotion data. In this step, emotion recognition algorithms are applied.

[0591] Step 8:

[0592] The emotion data is fed back into the machine learning model and reflected in the generated natural language text. The input is emotion data and normal frame data for the generated text, and the output is the final natural language text. Here, the text generation process integrates emotion data.

[0593] Step 9:

[0594] The server returns the generated natural language text to the user. The input is the final natural language text and the output is the text that is displayed on the terminal. This step includes displaying the text in a user interface.

[0595] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0596] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0597] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.

[0598] [Third embodiment]

[0599] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.

[0600] 5, the data processing system 310 includes the data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.

[0601] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0602] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.

[0603] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0604] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0605] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0606] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0607] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0608] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0609] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0610] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."

[0611] The present invention is a system that analyzes a video file provided by a user and generates natural language text based on the video file. The present invention will be described with specific examples.

[0612] First, a user uses a terminal to provide a path to a video file to the server. The video file is any video data, such as a dance practice video.

[0613] The server then reads the provided video file. Once the video file is read, the server extracts each frame to analyze its content. Each frame is a separate still image, which is then resized to a specific size, for example, 224x224 pixels.

[0614] Each resized frame is then normalized, a process that scales the frame's data to the range 0 to 1, allowing machine learning models to process the data more efficiently.

[0615] Next, the normalized frames are extracted as features. These features are numerical representations of specific information contained in the data. The extracted features are used as input for machine learning models.

[0616] The server uses these features to generate natural language text using a pre-trained video analysis model and natural language generation model. First, the video analysis model analyzes the features and extracts the necessary information. Next, that information is input into the natural language generation model, which generates natural language text based on the content of the video.

[0617] For example, based on a dance video provided by the user, a response such as, "This dance consists of basic hip-hop steps, and it's important to step left and right in time with the rhythm" is generated.

[0618] Finally, the server returns the generated natural language text to the user, who can then check and refer to the generated text through their terminal.

[0619] Through this series of processes, the present invention realizes a technology that analyzes non-verbal information and enables responses in natural language. In particular, the use of a pre-trained model enables more accurate analysis and responses.

[0620] The processing flow will be explained below.

[0621] Step 1:

[0622] The user uses the terminal to provide the server with the path of the video file to be analyzed, for example, the path of a dance practice video file.

[0623] Step 2:

[0624] The server receives the path to the video file provided by the user and opens the file using the cv2.VideoCapture function.

[0625] Step 3:

[0626] The server loads the video file frame by frame. For each frame, it does the following until it reaches the end of the video:

[0627] Step 4:

[0628] Use the cv2.resize function to resize the frame loaded by the server to a given size, for example 224x224 pixels.

[0629] Step 5:

[0630] The server normalizes the resized frame by dividing each pixel's value by 255, scaling it to the range 0 to 1.

[0631] Step 6:

[0632] The server adds the normalized frame to the list as a feature. This operation is performed for all frames.

[0633] Step 7:

[0634] Once the server has finished processing all frames, it converts the feature list into a NumPy array.

[0635] Step 8:

[0636] The pad_sequences function is used to pad the feature data acquired by the server to a fixed length, thereby preparing the input format for the machine learning model.

[0637] Step 9:

[0638] The server inputs the feature data into a pre-trained video analysis model and extracts the necessary information.

[0639] Step 10:

[0640] The server generates natural language text based on the extracted information using a natural language generation model.

[0641] Step 11:

[0642] The server returns the generated natural language text to the user, for example, providing advice or explanation based on the dance content.

[0643] Step 12:

[0644] The user can view the natural language text generated through the device and take action based on it, such as improving their activities.

[0645] Example 1

[0646] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[0647] Conventional systems have had difficulty extracting useful information from video files and providing it to users in natural language. Furthermore, due to the low accuracy of video analysis, the generated natural language text was often inaccurate. Furthermore, processing efficiency was low, and real-time performance was lacking.

[0648] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0649] In this invention, the server includes means for reading a video file provided by a user, means for extracting each frame of the video file, resizing it to a predetermined size, and normalizing it, means for extracting the normalized frames as features, means for inputting the features into a machine learning model to generate natural language text, and means for returning the generated natural language text. This makes it possible to efficiently and accurately extract necessary information from a video and generate natural language text based on that information.

[0650] "User" refers to a person who provides a video file to the system.

[0651] "Video file" refers to a file containing various video data that a user provides to a server.

[0652] A "frame" refers to the individual still images that make up a video file.

[0653] "Resizing" refers to the process of changing the size of a frame to a specified dimension.

[0654] "Normalization" refers to the process of scaling data to the range 0 to 1.

[0655] "Features" refer to data that expresses specific information contained in a frame as numerical data.

[0656] A "machine learning model" refers to an algorithm or system designed to perform a specific task by analyzing data and learning.

[0657] "Video analysis model" refers to a machine learning model trained to analyze video data and understand its content.

[0658] "Natural Language Generation Model" refers to a machine learning model trained to generate natural language text from data.

[0659] "Return" refers to the process by which the server provides the generated natural language text to the user.

[0660] The present invention is a system for analyzing a video file provided by a user and generating natural language text based on the video file. Specific embodiments of the present invention will be described below.

[0661] First, the user uses their device to provide the server with the path to a video file. The video file can be any video data, such as a dance practice video. The user enters the path to the video file and sends it to the server. This operation is performed through a form on a web browser or a dedicated application.

[0662] Next, the server retrieves the provided video file and prepares its contents for analysis. Specifically, it uses a video processing library such as FFmpeg to extract each frame of the video file. The extracted frames are individual still images. These frames are then resized to a predetermined size, for example, 224x224 pixels.

[0663] Each resized frame is then normalized, which is the process of scaling a frame's data to the range 0 to 1, allowing the machine learning model to process the data more efficiently.

[0664] The server then extracts features from the normalized frames. These features are numerical representations of specific information contained in the frames. Typically, deep learning models (e.g., ResNet) are used to extract the features.

[0665] The server then inputs the extracted features into a video analysis model and a natural language generation model. The video analysis model is a machine learning model trained to analyze video data and extract specific information, while the natural language generation model is a machine learning model trained to generate natural language text based on that information. For example, generative AI models such as GPT-3 are often used.

[0666] An example of a prompt to be input to a generative AI model is, "Please analyze the following dance practice video and generate natural language text based on its content," followed by information extracted by the video analysis model.

[0667] Finally, the server returns the generated natural language text to the user. The user can check the generated text through their device and use it as reference. For example, based on a dance video provided by the user, a natural language text such as "This dance consists of basic hip-hop steps, and it is important to step left and right in time with the rhythm" may be generated.

[0668] Through this series of processes, the present invention realizes a technology that can analyze non-verbal information such as video and generate natural language responses. In particular, the use of a pre-trained model enables highly accurate analysis and natural language generation.

[0669] The above is a specific embodiment for carrying out the present invention.

[0670] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0671] Step 1:

[0672] The user uses a terminal to provide the server with the path to the video file. The user enters the path to the video file into the input form on the terminal and presses the send button. This operation sends the path to the video file to the server. The input is the video file path specified by the user, and the output is the path sent to the server.

[0673] Step 2:

[0674] The server retrieves the video file and extracts each frame. The server reads the video file using the received video file path and splits the video into frames using a video processing library such as FFmpeg. The input is the video file path, and the output is the image data of each extracted frame.

[0675] Step 3:

[0676] The server resizes and normalizes each frame to a specified size. The server resizes the extracted frames to a specified size (e.g., 224x224 pixels) and normalizes the value of each pixel to a range of 0 to 1. The input is the image data of the extracted frames, and the output is the resized and normalized frame data. This process allows the machine learning model to process the data efficiently.

[0677] Step 4:

[0678] The server extracts features from the normalized frames. The server uses a deep learning model (e.g., ResNet) to extract features from the normalized frames. The input is the normalized frame data, and the output is feature data for each frame. Features are numerical representations of important information contained in the frame.

[0679] Step 5:

[0680] The server inputs the features into a video analysis model and extracts the necessary information. The server inputs the features into a pre-trained video analysis model and extracts information based on the content of the video. The input is feature data, and the output is the information extracted by the video analysis model.

[0681] Step 6:

[0682] The server inputs information from the video analysis model into a natural language generation model to generate natural language text. The server inputs information obtained from the video analysis model into a natural language generation model (e.g., GPT-3) and generates natural language text based on that information. The input is information from the video analysis model, and the output is the generated natural language text.

[0683] Step 7:

[0684] The server returns the generated natural language text to the user. The server sends the generated natural language text to the user's terminal so that the user can check it. The input is the generated natural language text, and the output is the text displayed on the user's terminal. The user can refer to this text.

[0685] (Application example 1)

[0686] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[0687] In recent years, the increase in video content has created a demand for users to quickly understand the content of videos. However, systems that automatically summarize video content and provide it as text are not widely available, meaning users are unable to save time watching videos. Therefore, there is a need for a system that automatically generates video summaries and provides them to users.

[0688] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0689] In this invention, the server includes means for reading a video file provided by a user, means for extracting frames from the video file, resizing the frames to a predetermined size, and normalizing the frames, means for extracting features from the normalized frames, means for inputting the features into a machine learning model to generate a natural language text summary, and means for returning the generated natural language text summary, thereby enabling the content of the video provided by the user to be automatically analyzed and presented as a summary.

[0690] "User" means an individual or legal entity that uses the service or application.

[0691] A "video file" is a digital file that contains video data, and may also contain audio and metadata.

[0692] "Each frame" refers to an individual still image extracted from a video file.

[0693] "Resizing" is the process of changing the width and height of an image (frame) to fit a new size.

[0694] "Normalization" is the process of scaling data to a particular range (typically between 0 and 1).

[0695] "Features" are important numerical data or patterns extracted from data and input into machine learning models.

[0696] A "machine learning model" is an algorithm or model that is trained to analyze data and perform a specific task.

[0697] A "video analysis model" is a machine learning model trained to perform a specific task to extract and analyze features from video data.

[0698] A "natural language generation model" is a machine learning model trained to generate coherent natural language text based on input data.

[0699] A "natural language text summary" is a document that condenses the content of the original video and presents it in an easy-to-understand format.

[0700] "Returning means" refers to the process or mechanism for providing the generated data or information to the user.

[0701] "Padding" is an operation to fill in missing parts to make data a fixed length.

[0702] The present invention is a system that analyzes video files provided by users and generates natural language text summaries based on the files. This system is realized through the interaction of a server, a terminal, and a user.

[0703] First, a user uses a terminal to provide a video file to the server. The provided video file can be any video data, such as a dance practice video or a cooking recipe video.

[0704] The server reads the received video file. Once the video file is read, the server extracts each frame to analyze its content. Each frame is an individual still image, which is then resized to a predetermined size (e.g., 224x224 pixels). Each resized frame is then normalized, scaling the data range from 0 to 1.

[0705] Next, the normalized frames are extracted as features. Features are numerical representations of important information contained in the video frames. Based on these features, the server generates data to be input into the machine learning model. This machine learning model consists of a pre-trained video analysis model and a natural language generation model.

[0706] The video analysis model analyzes the feature data and extracts important information. This information is input into the natural language generation model, which generates a natural language text summary based on the content of the video. For example, based on a dance practice video, the summary generated might be, "This dance consists of basic hip-hop steps. It is important to step left and right in time with the rhythm."

[0707] The server then returns the generated natural language text summary to the user, who can then review the generated text and use it as a reference. Through this process, the present invention realizes a technology for analyzing non-verbal information and providing a summary in natural language.

[0708] The present invention uses the following hardware and software:

[0709] OpenCV is used to read the video file and extract frames.

[0710] A pre-trained model using TensorFlow is used to extract features.

[0711] For natural language generation, generative AI models such as GPT-2 are used.

[0712] For example, the prompt for the dance practice video is as follows:

[0713] "This video teaches basic dance steps."

[0714] For recipe videos:

[0715] This video shows the cooking steps.

[0716] By using the above method, the present invention can efficiently summarize the contents of a video and provide it to the user.

[0717] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0718] Step 1:

[0719] A user uses a terminal to provide a video file to the server. The video file is any video data, and the user specifies a file path of their choice. The input here is the path of the video file provided by the user. As an output, the server obtains the specified video file.

[0720] Step 2:

[0721] The server reads the video file and extracts each frame. Each extracted frame is an individual still image, and it uses a library such as OpenCV to extract frames from the video file. The input to this process is the video file obtained in step 1, and the output is the individual extracted frames.

[0722] Step 3:

[0723] The server resizes and normalizes each frame to a predetermined size. Specifically, it resizes each frame to 224x224 pixels and then scales the pixel data to the range of 0 to 1. Data processing involves resizing and normalization. The input is the frame extracted in step 2, and the output is the resized and normalized frame.

[0724] Step 4:

[0725] The server extracts features from the resized and normalized frames. A pre-trained model using TensorFlow is used for feature extraction. Specifically, each frame is input to a specific neural network, which outputs a feature vector. The input is the frame processed in step 3, and the output is a feature vector.

[0726] Step 5:

[0727] The server inputs the feature vectors into a machine learning model to generate a natural language text summary. A video analysis model analyzes the feature vectors and extracts important information. A natural language generation model (e.g., GPT-2) then generates a text summary based on that information. Data calculations involve converting features to text. The input is the feature vector, and the output is the natural language text summary.

[0728] Step 6:

[0729] The server returns the generated natural language text summary to the user. The generated text summary is provided in a format that the user can view through their terminal. The input is the natural language text summary generated in step 5, and the output is the text summary returned to the user's terminal. This allows the user to refer to the provided summary.

[0730] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[0731] The present invention combines a system that analyzes video files provided by users and generates natural language text from them with an emotion engine that recognizes the user's emotions. The present invention will be described below with specific examples.

[0732] First, the user uses the terminal to provide the server with the path to the video file to be analyzed. The video file can be any video data, such as a dance practice video or a presentation practice video.

[0733] The server receives the video file path provided by the user and opens the file. Once the video file is read, the server extracts each frame to analyze its content. Each frame is resized, for example to 224x224 pixels. Each resized frame is then normalized. Normalization is the process of scaling the frame's data to a range of 0 to 1, allowing the machine learning model to process the data more efficiently.

[0734] The normalized frames are extracted as features. These features represent specific information contained in the data as numerical data. The extracted features are used as input to a machine learning model. The server uses these features to generate natural language text using a pre-trained video analysis model and natural language generation model.

[0735] A distinctive feature of the present invention is the emotion engine. The emotion engine analyzes emotions from the user's facial expressions and voice. This analyzed emotion data is fed back to the machine learning model and reflected in the generated natural language text. For example, if the user is nervous, the generated advice is "Relax and try again."

[0736] For example, based on a dance video provided by the user, a response such as "This dance consists of basic hip-hop steps, and it is important to step left and right in time with the rhythm" is generated. Furthermore, if the user shows a happy emotion, positive feedback such as "You're doing great!" is added.

[0737] Finally, the server returns the generated natural language text to the user, who can then review the text and take action, such as improving their activities, based on the text.

[0738] Through this series of processes, the present invention realizes a technology that analyzes non-verbal information and enables responses in natural language that take the user's emotions into account. In particular, by combining a pre-trained model with an emotion engine, more accurate analysis and responses that take emotions into account become possible.

[0739] The processing flow will be explained below.

[0740] Step 1:

[0741] The user uses the terminal to provide the server with the path of the video file to be analyzed, for example, the path of a dance practice video file.

[0742] Step 2:

[0743] The server receives the path to the video file provided by the user and opens the file using the cv2.VideoCapture function.

[0744] Step 3:

[0745] The server loads the video file frame by frame. For each frame, it does the following until it reaches the end of the video:

[0746] Step 4:

[0747] Use the cv2.resize function to resize the frame loaded by the server to a given size, for example 224x224 pixels.

[0748] Step 5:

[0749] The server normalizes the resized frame by dividing each pixel's value by 255, scaling it to the range 0 to 1.

[0750] Step 6:

[0751] The server adds the normalized frame to the list as a feature. This operation is performed for all frames.

[0752] Step 7:

[0753] Once the server has finished processing all frames, it converts the feature list into a NumPy array.

[0754] Step 8:

[0755] The pad_sequences function is used to pad the feature data acquired by the server to a fixed length, thereby preparing the input format for the machine learning model.

[0756] Step 9:

[0757] The server inputs the feature data into a pre-trained video analysis model and extracts the necessary information.

[0758] Step 10:

[0759] The server runs an emotion engine using the user's real-time video and audio data to analyze the user's emotions, for example, by using facial recognition technology or voice analysis technology to identify the user's emotions.

[0760] Step 11:

[0761] The server then feeds the analyzed emotion data back to the machine learning model, which then uses that data to generate natural language text. For example, if the user is nervous, the system generates advice like, "Relax and try again."

[0762] Step 12:

[0763] The server returns the generated natural language text to the user, providing advice and explanations based on the dance content, as well as emotionally sensitive feedback.

[0764] Step 13:

[0765] The user can review the generated natural language text through the device and take action based on it to improve their activity, for example, practicing a dance step again in accordance with the new advice.

[0766] Example 2

[0767] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[0768] Conventional video analysis systems were unable to consider the user's emotions when generating text from video content. As a result, the generated text could not provide advice or feedback appropriate to the user's situation, resulting in poor usability. A particular issue was the inability to properly reflect the user's emotions, such as tension or joy.

[0769] The identification process by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes: means for reading a video file provided by a user; means for extracting each frame of the video file, resizing it to a predetermined size, and normalizing it; means for extracting the normalized frames as features; means for inputting the features into a machine learning model and generating natural language text; means for returning the generated natural language text; means including an emotion engine for analyzing the user's facial expressions and voice and recognizing their emotions; and means for feeding back the emotion data to the machine learning model and reflecting it in the natural language text. This makes it possible to generate natural language text that reflects the user's emotions in the analysis results of the video content.

[0770] "User" means an individual or organization that uses the system to provide video files and receive analysis results.

[0771] A "video file" is a digital file in a video data format that a user provides to the system.

[0772] A "frame" refers to one of the still images that make up a video file, and a video is made up of a series of many of these frames.

[0773] "Resizing" is the process of changing the resolution of each frame of a video file to a specific size.

[0774] "Normalization" is the process of scaling the range of data to a certain range, typically to a range from 0 to 1.

[0775] A "feature" is a specific pattern or piece of information in data that has been quantified and is used as input into a machine learning model.

[0776] A "machine learning model" is a pre-trained algorithm or network that analyzes data and makes predictions or classifications.

[0777] "Natural language text" is text in a language format that humans can understand, generated based on the results of video analysis.

[0778] An "emotion engine" is an algorithm or system that analyzes and recognizes emotions from a user's facial expressions, voice, etc.

[0779] "Feedback" is the process of inputting analysis results and data back into the system and reflecting the results.

[0780] MODE FOR CARRYING OUT THE INVENTION

[0781] The present invention combines a system that analyzes video files provided by users and generates natural language text from them with an emotion engine that recognizes the user's emotions. The present invention will be described below with specific examples.

[0782] First, the user uses their device to provide the server with the path to the video file to be analyzed. The video file can be any video data, such as a dance practice video or a presentation practice video. The server receives the path to the video file provided by the user and opens the file. Once the video file is read, the server extracts each frame to analyze its content. Each frame is resized, for example, to 224x224 pixels. Each resized frame is then normalized. Normalization is the process of scaling the frame data to a range of 0 to 1, which allows the machine learning model to process the data more efficiently.

[0783] The normalized frames are extracted as features. These features represent specific information contained in the data as numerical data. The server then inputs these features into a pre-trained video analysis model to analyze the content of the video. Based on the results of the video analysis, the server then generates natural language text using a pre-trained natural language generation model.

[0784] A distinctive feature of this system is the emotion engine, which analyzes emotions from the user's facial expressions and voice. This analyzed emotion data is fed back to the machine learning model and reflected in the generated natural language text. As a specific example, based on a dance video provided by the user, a response such as "This dance consists of basic hip-hop steps, and it is important to step left and right in time with the rhythm" is generated. Furthermore, if the user shows a happy emotion, positive feedback such as "You're doing great!" is added.

[0785] Finally, the server returns the generated natural language text to the user, who can review the text through their device and take action based on it, such as improving their activities.

[0786] Hardware and software used

[0787] This system consists of a server, a terminal, and related software components. Specifically, the following hardware and software are used:

[0788] Hardware:

[0789] Server: A high-performance computer that analyzes video files and generates natural language text.

[0790] Device: The device (PC, tablet, smartphone, etc.) where the user uploads the video file and reviews the generated text.

[0791] software:

[0792] Video analysis libraries: Libraries such as OpenCV for extracting, resizing, and normalizing frames from video files.

[0793] Machine learning frameworks, such as TensorFlow and PyTorch, are used to implement and run video analysis and natural language generation models.

[0794] Emotion analysis engine: A dedicated library or API for analyzing a user's facial expressions and voice.

[0795] Prompt Sentence Examples

[0796] "What happens if a user uploads a dance practice video?"

[0797] In response to this prompt, the system will generate an explanation like this:

[0798] "When a user uploads a dance practice video, the server analyzes the video, extracts each frame, and resizes and normalizes it. It then extracts features and generates natural language text based on them using a pre-trained model. If the user's sentiment is positive, it will add feedback like, 'You're doing great!'"

[0799] Through this series of processes, the present invention realizes a technology that analyzes non-verbal information and generates a response in natural language that takes into account the user's emotions.

[0800] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0801] Step 1:

[0802] The user inputs the path of the video file to be analyzed using the device. The input video file path is, for example, "C:\Users\SampleUser\Videos\dance_practice.mp4". This is the input.

[0803] Step 2:

[0804] The device sends the path of the input video file to the server. Specifically, it sends the path information using an HTTP request or other communication method. The output is the path information received by the server.

[0805] Step 3:

[0806] Opens the specified video file based on the video file path received by the server. Here, the input is the video file path and the output is the video file in its open state.

[0807] Step 4:

[0808] The server extracts each frame of the video file. Specifically, it uses a library such as OpenCV to obtain image data for each frame. The input is the video file, and the output is the extracted frames.

[0809] Step 5:

[0810] The server resizes each extracted frame to 224x224 pixels, for example using the OpenCV cv2.resize function. The input is the image data of the frame, and the output is the resized frames.

[0811] Step 6:

[0812] The server normalizes each resized frame, scaling each pixel value to the range 0 to 1 by dividing it by 255. The input is a set of resized frames, and the output is a set of normalized frames.

[0813] Step 7:

[0814] The server extracts features from the normalized frames. For example, it uses a convolutional neural network (CNN) to extract specific features from the frames as numerical data. The input is the normalized frames, and the output is the feature data.

[0815] Step 8:

[0816] The server analyzes the features using a pre-trained video analysis model. To obtain the analysis results, the feature data is input into the model. The input is the feature data, and the output is the analysis results.

[0817] Step 9:

[0818] The server generates text based on the analysis results using a pre-trained natural language generation model. The input is the analysis results, and the output is the generated natural language text.

[0819] Step 10:

[0820] The server uses an emotion engine to analyze the user's facial expressions and voice to obtain emotion data. The input is the user's facial expression image and voice data, and the output is emotion data.

[0821] Step 11:

[0822] The server feeds the emotion data back to the machine learning model and reflects it in the natural language text. Specifically, it adds advice and feedback according to the emotion to the text. The input is the emotion data and the initial natural language text, and the output is the final natural language text that takes emotion into account.

[0823] Step 12:

[0824] The server returns the generated natural language text to the user. Specifically, it returns the text using an HTTP response, etc. The input is the final natural language text, and the output is the text displayed on the user's terminal.

[0825] Step 13:

[0826] The user views the generated natural language text through a terminal. The input is the returned text, and the output is the user's understanding and action based on it.

[0827] The above are the processing steps of the present invention.

[0828] (Application example 2)

[0829] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[0830] Conventional video analysis systems simply analyze video data and generate text without considering the user's emotions. As a result, in fields such as customer service and product recommendations, they are unable to generate responses that appropriately reflect the customer's emotions and reactions, making it difficult to improve customer satisfaction and communicate effectively. The objective of this invention is to provide more effective and personalized responses by analyzing the user's emotions and generating natural language text that reflects them.

[0831] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[0832] In this invention, the server includes means for reading a video file provided by a user, means for extracting each frame of the video file, resizing it to a predetermined size, and normalizing it, means for extracting the normalized frames as features, means for inputting the features into a machine learning model to generate natural language text, means for returning the generated natural language text, means for creating a context for generating appropriate responses and suggestions for customer service, and means for analyzing the user's emotions and reflecting them in the context. This makes it possible to generate natural language text that reflects the user's emotions, thereby enabling more effective and personalized customer service.

[0833] A "user" is an entity that uses this system to provide videos and receive analysis results.

[0834] "Video file" refers to video data provided by the user, and is the media data to be analyzed.

[0835] The term "means" refers to a method or technique for realizing the functions or processes of the present invention.

[0836] A "frame" is an individual still image that makes up a video file.

[0837] "Resizing" is a process of changing the size of an image or frame to a predetermined size.

[0838] "Normalization" is the process of scaling data to a certain range, primarily used to make data easier to handle uniformly in machine learning.

[0839] A "feature" is specific information extracted from data expressed as numerical data.

[0840] A "machine learning model" refers to an algorithm or mathematical model that learns patterns from data and makes predictions and classifications.

[0841] "Natural language text" refers to human-understandable sentences generated from analyzed video data.

[0842] "Context" refers to the background information and circumstances for generating customer service responses and suggestions, and also includes the results of user sentiment analysis.

[0843] "Means for analyzing emotions" refers to a technology or method that detects the user's emotional state from facial expressions, voice, etc., and uses that information within the system.

[0844] "Means for returning" refers to the method or technique for presenting the generated natural language text to the user.

[0845] The present invention is a system that analyzes a video file provided by a user, generates natural language text from the video file, and further combines it with an emotion engine that recognizes the user's emotions. Specific embodiments of the present invention will be described in detail below.

[0846] First, the user uses their device to provide the server with the path to the video file to be analyzed. The video file can be any video data, such as a recorded video of a customer service session. The server receives the path to the video file provided by the user and opens the file. Once the video file is read, the server extracts each frame to analyze its content. Each frame is resized, for example, to 224x224 pixels. Each resized frame is then normalized. Normalization is the process of scaling the frame data to a range of 0 to 1, which allows the machine learning model to process the data more efficiently.

[0847] The normalized frames are extracted as features. These features represent specific information contained in the data as numerical data. The extracted features are used as input to a machine learning model. The server uses these features to generate natural language text using a pre-trained video analysis model and natural language generation model.

[0848] A distinctive feature of the present invention is the emotion engine. The emotion engine analyzes emotions from the user's facial expressions and voice. This analyzed emotion data is fed back to the machine learning model and reflected in the generated natural language text. For example, if the user shows interest, a response such as "This product is new this season and is particularly popular" is generated. If the user loses interest, the response switches to "Can you tell me what other designs you like?"

[0849] Specifically, the server uses OpenCV to extract frames from the video (assuming a video capture camera), resizes the frames to 224x224 pixels, and normalizes them. The normalized data is then input into machine learning models: EmotionEngine (emotion recognition) and TextGenerationModel (natural language generation). The data obtained from sentiment analysis influences the generated text, resulting in natural language text that reflects the user's emotions.

[0850] Examples:

[0851] This application can be used in a scenario where a salesperson in a fashion store is wearing smart glasses and serving customers. When a customer sees an item of clothing that interests them, the smart glasses respond to their emotions and provide appropriate guidance, such as, "This item is new this season and is particularly popular. Please let us know if you need more information." If the customer appears to be losing interest, the smart glasses can switch to a suggestion such as, "Can you tell us what other designs you like?"

[0852] Example prompt:

[0853] "Customer interested" prompt:

[0854] Context: Your customer is interested.

[0855] Frame Data: [frame 1, frame 2, ...]

[0856] Generated text type: Product description

[0857] "Customer is losing interest" prompt:

[0858] Context: Your customer is losing interest.

[0859] Frame Data: [frame 1, frame 2, ...]

[0860] Type of generated text: Questions and Answers

[0861] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0862] Step 1:

[0863] The user uses the device to provide the server with the path to the video file to be analyzed. The path to the video file is specified as input, and the data is sent to the server as output. In this step, the specific actions are to select and send the video file.

[0864] Step 2:

[0865] The server receives the video file path provided by the user and opens the file. The input is the video file path, and the output is the video file itself. This step involves reading the target file from the file system.

[0866] Step 3:

[0867] The server extracts each frame to analyze the video file. The input is the video file, and the output is a list of extracted frames. Specifically, the server uses OpenCV to split the video into frames.

[0868] Step 4:

[0869] Each frame is resized and normalized to 224x224 pixels. The input is the extracted frame, and the output is the resized and normalized frame. Specifically, the resizing and normalization processes are performed.

[0870] Step 5:

[0871] The normalized frames are extracted as features. The input is the normalized frames and the output is the feature data. In this step, a feature extraction algorithm is applied.

[0872] Step 6:

[0873] The server inputs the feature data into a machine learning model and generates natural language text. The input is the feature data, and the output is the generated natural language text. This includes operations that utilize a pre-trained video analysis model and a natural language generation model.

[0874] Step 7:

[0875] The emotion engine analyzes emotions from the user's facial expressions and voice. The input is normalized frame and voice data, and the output is analyzed emotion data. In this step, emotion recognition algorithms are applied.

[0876] Step 8:

[0877] The emotion data is fed back into the machine learning model and reflected in the generated natural language text. The input is emotion data and normal frame data for the generated text, and the output is the final natural language text. Here, the text generation process integrates emotion data.

[0878] Step 9:

[0879] The server returns the generated natural language text to the user. The input is the final natural language text and the output is the text that is displayed on the terminal. This step includes displaying the text in a user interface.

[0880] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0881] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0882] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.

[0883] [Fourth embodiment]

[0884] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.

[0885] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[0886] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0887] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.

[0888] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0889] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0890] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0891] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.

[0892] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0893] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0894] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0895] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0896] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[0897] The present invention is a system that analyzes a video file provided by a user and generates natural language text based on the video file. The present invention will be described with specific examples.

[0898] First, a user uses a terminal to provide a path to a video file to the server. The video file is any video data, such as a dance practice video.

[0899] The server then reads the provided video file. Once the video file is read, the server extracts each frame to analyze its content. Each frame is a separate still image, which is then resized to a specific size, for example, 224x224 pixels.

[0900] Each resized frame is then normalized, a process that scales the frame's data to the range 0 to 1, allowing machine learning models to process the data more efficiently.

[0901] Next, the normalized frames are extracted as features. These features are numerical representations of specific information contained in the data. The extracted features are used as input for machine learning models.

[0902] The server uses these features to generate natural language text using a pre-trained video analysis model and natural language generation model. First, the video analysis model analyzes the features and extracts the necessary information. Next, that information is input into the natural language generation model, which generates natural language text based on the content of the video.

[0903] For example, based on a dance video provided by the user, a response such as, "This dance consists of basic hip-hop steps, and it's important to step left and right in time with the rhythm" is generated.

[0904] Finally, the server returns the generated natural language text to the user, who can then check and refer to the generated text through their terminal.

[0905] Through this series of processes, the present invention realizes a technology that analyzes non-verbal information and enables responses in natural language. In particular, the use of a pre-trained model enables more accurate analysis and responses.

[0906] The processing flow will be explained below.

[0907] Step 1:

[0908] The user uses the terminal to provide the server with the path of the video file to be analyzed, for example, the path of a dance practice video file.

[0909] Step 2:

[0910] The server receives the path to the video file provided by the user and opens the file using the cv2.VideoCapture function.

[0911] Step 3:

[0912] The server loads the video file frame by frame. For each frame, it does the following until it reaches the end of the video:

[0913] Step 4:

[0914] Use the cv2.resize function to resize the frame loaded by the server to a given size, for example 224x224 pixels.

[0915] Step 5:

[0916] The server normalizes the resized frame by dividing each pixel's value by 255, scaling it to the range 0 to 1.

[0917] Step 6:

[0918] The server adds the normalized frame to the list as a feature. This operation is performed for all frames.

[0919] Step 7:

[0920] Once the server has finished processing all frames, it converts the feature list into a NumPy array.

[0921] Step 8:

[0922] The pad_sequences function is used to pad the feature data acquired by the server to a fixed length, thereby preparing the input format for the machine learning model.

[0923] Step 9:

[0924] The server inputs the feature data into a pre-trained video analysis model and extracts the necessary information.

[0925] Step 10:

[0926] The server generates natural language text based on the extracted information using a natural language generation model.

[0927] Step 11:

[0928] The server returns the generated natural language text to the user, for example, providing advice or explanation based on the dance content.

[0929] Step 12:

[0930] The user can view the natural language text generated through the device and take action based on it, such as improving their activities.

[0931] Example 1

[0932] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[0933] Conventional systems have had difficulty extracting useful information from video files and providing it to users in natural language. Furthermore, due to the low accuracy of video analysis, the generated natural language text was often inaccurate. Furthermore, processing efficiency was low, and real-time performance was lacking.

[0934] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0935] In this invention, the server includes means for reading a video file provided by a user, means for extracting each frame of the video file, resizing it to a predetermined size, and normalizing it, means for extracting the normalized frames as features, means for inputting the features into a machine learning model to generate natural language text, and means for returning the generated natural language text. This makes it possible to efficiently and accurately extract necessary information from a video and generate natural language text based on that information.

[0936] "User" refers to a person who provides a video file to the system.

[0937] "Video file" refers to a file containing various video data that a user provides to a server.

[0938] A "frame" refers to the individual still images that make up a video file.

[0939] "Resizing" refers to the process of changing the size of a frame to a specified dimension.

[0940] "Normalization" refers to the process of scaling data to the range 0 to 1.

[0941] "Features" refer to data that expresses specific information contained in a frame as numerical data.

[0942] A "machine learning model" refers to an algorithm or system designed to perform a specific task by analyzing data and learning.

[0943] "Video analysis model" refers to a machine learning model trained to analyze video data and understand its content.

[0944] "Natural Language Generation Model" refers to a machine learning model trained to generate natural language text from data.

[0945] "Return" refers to the process by which the server provides the generated natural language text to the user.

[0946] The present invention is a system for analyzing a video file provided by a user and generating natural language text based on the video file. Specific embodiments of the present invention will be described below.

[0947] First, the user uses their device to provide the server with the path to a video file. The video file can be any video data, such as a dance practice video. The user enters the path to the video file and sends it to the server. This operation is performed through a form on a web browser or a dedicated application.

[0948] Next, the server retrieves the provided video file and prepares its contents for analysis. Specifically, it uses a video processing library such as FFmpeg to extract each frame of the video file. The extracted frames are individual still images. These frames are then resized to a predetermined size, for example, 224x224 pixels.

[0949] Each resized frame is then normalized, which is the process of scaling a frame's data to the range 0 to 1, allowing the machine learning model to process the data more efficiently.

[0950] The server then extracts features from the normalized frames. These features are numerical representations of specific information contained in the frames. Typically, deep learning models (e.g., ResNet) are used to extract the features.

[0951] The server then inputs the extracted features into a video analysis model and a natural language generation model. The video analysis model is a machine learning model trained to analyze video data and extract specific information, while the natural language generation model is a machine learning model trained to generate natural language text based on that information. For example, generative AI models such as GPT-3 are often used.

[0952] An example of a prompt to be input to a generative AI model is, "Please analyze the following dance practice video and generate natural language text based on its content," followed by information extracted by the video analysis model.

[0953] Finally, the server returns the generated natural language text to the user. The user can check the generated text through their device and use it as reference. For example, based on a dance video provided by the user, a natural language text such as "This dance consists of basic hip-hop steps, and it is important to step left and right in time with the rhythm" may be generated.

[0954] Through this series of processes, the present invention realizes a technology that can analyze non-verbal information such as video and generate natural language responses. In particular, the use of a pre-trained model enables highly accurate analysis and natural language generation.

[0955] The above is a specific embodiment for carrying out the present invention.

[0956] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0957] Step 1:

[0958] The user uses a terminal to provide the server with the path to the video file. The user enters the path to the video file into the input form on the terminal and presses the send button. This operation sends the path to the video file to the server. The input is the video file path specified by the user, and the output is the path sent to the server.

[0959] Step 2:

[0960] The server retrieves the video file and extracts each frame. The server reads the video file using the received video file path and splits the video into frames using a video processing library such as FFmpeg. The input is the video file path, and the output is the image data of each extracted frame.

[0961] Step 3:

[0962] The server resizes and normalizes each frame to a specified size. The server resizes the extracted frames to a specified size (e.g., 224x224 pixels) and normalizes the value of each pixel to a range of 0 to 1. The input is the image data of the extracted frames, and the output is the resized and normalized frame data. This process allows the machine learning model to process the data efficiently.

[0963] Step 4:

[0964] The server extracts features from the normalized frames. The server uses a deep learning model (e.g., ResNet) to extract features from the normalized frames. The input is the normalized frame data, and the output is feature data for each frame. Features are numerical representations of important information contained in the frame.

[0965] Step 5:

[0966] The server inputs the features into a video analysis model and extracts the necessary information. The server inputs the features into a pre-trained video analysis model and extracts information based on the content of the video. The input is feature data, and the output is the information extracted by the video analysis model.

[0967] Step 6:

[0968] The server inputs information from the video analysis model into a natural language generation model to generate natural language text. The server inputs information obtained from the video analysis model into a natural language generation model (e.g., GPT-3) and generates natural language text based on that information. The input is information from the video analysis model, and the output is the generated natural language text.

[0969] Step 7:

[0970] The server returns the generated natural language text to the user. The server sends the generated natural language text to the user's terminal so that the user can check it. The input is the generated natural language text, and the output is the text displayed on the user's terminal. The user can refer to this text.

[0971] (Application example 1)

[0972] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[0973] In recent years, the increase in video content has created a demand for users to quickly understand the content of videos. However, systems that automatically summarize video content and provide it as text are not widely available, meaning users are unable to save time watching videos. Therefore, there is a need for a system that automatically generates video summaries and provides them to users.

[0974] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0975] In this invention, the server includes means for reading a video file provided by a user, means for extracting frames from the video file, resizing the frames to a predetermined size, and normalizing the frames, means for extracting features from the normalized frames, means for inputting the features into a machine learning model to generate a natural language text summary, and means for returning the generated natural language text summary, thereby enabling the content of the video provided by the user to be automatically analyzed and presented as a summary.

[0976] "User" means an individual or legal entity that uses the service or application.

[0977] A "video file" is a digital file that contains video data, and may also contain audio and metadata.

[0978] "Each frame" refers to an individual still image extracted from a video file.

[0979] "Resizing" is the process of changing the width and height of an image (frame) to fit a new size.

[0980] "Normalization" is the process of scaling data to a particular range (typically between 0 and 1).

[0981] "Features" are important numerical data or patterns extracted from data and input into machine learning models.

[0982] A "machine learning model" is an algorithm or model that is trained to analyze data and perform a specific task.

[0983] A "video analysis model" is a machine learning model trained to perform a specific task to extract and analyze features from video data.

[0984] A "natural language generation model" is a machine learning model trained to generate coherent natural language text based on input data.

[0985] A "natural language text summary" is a document that condenses the content of the original video and presents it in an easy-to-understand format.

[0986] "Returning means" refers to the process or mechanism for providing the generated data or information to the user.

[0987] "Padding" is an operation to fill in missing parts to make data a fixed length.

[0988] The present invention is a system that analyzes video files provided by users and generates natural language text summaries based on the files. This system is realized through the interaction of a server, a terminal, and a user.

[0989] First, a user uses a terminal to provide a video file to the server. The provided video file can be any video data, such as a dance practice video or a cooking recipe video.

[0990] The server reads the received video file. Once the video file is read, the server extracts each frame to analyze its content. Each frame is an individual still image, which is then resized to a predetermined size (e.g., 224x224 pixels). Each resized frame is then normalized, scaling the data range from 0 to 1.

[0991] Next, the normalized frames are extracted as features. Features are numerical representations of important information contained in the video frames. Based on these features, the server generates data to be input into the machine learning model. This machine learning model consists of a pre-trained video analysis model and a natural language generation model.

[0992] The video analysis model analyzes the feature data and extracts important information. This information is input into the natural language generation model, which generates a natural language text summary based on the content of the video. For example, based on a dance practice video, the summary generated might be, "This dance consists of basic hip-hop steps. It is important to step left and right in time with the rhythm."

[0993] The server then returns the generated natural language text summary to the user, who can then review the generated text and use it as a reference. Through this process, the present invention realizes a technology for analyzing non-verbal information and providing a summary in natural language.

[0994] The present invention uses the following hardware and software:

[0995] OpenCV is used to read the video file and extract frames.

[0996] A pre-trained model using TensorFlow is used to extract features.

[0997] For natural language generation, generative AI models such as GPT-2 are used.

[0998] For example, the prompt for the dance practice video is as follows:

[0999] "This video teaches basic dance steps."

[1000] For recipe videos:

[1001] This video shows the cooking steps.

[1002] By using the above method, the present invention can efficiently summarize the contents of a video and provide it to the user.

[1003] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[1004] Step 1:

[1005] A user uses a terminal to provide a video file to the server. The video file is any video data, and the user specifies a file path of their choice. The input here is the path of the video file provided by the user. As an output, the server obtains the specified video file.

[1006] Step 2:

[1007] The server reads the video file and extracts each frame. Each extracted frame is an individual still image, and it uses a library such as OpenCV to extract frames from the video file. The input to this process is the video file obtained in step 1, and the output is the individual extracted frames.

[1008] Step 3:

[1009] The server resizes and normalizes each frame to a predetermined size. Specifically, it resizes each frame to 224x224 pixels and then scales the pixel data to the range of 0 to 1. Data processing involves resizing and normalization. The input is the frame extracted in step 2, and the output is the resized and normalized frame.

[1010] Step 4:

[1011] The server extracts features from the resized and normalized frames. A pre-trained model using TensorFlow is used for feature extraction. Specifically, each frame is input to a specific neural network, which outputs a feature vector. The input is the frame processed in step 3, and the output is a feature vector.

[1012] Step 5:

[1013] The server inputs the feature vectors into a machine learning model to generate a natural language text summary. A video analysis model analyzes the feature vectors and extracts important information. A natural language generation model (e.g., GPT-2) then generates a text summary based on that information. Data calculations involve converting features to text. The input is the feature vector, and the output is the natural language text summary.

[1014] Step 6:

[1015] The server returns the generated natural language text summary to the user. The generated text summary is provided in a format that the user can view through their terminal. The input is the natural language text summary generated in step 5, and the output is the text summary returned to the user's terminal. This allows the user to refer to the provided summary.

[1016] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[1017] The present invention combines a system that analyzes video files provided by users and generates natural language text from them with an emotion engine that recognizes the user's emotions. The present invention will be described below with specific examples.

[1018] First, the user uses the terminal to provide the server with the path to the video file to be analyzed. The video file can be any video data, such as a dance practice video or a presentation practice video.

[1019] The server receives the video file path provided by the user and opens the file. Once the video file is read, the server extracts each frame to analyze its content. Each frame is resized, for example to 224x224 pixels. Each resized frame is then normalized. Normalization is the process of scaling the frame's data to a range of 0 to 1, allowing the machine learning model to process the data more efficiently.

[1020] The normalized frames are extracted as features. These features represent specific information contained in the data as numerical data. The extracted features are used as input to a machine learning model. The server uses these features to generate natural language text using a pre-trained video analysis model and natural language generation model.

[1021] A distinctive feature of the present invention is the emotion engine. The emotion engine analyzes emotions from the user's facial expressions and voice. This analyzed emotion data is fed back to the machine learning model and reflected in the generated natural language text. For example, if the user is nervous, the generated advice is "Relax and try again."

[1022] For example, based on a dance video provided by the user, a response such as "This dance consists of basic hip-hop steps, and it is important to step left and right in time with the rhythm" is generated. Furthermore, if the user shows a happy emotion, positive feedback such as "You're doing great!" is added.

[1023] Finally, the server returns the generated natural language text to the user, who can then review the text and take action, such as improving their activities, based on the text.

[1024] Through this series of processes, the present invention realizes a technology that analyzes non-verbal information and enables responses in natural language that take the user's emotions into account. In particular, by combining a pre-trained model with an emotion engine, more accurate analysis and responses that take emotions into account become possible.

[1025] The processing flow will be explained below.

[1026] Step 1:

[1027] The user uses the terminal to provide the server with the path of the video file to be analyzed, for example, the path of a dance practice video file.

[1028] Step 2:

[1029] The server receives the path to the video file provided by the user and opens the file using the cv2.VideoCapture function.

[1030] Step 3:

[1031] The server loads the video file frame by frame. For each frame, it does the following until it reaches the end of the video:

[1032] Step 4:

[1033] Use the cv2.resize function to resize the frame loaded by the server to a given size, for example 224x224 pixels.

[1034] Step 5:

[1035] The server normalizes the resized frame by dividing each pixel's value by 255, scaling it to the range 0 to 1.

[1036] Step 6:

[1037] The server adds the normalized frame to the list as a feature. This operation is performed for all frames.

[1038] Step 7:

[1039] Once the server has finished processing all frames, it converts the feature list into a NumPy array.

[1040] Step 8:

[1041] The pad_sequences function is used to pad the feature data acquired by the server to a fixed length, thereby preparing the input format for the machine learning model.

[1042] Step 9:

[1043] The server inputs the feature data into a pre-trained video analysis model and extracts the necessary information.

[1044] Step 10:

[1045] The server runs an emotion engine using the user's real-time video and audio data to analyze the user's emotions, for example, by using facial recognition technology or voice analysis technology to identify the user's emotions.

[1046] Step 11:

[1047] The server then feeds the analyzed emotion data back to the machine learning model, which then uses that data to generate natural language text. For example, if the user is nervous, the system generates advice like, "Relax and try again."

[1048] Step 12:

[1049] The server returns the generated natural language text to the user, providing advice and explanations based on the dance content, as well as emotionally sensitive feedback.

[1050] Step 13:

[1051] The user can review the generated natural language text through the device and take action based on it to improve their activity, for example, practicing a dance step again in accordance with the new advice.

[1052] Example 2

[1053] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1054] Conventional video analysis systems were unable to consider the user's emotions when generating text from video content. As a result, the generated text could not provide advice or feedback appropriate to the user's situation, resulting in poor usability. A particular issue was the inability to properly reflect the user's emotions, such as tension or joy.

[1055] The identification process by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes: means for reading a video file provided by a user; means for extracting each frame of the video file, resizing it to a predetermined size, and normalizing it; means for extracting the normalized frames as features; means for inputting the features into a machine learning model and generating natural language text; means for returning the generated natural language text; means including an emotion engine for analyzing the user's facial expressions and voice and recognizing their emotions; and means for feeding back the emotion data to the machine learning model and reflecting it in the natural language text. This makes it possible to generate natural language text that reflects the user's emotions in the analysis results of the video content.

[1056] "User" means an individual or organization that uses the system to provide video files and receive analysis results.

[1057] A "video file" is a digital file in a video data format that a user provides to the system.

[1058] A "frame" refers to one of the still images that make up a video file, and a video is made up of a series of many of these frames.

[1059] "Resizing" is the process of changing the resolution of each frame of a video file to a specific size.

[1060] "Normalization" is the process of scaling the range of data to a certain range, typically to a range from 0 to 1.

[1061] A "feature" is a specific pattern or piece of information in data that has been quantified and is used as input into a machine learning model.

[1062] A "machine learning model" is a pre-trained algorithm or network that analyzes data and makes predictions or classifications.

[1063] "Natural language text" is text in a language format that humans can understand, generated based on the results of video analysis.

[1064] An "emotion engine" is an algorithm or system that analyzes and recognizes emotions from a user's facial expressions, voice, etc.

[1065] "Feedback" is the process of inputting analysis results and data back into the system and reflecting the results.

[1066] MODE FOR CARRYING OUT THE INVENTION

[1067] The present invention combines a system that analyzes video files provided by users and generates natural language text from them with an emotion engine that recognizes the user's emotions. The present invention will be described below with specific examples.

[1068] First, the user uses their device to provide the server with the path to the video file to be analyzed. The video file can be any video data, such as a dance practice video or a presentation practice video. The server receives the path to the video file provided by the user and opens the file. Once the video file is read, the server extracts each frame to analyze its content. Each frame is resized, for example, to 224x224 pixels. Each resized frame is then normalized. Normalization is the process of scaling the frame data to a range of 0 to 1, which allows the machine learning model to process the data more efficiently.

[1069] The normalized frames are extracted as features. These features represent specific information contained in the data as numerical data. The server then inputs these features into a pre-trained video analysis model to analyze the content of the video. Based on the results of the video analysis, the server then generates natural language text using a pre-trained natural language generation model.

[1070] A distinctive feature of this system is the emotion engine, which analyzes emotions from the user's facial expressions and voice. This analyzed emotion data is fed back to the machine learning model and reflected in the generated natural language text. As a specific example, based on a dance video provided by the user, a response such as "This dance consists of basic hip-hop steps, and it is important to step left and right in time with the rhythm" is generated. Furthermore, if the user shows a happy emotion, positive feedback such as "You're doing great!" is added.

[1071] Finally, the server returns the generated natural language text to the user, who can review the text through their device and take action based on it, such as improving their activities.

[1072] Hardware and software used

[1073] This system consists of a server, a terminal, and related software components. Specifically, the following hardware and software are used:

[1074] Hardware:

[1075] Server: A high-performance computer that analyzes video files and generates natural language text.

[1076] Device: The device (PC, tablet, smartphone, etc.) where the user uploads the video file and reviews the generated text.

[1077] software:

[1078] Video analysis libraries: Libraries such as OpenCV for extracting, resizing, and normalizing frames from video files.

[1079] Machine learning frameworks, such as TensorFlow and PyTorch, are used to implement and run video analysis and natural language generation models.

[1080] Emotion analysis engine: A dedicated library or API for analyzing a user's facial expressions and voice.

[1081] Prompt Sentence Examples

[1082] "What happens if a user uploads a dance practice video?"

[1083] In response to this prompt, the system will generate an explanation like this:

[1084] "When a user uploads a dance practice video, the server analyzes the video, extracts each frame, and resizes and normalizes it. It then extracts features and generates natural language text based on them using a pre-trained model. If the user's sentiment is positive, it will add feedback like, 'You're doing great!'"

[1085] Through this series of processes, the present invention realizes a technology that analyzes non-verbal information and generates a response in natural language that takes into account the user's emotions.

[1086] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1087] Step 1:

[1088] The user inputs the path of the video file to be analyzed using the device. The input video file path is, for example, "C:\Users\SampleUser\Videos\dance_practice.mp4". This is the input.

[1089] Step 2:

[1090] The device sends the path of the input video file to the server. Specifically, it sends the path information using an HTTP request or other communication method. The output is the path information received by the server.

[1091] Step 3:

[1092] Opens the specified video file based on the video file path received by the server. Here, the input is the video file path and the output is the video file in its open state.

[1093] Step 4:

[1094] The server extracts each frame of the video file. Specifically, it uses a library such as OpenCV to obtain image data for each frame. The input is the video file, and the output is the extracted frames.

[1095] Step 5:

[1096] The server resizes each extracted frame to 224x224 pixels, for example using the OpenCV cv2.resize function. The input is the image data of the frame, and the output is the resized frames.

[1097] Step 6:

[1098] The server normalizes each resized frame, scaling each pixel value to the range 0 to 1 by dividing it by 255. The input is a set of resized frames, and the output is a set of normalized frames.

[1099] Step 7:

[1100] The server extracts features from the normalized frames. For example, it uses a convolutional neural network (CNN) to extract specific features from the frames as numerical data. The input is the normalized frames, and the output is the feature data.

[1101] Step 8:

[1102] The server analyzes the features using a pre-trained video analysis model. To obtain the analysis results, the feature data is input into the model. The input is the feature data, and the output is the analysis results.

[1103] Step 9:

[1104] The server generates text based on the analysis results using a pre-trained natural language generation model. The input is the analysis results, and the output is the generated natural language text.

[1105] Step 10:

[1106] The server uses an emotion engine to analyze the user's facial expressions and voice to obtain emotion data. The input is the user's facial expression image and voice data, and the output is emotion data.

[1107] Step 11:

[1108] The server feeds the emotion data back to the machine learning model and reflects it in the natural language text. Specifically, it adds advice and feedback according to the emotion to the text. The input is the emotion data and the initial natural language text, and the output is the final natural language text that takes emotion into account.

[1109] Step 12:

[1110] The server returns the generated natural language text to the user. Specifically, it returns the text using an HTTP response, etc. The input is the final natural language text, and the output is the text displayed on the user's terminal.

[1111] Step 13:

[1112] The user views the generated natural language text through a terminal. The input is the returned text, and the output is the user's understanding and action based on it.

[1113] The above are the processing steps of the present invention.

[1114] (Application example 2)

[1115] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1116] Conventional video analysis systems simply analyze video data and generate text without considering the user's emotions. As a result, in fields such as customer service and product recommendations, they are unable to generate responses that appropriately reflect the customer's emotions and reactions, making it difficult to improve customer satisfaction and communicate effectively. The objective of this invention is to provide more effective and personalized responses by analyzing the user's emotions and generating natural language text that reflects them.

[1117] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[1118] In this invention, the server includes means for reading a video file provided by a user, means for extracting each frame of the video file, resizing it to a predetermined size, and normalizing it, means for extracting the normalized frames as features, means for inputting the features into a machine learning model to generate natural language text, means for returning the generated natural language text, means for creating a context for generating appropriate responses and suggestions for customer service, and means for analyzing the user's emotions and reflecting them in the context. This makes it possible to generate natural language text that reflects the user's emotions, thereby enabling more effective and personalized customer service.

[1119] A "user" is an entity that uses this system to provide videos and receive analysis results.

[1120] "Video file" refers to video data provided by the user, and is the media data to be analyzed.

[1121] The term "means" refers to a method or technique for realizing the functions or processes of the present invention.

[1122] A "frame" is an individual still image that makes up a video file.

[1123] "Resizing" is a process of changing the size of an image or frame to a predetermined size.

[1124] "Normalization" is the process of scaling data to a certain range, primarily used to make data easier to handle uniformly in machine learning.

[1125] A "feature" is specific information extracted from data expressed as numerical data.

[1126] A "machine learning model" refers to an algorithm or mathematical model that learns patterns from data and makes predictions and classifications.

[1127] "Natural language text" refers to human-understandable sentences generated from analyzed video data.

[1128] "Context" refers to the background information and circumstances for generating customer service responses and suggestions, and also includes the results of user sentiment analysis.

[1129] "Means for analyzing emotions" refers to a technology or method that detects the user's emotional state from facial expressions, voice, etc., and uses that information within the system.

[1130] "Means for returning" refers to the method or technique for presenting the generated natural language text to the user.

[1131] The present invention is a system that analyzes a video file provided by a user, generates natural language text from the video file, and further combines it with an emotion engine that recognizes the user's emotions. Specific embodiments of the present invention will be described in detail below.

[1132] First, the user uses their device to provide the server with the path to the video file to be analyzed. The video file can be any video data, such as a recorded video of a customer service session. The server receives the path to the video file provided by the user and opens the file. Once the video file is read, the server extracts each frame to analyze its content. Each frame is resized, for example, to 224x224 pixels. Each resized frame is then normalized. Normalization is the process of scaling the frame data to a range of 0 to 1, which allows the machine learning model to process the data more efficiently.

[1133] The normalized frames are extracted as features. These features represent specific information contained in the data as numerical data. The extracted features are used as input to a machine learning model. The server uses these features to generate natural language text using a pre-trained video analysis model and natural language generation model.

[1134] A distinctive feature of the present invention is the emotion engine. The emotion engine analyzes emotions from the user's facial expressions and voice. This analyzed emotion data is fed back to the machine learning model and reflected in the generated natural language text. For example, if the user shows interest, a response such as "This product is new this season and is particularly popular" is generated. If the user loses interest, the response switches to "Can you tell me what other designs you like?"

[1135] Specifically, the server uses OpenCV to extract frames from the video (assuming a video capture camera), resizes the frames to 224x224 pixels, and normalizes them. The normalized data is then input into machine learning models: EmotionEngine (emotion recognition) and TextGenerationModel (natural language generation). The data obtained from sentiment analysis influences the generated text, resulting in natural language text that reflects the user's emotions.

[1136] Examples:

[1137] This application can be used in a scenario where a salesperson in a fashion store is wearing smart glasses and serving customers. When a customer sees an item of clothing that interests them, the smart glasses respond to their emotions and provide appropriate guidance, such as, "This item is new this season and is particularly popular. Please let us know if you need more information." If the customer appears to be losing interest, the smart glasses can switch to a suggestion such as, "Can you tell us what other designs you like?"

[1138] Example prompt:

[1139] "Customer interested" prompt:

[1140] Context: Your customer is interested.

[1141] Frame Data: [frame 1, frame 2, ...]

[1142] Generated text type: Product description

[1143] "Customer is losing interest" prompt:

[1144] Context: Your customer is losing interest.

[1145] Frame Data: [frame 1, frame 2, ...]

[1146] Type of generated text: Questions and Answers

[1147] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1148] Step 1:

[1149] The user uses the device to provide the server with the path to the video file to be analyzed. The path to the video file is specified as input, and the data is sent to the server as output. In this step, the specific actions are to select and send the video file.

[1150] Step 2:

[1151] The server receives the video file path provided by the user and opens the file. The input is the video file path, and the output is the video file itself. This step involves reading the target file from the file system.

[1152] Step 3:

[1153] The server extracts each frame to analyze the video file. The input is the video file, and the output is a list of extracted frames. Specifically, the server uses OpenCV to split the video into frames.

[1154] Step 4:

[1155] Each frame is resized and normalized to 224x224 pixels. The input is the extracted frame, and the output is the resized and normalized frame. Specifically, the resizing and normalization processes are performed.

[1156] Step 5:

[1157] The normalized frames are extracted as features. The input is the normalized frames and the output is the feature data. In this step, a feature extraction algorithm is applied.

[1158] Step 6:

[1159] The server inputs the feature data into a machine learning model and generates natural language text. The input is the feature data, and the output is the generated natural language text. This includes operations that utilize a pre-trained video analysis model and a natural language generation model.

[1160] Step 7:

[1161] The emotion engine analyzes emotions from the user's facial expressions and voice. The input is normalized frame and voice data, and the output is analyzed emotion data. In this step, emotion recognition algorithms are applied.

[1162] Step 8:

[1163] The emotion data is fed back into the machine learning model and reflected in the generated natural language text. The input is emotion data and normal frame data for the generated text, and the output is the final natural language text. Here, the text generation process integrates emotion data.

[1164] Step 9:

[1165] The server returns the generated natural language text to the user. The input is the final natural language text and the output is the text that is displayed on the terminal. This step includes displaying the text in a user interface.

[1166] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.

[1167] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1168] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.

[1169] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[1170] FIG. 9 is a diagram illustrating an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and actions arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.

[1171] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.

[1172] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).

[1173] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.

[1174] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."

[1175] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values ​​indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.

[1176] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).

[1177] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.

[1178] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.

[1179] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[1180] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.

[1181] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.

[1182] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.

[1183] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.

[1184] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.

[1185] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.

[1186] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.

[1187] The following is further disclosed regarding the above embodiment.

[1188] (Claim 1)

[1189] means for reading a video file provided by a user;

[1190] means for extracting each frame of the video file, resizing it to a predetermined size, and normalizing it;

[1191] means for extracting the normalized frames as features;

[1192] a means for inputting the feature quantity into a machine learning model to generate natural language text;

[1193] means for returning the generated natural language text;

[1194] A system including:

[1195] (Claim 2)

[1196] 10. The system of claim 1, wherein the machine learning models include a pre-trained video analysis model and a natural language generation model.

[1197] (Claim 3)

[1198] 2. The system of claim 1, further comprising: means for padding the features to a fixed length.

[1199] "Example 1"

[1200] (Claim 1)

[1201] means for reading a video file provided by a user;

[1202] means for extracting each frame of the video file, resizing it to a predetermined size, and normalizing it;

[1203] means for extracting the normalized frames as features;

[1204] a means for inputting the feature quantity into a machine learning model to generate natural language text;

[1205] means for returning the generated natural language text;

[1206] A system including:

[1207] (Claim 2)

[1208] 10. The system of claim 1, wherein the machine learning models include a pre-trained video analysis model and a natural language generation model.

[1209] (Claim 3)

[1210] 2. The system of claim 1, wherein the normalization process scales each frame of the video to a range of 0 to 1.

[1211] (Claim 4)

[1212] 2. The system of claim 1, further comprising: means for padding the features to a fixed length.

[1213] (Claim 5)

[1214] 2. The system of claim 1, wherein the natural language text returned to the user includes explanations and comments based on the video content.

[1215] "Application Example 1"

[1216] (Claim 1)

[1217] means for reading a video file provided by a user;

[1218] means for extracting each frame of the video file, resizing it to a predetermined size, and normalizing it;

[1219] means for extracting the normalized frames as features;

[1220] means for inputting the features into a machine learning model to generate a natural language text summary;

[1221] means for returning the generated natural language text summary;

[1222] A system including:

[1223] (Claim 2)

[1224] 10. The system of claim 1, wherein the machine learning models include a pre-trained video analysis model and a natural language generation model.

[1225] (Claim 3)

[1226] 2. The system of claim 1, further comprising: means for padding the features to a fixed length.

[1227] "Example 2: Combining Emotion Engines"

[1228] (Claim 1)

[1229] means for reading a video file provided by a user;

[1230] means for extracting each frame of the video file, resizing it to a predetermined size, and normalizing it;

[1231] means for extracting the normalized frames as features;

[1232] a means for inputting the feature quantity into a machine learning model to generate natural language text;

[1233] means for returning the generated natural language text;

[1234] means including an emotion engine that analyzes a user's facial expression and voice and recognizes their emotions;

[1235] a means for feeding back the emotion data to a machine learning model and reflecting the emotion data in natural language text;

[1236] A system including:

[1237] (Claim 2)

[1238] 10. The system of claim 1, wherein the machine learning models include a pre-trained video analysis model and a natural language generation model.

[1239] (Claim 3)

[1240] 2. The system of claim 1, further comprising: means for padding the features to a fixed length.

[1241] "Application example 2 when combining emotion engines"

[1242] (Claim 1)

[1243] means for reading a video file provided by a user;

[1244] means for extracting each frame of the video file, resizing it to a predetermined size, and normalizing it;

[1245] means for extracting the normalized frames as features;

[1246] a means for inputting the feature quantity into a machine learning model to generate natural language text;

[1247] means for returning the generated natural language text;

[1248] A means of creating context for generating appropriate customer responses and suggestions;

[1249] means for analyzing a user's emotion and reflecting the emotion in the context;

[1250] A system including:

[1251] (Claim 2)

[1252] 10. The system of claim 1, wherein the machine learning models include a pre-trained video analysis model and a natural language generation model.

[1253] (Claim 3)

[1254] 2. The system of claim 1, further comprising: means for padding the features to a fixed length. [Explanation of symbols]

[1255] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Device 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robot< / url:> < / url:> < / url:> < / url:>

Claims

1. means for reading a video file provided by a user; means for extracting each frame of the video file, resizing it to a predetermined size, and normalizing it; means for extracting the normalized frames as features; a means for inputting the feature quantity into a machine learning model to generate natural language text; means for returning the generated natural language text; A system including:

2. 10. The system of claim 1, wherein the machine learning models include a pre-trained video analysis model and a natural language generation model.

3. The system of claim 1 further comprising means for padding the features to a fixed length.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A