System

The system addresses the limitations of single-modal data analysis by integrating multimodal input processing and emotion recognition to provide personalized content, enhancing user satisfaction and corporate competitiveness.

JP2026019062APending Publication Date: 2026-02-05SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024120471
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-07-25
Publication Date
2026-02-05

AI Technical Summary

Technical Problem

Conventional data analysis and customer service systems are limited to specific modalities (e.g., text or voice), making it difficult to integrate and analyze multimodal input data, and lack emotion recognition, leading to insufficient individualization of customer service and educational content, which hampers user satisfaction and corporate competitiveness.

Method used

A system that collects, preprocesses, and analyzes multimodal input data (text, audio, image, video) to identify user emotions and needs, generates customized content, and collects feedback to improve performance, using technologies like natural language processing, voice recognition, and facial recognition.

Benefits of technology

Enables quick generation of personalized content that accurately reflects user emotions and needs, improving customer service efficiency and training effectiveness.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026019062000001_ABST
    Figure 2026019062000001_ABST
Patent Text Reader

Abstract

A system is provided.SOLUTION: A system, comprising: means for collecting multi-modal input data; means for pre-processing the collected multi-modal input data; means for analyzing emotions and needs using the pre-processed data; means for generating customized content based on the analysis results; means for providing the generated content to users; and means for collecting feedback from users to improve the performance of the system.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The technology of the present disclosure relates to a system. [Background technology]

[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]

[0004] Conventional data analysis and customer service systems are often limited to specific modalities (e.g., text or voice), making it difficult to integrate and analyze multimodal input data. Furthermore, emotion recognition technology has not yet been fully implemented, making it difficult to accurately reflect user emotions and needs in real time. As a result, the individualization of customer service and educational content has been insufficient, creating challenges in improving user satisfaction and strengthening corporate competitiveness. [Means for solving the problem]

[0005] The present invention solves the above problems by the following means.

[0006] Specifically, the present invention provides a system that includes a means for collecting multimodal input data, a means for preprocessing the collected multimodal input data, a means for analyzing emotions and needs using the preprocessed data, a means for generating customized content based on the analysis results, a means for providing the generated content to users, and a means for collecting user feedback and improving system performance. This system enables integrated analysis of multiple modal data, such as text data, audio data, image data, and video data, to accurately grasp users' emotions and needs in real time. As a result, personalized content can be generated and provided quickly, improving customer service efficiency and training effectiveness.

[0007] "Multimodal input data" refers to data provided in multiple different modalities, such as text data, audio data, image data, and video data.

[0008] "Means for collecting" refers to an interface or device for receiving multimodal input data from a user and incorporating the data into the system.

[0009] "Preprocessing means" refers to a function that performs preparatory steps such as format conversion and data normalization to make the collected multimodal data easier to analyze.

[0010] "Means of analysis" refers to algorithms and AI models that use pre-processed data to identify and assess emotions and needs.

[0011] "Customized content" refers to information and media that is individually generated to address the needs and emotions of a particular user based on analyzed data.

[0012] "Means for providing" refers to a method or device for displaying or transmitting the generated customized content to a user.

[0013] "Means of collecting feedback" refers to the functions and processes for obtaining user evaluations and reactions and using them to improve the system.

[0014] "Means to improve performance" refers to the ability to retrain algorithms and models based on collected feedback to improve the accuracy and efficiency of the system.

[0015] "Natural language processing technology" refers to technology for analyzing text data and understanding its content and emotions.

[0016] "Voice recognition technology" refers to the technology for converting voice data into text and analyzing its content and emotions.

[0017] "Facial recognition technology" refers to technology that identifies faces from image and video data and analyzes emotions and facial expressions.

[0018] "Facial expression analysis technology" refers to technology that identifies facial expressions from image and video data and recognizes the emotions they express. [Brief explanation of the drawings]

[0019] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7]FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION

[0020] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.

[0021] First, the terms used in the following description will be explained.

[0022] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).

[0023] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.

[0024] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.

[0025] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.

[0026] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."

[0027] [First embodiment]

[0028] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.

[0029] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0030] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0031] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.

[0032] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0033] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0034] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.

[0035] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0036] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0037] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0038] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0039] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0040] The present invention relates to a system that uses multimodal inputs such as text, voice, images, and videos to understand the needs and emotions of users and provide customized content and solutions. The program of this system is explained below in natural language.

[0041] System Overview

[0042] This system consists of three entities: a server, a terminal, and a user. The server is responsible for data collection, preprocessing, analysis, content generation, and feedback collection and performance improvement, while the terminal functions as an interface with the user. As a user of the system, the user provides multimodal input data and receives customized content as a result.

[0043] Specific examples

[0044] For example, a case where this system is applied to customer support for small and medium-sized enterprises will be described.

[0045] 1. User Input

[0046] A user sends a text message through the chat window saying, "I don't know how to use the product," and then uploads a video showing their confused expression.

[0047] 2. Data collection

[0048] The terminal temporarily stores the text, audio, image, and video data sent by the user and transmits it to the server in real time.

[0049] 3. Data Preprocessing

[0050] The server preprocesses the various data it receives, specifically by segmenting text data, normalizing words and phrases, adjusting the sampling rate of audio data, and converting image and video formats.

[0051] 4. Emotion and Needs Analysis

[0052] The server analyzes the pre-processed data, using natural language processing technology to perform sentiment analysis of the text, speech recognition technology to convert the voice data into text and analyze its tone, and computer vision technology to perform facial recognition and facial expression analysis from image and video data to identify the user's emotional state.

[0053] 5. Customized Content Generation

[0054] Based on the analysis, the server generates customized content that addresses the user's needs and emotions. In this example, it creates a text message detailing how to use the product and a reassuring video guide.

[0055] 6. Providing Feedback

[0056] The server sends the generated content to the terminal,

[0057] The device displays this information to the user, who then takes action to resolve the problem based on the information provided.

[0058] 7. Continuous learning and improvement

[0059] The server collects feedback data from users and retrains the AI ​​model, improving the accuracy of sentiment analysis and needs prediction.

[0060] In this way, this system can quickly provide highly customized content based on the user's complex input data. This will improve the efficiency of customer service, increase customer satisfaction, and strengthen the competitiveness of small and medium-sized enterprises. Similar benefits can also be expected in the fields of education and research and development.

[0061] The processing flow will be explained below.

[0062] Step 1:

[0063] A user inputs multimodal data such as text, voice, images, and video. For example, a user may enter text such as "I don't know how to use the product" into a chat box and upload a video of a confused expression.

[0064] Step 2:

[0065] The device temporarily stores the text, voice, image, and video data entered by the user and transmits it to the server in real time. At this time, the data is temporarily cached to prevent data loss.

[0066] Step 3:

[0067] The server performs preprocessing on the received multimodal data, specifically tokenizing and normalizing text data, sample rate conversion for audio data, and format conversion for image and video data.

[0068] Step 4:

[0069] The server prepares the preprocessed data for analysis, for example, using natural language processing (NLP) to perform sentiment analysis on the text, using speech recognition technology to convert audio data into text, and using computer vision technology to perform facial recognition and facial expression analysis on image and video data.

[0070] Step 5:

[0071] The server identifies the user's emotional state and needs based on the analysis results. For example, the result of sentiment analysis may be "confused," and facial expression analysis may also detect a confused expression.

[0072] Step 6:

[0073] The server generates customized content tailored to the user based on the identified emotional state and needs, such as a text message detailing how to use a product or a reassuring video guide.

[0074] Step 7:

[0075] The server transmits the generated customized content to the terminal, compressing the data as necessary during transmission to ensure efficient data transmission.

[0076] Step 8:

[0077] The device displays the received content to the user. For example, a text message is displayed on the chat screen, and a video guidance is played on a player.

[0078] Step 9:

[0079] The user reviews the customized content provided and takes the necessary action to resolve the issue.

[0080] Step 10:

[0081] The server collects feedback data from users, including text and audio ratings.

[0082] Step 11:

[0083] The server analyzes the collected feedback, evaluates the performance of the AI ​​model, and identifies areas for improvement, which are then used for retraining to improve the accuracy of sentiment analysis and needs prediction.

[0084] In this way, through a series of steps, the system can quickly and accurately provide customized content based on the user's multiple input data.

[0085] Example 1

[0086] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0087] Many modern systems rely on single-modal input data from users (text only or voice only), which makes it difficult to fully understand users' emotions and needs, making it difficult to respond appropriately or provide customized content. Furthermore, there is a lack of means to properly collect user feedback and improve system performance, making continuous improvement difficult.

[0088] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0089] In this invention, the server includes a means for a user to input input data into the terminal, a means for the terminal to temporarily store the input data and send it to the server, and a means for the server to pre-process the collected multimodal input data. This makes it possible to analyze the user's emotions and needs from multiple angles and quickly provide highly customized content. It is also possible to collect feedback from users and continuously improve the system's performance.

[0090] A "user" is a person or organization that utilizes the system to input data and receive services.

[0091] A "terminal" is a device that allows a user to input data and send it to a server. Examples include computers, smartphones, and tablets.

[0092] "Server" means a device or system that receives data sent from a terminal, pre-processes and analyzes it, and generates and provides customized content.

[0093] "Multimodal input data" refers to data in multiple formats, including text data, audio data, image data, and video data.

[0094] "Preprocessing" is the process of preparing received data so that it can be easily analyzed, since it is difficult to handle as it is. This includes morphological analysis and normalization of text, adjusting the sampling rate of audio, and converting the format of images and videos.

[0095] "Emotion and needs analysis" refers to identifying a user's emotional state and the information and solutions they need based on pre-processed data. This analysis uses natural language processing, voice recognition, facial recognition, and facial expression analysis technologies.

[0096] "Customized content" refers to information or guidance generated based on the analysis results to address the user's individual needs and emotions, such as text messages or video guidance.

[0097] "Feedback" refers to the reactions and evaluations given by users to the content provided, and is information that can be used to improve the performance of the system.

[0098] "Natural language processing technology" is a technology that processes and analyzes human language using a computer. It includes morphological analysis and sentiment analysis.

[0099] "Speech recognition technology" is a technology that converts voice data into text data and analyzes the content and tone of the text.

[0100] "Facial recognition technology" is a technology that detects faces from images and videos and analyzes their features.

[0101] "Facial expression analysis technology" is a technology that uses facial recognition technology to identify a user's emotional state from their facial expression.

[0102] A "generative AI model" is an artificial intelligence model that generates new information based on given data. Examples include GPT-3.

[0103] A "prompt sentence" is the text of an instruction or question that a user enters into a system.

[0104] MODE FOR CARRYING OUT THE INVENTION

[0105] The present invention relates to a system that uses multimodal input data such as text, audio, images, and videos to understand the needs and emotions of users and provide customized content. Specific embodiments of this system are described below.

[0106] System Overview

[0107] This system consists of three entities: a server, a terminal, and a user. The server is responsible for data collection, preprocessing, analysis, generation of customized content, feedback collection, and performance improvement. The terminal functions as an interface with the user, who, as a user of the system, provides multimodal input data and receives customized content as a result.

[0108] Hardware and software used

[0109] Hardware: Server machines, user devices (computers, smartphones, tablets, etc.)

[0110] software:

[0111] Natural language processing libraries: NLTK, spaCy

[0112] Speech Recognition Library: Google Speech-to-Text API

[0113] Image processing library: OpenCV

[0114] Facial Expression Recognition API:Facial Expression Recognition API

[0115] Generative AI model: GPT-3

[0116] User Input

[0117] Users can use the chat window on their device to type text messages, or upload audio messages and video files. For example, a user might type "I don't know how to use this product" and attach a video of themselves looking confused.

[0118] Prompt Sentence Examples

[0119] I don't know how to use the new product. How do I set it up?

[0120] Video file example

[0121] Approximately 30 seconds of video file (mp4 format) containing a confused expression

[0122] Data collection and transmission

[0123] The device temporarily stores the text, audio, image, and video data entered by the user and transmits it to the server in real time. Specifically, the device transmits this data to the server using an HTTP POST request.

[0124] Data Preprocessing

[0125] The server preprocesses the various types of data it receives. For text data, it uses a natural language processing library (e.g., NLTK or spaCy) to perform morphological analysis and normalization. For audio data, it adjusts the sampling rate and removes noise. For image and video data, it converts them into an appropriate format using OpenCV and performs preprocessing for face detection.

[0126] Emotion and Needs Analysis

[0127] The server analyzes the preprocessed data. For text data, it performs sentiment analysis using models such as the BERT model. For audio data, it converts the audio to text using the Google Speech-to-Text API and analyzes its tone. For image and video data, it uses the Facial Expression Recognition API for facial recognition and expression analysis.

[0128] Customized content generation

[0129] Based on the analysis results, the server generates customized content that addresses the user's needs and emotions. Using a generative AI model (e.g., GPT-3), it creates text messages detailing how to use the product and reassuring video guidance.

[0130] Providing content and gathering feedback

[0131] The server sends the generated content to the device, which displays it to the user. The user then takes action to solve the problem based on the information provided. The server also collects feedback data from users to continuously improve the system's performance.

[0132] In this way, this system can quickly provide highly customized content based on the user's complex input data, resulting in more efficient customer service and improved customer satisfaction, thereby enhancing a company's competitiveness.

[0133] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0134] Step 1:

[0135] The user provides multimodal input data such as text, voice, images, and video to the device. The input data includes a prompt such as "I don't know how to use this new product. How do I set it up?" and a video file (mp4 format) containing a confused facial expression. The device receives this data based on the user's input.

[0136] Step 2:

[0137] The device temporarily stores the received text, audio, image, and video data and sends it to the server in real time using an HTTP POST request. Input data is stored and sent as is.

[0138] Step 3:

[0139] The server preprocesses the received multimodal input data, specifically by:

[0140] Text data: Perform morphological analysis and normalization using a natural language processing library (e.g., NLTK, spaCy).

[0141] Audio data: Adjust the sampling rate and remove noise.

[0142] Image and video data: Use OpenCV to convert them into the appropriate format and pre-process them for face detection.

[0143] The output is preprocessed text, clear audio data, and format-converted images and videos.

[0144] Step 4:

[0145] The server analyzes the preprocessed data, specifically:

[0146] Text data: Sentiment analysis is performed using models such as BERT.

[0147] Audio data: Converted to text using the Google Speech-to-Text API, and tone analysis is performed based on that text.

[0148] Image and video data: Use the Facial Expression Recognition API for facial recognition and facial expression analysis.

[0149] The output is an analysis that indicates the user's emotional state and needs.

[0150] Step 5:

[0151] The server generates customized content based on the analysis results. Using a generative AI model (e.g., GPT-3), it creates text messages with specific actions to take based on the user's needs and reassuring video guidance. The output is a customized text message and video guidance.

[0152] Step 6:

[0153] The server then sends the generated customized content to the device, which then receives it and displays it to the user. Specifically, the device displays text messages in a chat window and plays video guidance. The user can then take action to resolve the problem based on this content.

[0154] Step 7:

[0155] The server collects feedback data from users. Specifically, it collects the ratings and opinions users have given about the content provided and sends that data to the server. The server then retrains the generative AI model based on the collected feedback data, continuously improving the system's performance. This improves the accuracy of user sentiment analysis and need prediction, providing a better user experience.

[0156] (Application example 1)

[0157] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0158] Conventional customer support systems are unable to fully utilize the complex input data (text, voice, images, video, etc.) from users, making it difficult to accurately grasp users' emotions and specific needs. Food delivery services, in particular, are required to provide quick and accurate solutions to issues such as cold food, but current systems have limitations. Therefore, there is an urgent need to develop a system that can analyze user input data at multiple stages and provide appropriate solutions.

[0159] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0160] In this invention, the server includes means for collecting multimodal input data, means for preprocessing the collected multimodal input data, means for analyzing emotions and needs using the preprocessed data, means for generating customized content based on the analysis results, means for providing the generated content to users, means for collecting feedback from users and improving system performance, and means for analyzing text and video data entered by users and providing corresponding solutions, including generating guide videos including instructions on how to reheat food when it has cooled and an apology message. This makes it possible to quickly and accurately solve specific user problems and improve customer satisfaction.

[0161] "Multimodal input data" refers to data in multiple different formats, such as text data, audio data, image data, and video data.

[0162] "Means of collection" refers to the mechanisms and technologies for acquiring multimodal input data provided by users.

[0163] "Preprocessing means" refers to the techniques and algorithms used to prepare collected multimodal input data in an analyzable form.

[0164] "Means for emotion and needs analysis" refers to analytical tools and techniques for discovering user emotions and specific needs based on processed data.

[0165] "Means for generating customized content" refers to technology that creates specific information or guides tailored to the user's needs based on the analysis results.

[0166] "Means for providing" refers to the mechanisms and technologies for displaying or notifying the user of the generated customized content.

[0167] "Means of collecting feedback" refers to methods and techniques for obtaining responses and reactions from users and using them to improve the system.

[0168] "Means of providing solutions" refers to mechanisms and technologies for presenting users with specific solutions and proposals for solving problems or issues.

[0169] "Means for generating guide videos" refers to technology for creating video content that includes guidance and instructions that meet the needs of users.

[0170] System configuration

[0171] This system is composed of three elements: a server, a terminal, and a user. Each element is explained in detail below.

[0172] server

[0173] The server is responsible for data collection, preprocessing, analysis, content generation, and feedback collection and performance improvement. Specific software used includes TensorFlow, PyTorch, OpenCV, NLTK, and Transformer-based NLP models (e.g., BERT).

[0174] Terminal

[0175] The terminal is a device such as a smartphone that functions as an interface with the user. Specifically, it is built using React Native and plays a role in sending data entered by the user to the server and displaying the content received from the server.

[0176] User

[0177] Users are the users of the system and provide multimodal input data (text, audio, images, video), which is analyzed and customized content is returned to the user.

[0178] Processing flow

[0179] 1. User input:

[0180] The user inputs a text message and uploads video data via the device. For example, the user can provide a message such as "My food has arrived but it's cold and I'm worried" along with a video of the food being cold.

[0181] 2. Data Collection:

[0182] The device temporarily stores the text and video data sent by the user and then transmits it to the server. Data transmission is performed using real-time communication.

[0183] 3. Data preprocessing:

[0184] The server preprocesses the received data: text data is segmented using NLTK, and video data is converted to the appropriate format using OpenCV.

[0185] 4. Emotion and Needs Analysis:

[0186] The server analyzes the preprocessed data: text data is subjected to sentiment analysis using natural language processing techniques such as the BERT model, and video data is subjected to facial recognition and facial expression analysis using OpenCV.

[0187] 5. Customized Content Generation:

[0188] Based on the analysis results, it generates customized content that addresses the user's needs and emotions. In this example, it generates a guide video explaining how to reheat cold food and an apology message.

[0189] 6. Providing Feedback:

[0190] The generated content is sent to the device, which displays it to the user, who then takes action to solve the problem based on the information provided.

[0191] 7. Continuous learning and improvement:

[0192] The server collects feedback data from users and retrains the AI ​​model, improving the accuracy of sentiment analysis and needs prediction.

[0193] Examples and prompts

[0194] Specific examples

[0195] If a user texts, "My food arrived but it's cold and I'm worried," and uploads a video showing the food cold, the system will provide a video guide on how to reheat it and an apology message.

[0196] Prompt Sentence Examples

[0197] User input: The text "My food arrived but it's cold and I'm worried" and a video of the food being cold.

[0198] Desired output: An apology message and a guided video with specific steps for reheating the food.

[0199] The system of the present invention makes it possible to quickly and accurately solve specific user problems and improve customer satisfaction.

[0200] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0201] Step 1:

[0202] User input:

[0203] The user inputs a text message and uploads video data through the device. For example, the user can provide a text message saying, "My food arrived, but it's cold and I'm worried," along with a video of the cold food. The device then receives this input and prepares it for the next step.

[0204] input:

[0205] Text messages, video data

[0206] output:

[0207] Temporarily saved text messages and video data

[0208] Step 2:

[0209] Data collection:

[0210] The device temporarily stores the text and video data sent by the user and transmits it to the server in real time. This data transmission requires a stable internet connection.

[0211] input:

[0212] Temporarily saved text messages and video data

[0213] output:

[0214] Text messages and video data sent to the server

[0215] Step 3:

[0216] Data preprocessing:

[0217] The server preprocesses the received data: text data is segmented using natural language processing technology (NLTK), and video data is converted into a format that can be analyzed frame by frame using OpenCV.

[0218] input:

[0219] Text messages and video data sent to the server

[0220] output:

[0221] Segmented text data, converted video data

[0222] Step 4:

[0223] Emotion and Needs Analysis:

[0224] The server analyzes emotions and needs based on the preprocessed data. It uses the BERT model for sentiment analysis on text data, and OpenCV for facial recognition and facial expression analysis on video data to identify the user's emotional state.

[0225] input:

[0226] Segmented text data, converted video data

[0227] output:

[0228] Analyzed emotion and needs data

[0229] Step 5:

[0230] Customized Content Generation:

[0231] Based on the analysis results, the system generates customized content that responds to the user's needs and emotions, such as a guide video explaining how to reheat cold food and an apology message.

[0232] input:

[0233] Analyzed emotion and needs data

[0234] output:

[0235] Guide video, apology message

[0236] Step 6:

[0237] Providing feedback:

[0238] The server sends the generated content to the device, which displays it to the user, who then acts to solve the problem based on the information provided.

[0239] input:

[0240] Guide video, apology message

[0241] output:

[0242] Guide video and apology message displayed to users

[0243] Step 7:

[0244] Continuous learning and improvement:

[0245] The server collects feedback data from users and uses it to retrain the AI ​​model, improving the accuracy of sentiment analysis and needs prediction.

[0246] input:

[0247] User feedback data

[0248] output:

[0249] Improved AI models

[0250] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.

[0251] This invention relates to a system that provides customized content by utilizing multimodal inputs such as text, voice, images, and video, and by performing advanced analysis of user needs and emotions. In particular, by combining an emotion engine, it is possible to accurately recognize the user's emotional state and generate appropriate responses.

[0252] System Overview

[0253] This system is composed of three entities: a server, a terminal, and a user. The server collects data, preprocesses it, analyzes it, generates customized content, and collects feedback and improves performance, while the terminal functions as an interface with the user. As a user of the system, the user provides multimodal input data and receives customized content based on the analysis results.

[0254] Incorporating an emotion engine

[0255] One of the features of this system is the incorporation of an emotion engine, which integrates natural language processing (NLP), speech recognition, facial recognition, and facial expression analysis technologies to analyze the user's emotions in real time.

[0256] Specific examples

[0257] The application of this system to customer support for small and medium-sized enterprises will be explained.

[0258] 1. User Input

[0259] A user texts the support chat saying, "I don't know how to use the product," and then uploads a video showing a confused expression.

[0260] 2. Data collection

[0261] The device temporarily stores the multimodal data sent by the user, including text, audio, images, and video, and transmits it to the server in real time. The data is temporarily cached to prevent data loss.

[0262] 3. Data Preprocessing

[0263] The server preprocesses the various data it receives, specifically tokenizing and normalizing text data, converting the sampling rate of audio data, and converting the format of image and video data.

[0264] 4. Emotion and Needs Analysis

[0265] The emotion engine uses the pre-processed data to analyze the user's emotional state and needs, using natural language processing technology to perform sentiment analysis of the text, speech recognition technology to convert the voice data into text and analyze its tone, and computer vision technology to perform facial recognition and facial expression analysis from images and videos.

[0266] 5. Customized Content Generation

[0267] Based on the analysis, the server generates customized content that responds to the user's emotional state and needs, such as a text message detailing how to use the product and a reassuring video guide.

[0268] 6. Providing Feedback

[0269] The server transmits the generated customized content to the terminal, compressing the data as necessary to transmit the data efficiently.

[0270] The device displays the received content to the user, displaying text messages on the screen and playing video guidance.

[0271] 7. Continuous learning and improvement

[0272] The server collects feedback from the user, which may include a text or audio rating.

[0273] The server analyzes the collected feedback and retrains the emotion engine algorithm, thereby improving the system's accuracy and response quality.

[0274] This system can accurately grasp the user's emotional state and needs in real time based on a variety of input data, and quickly provide highly customized content. This can improve the efficiency of customer service and customer satisfaction, strengthening the competitiveness of small and medium-sized enterprises. Similar benefits can be expected in the fields of education and research and development.

[0275] The processing flow will be explained below.

[0276] Step 1:

[0277] A user inputs multimodal data such as text, audio, images, and video. For example, a user texts a support chat message saying, "I don't know how to use the product," and uploads a video showing a confused expression.

[0278] Step 2:

[0279] The device temporarily stores the text, voice, image, and video data entered by the user and transmits it to the server in real time, using a cache process to prevent data loss.

[0280] Step 3:

[0281] The server preprocesses the multimodal data it receives, specifically tokenizing and normalizing text data, converting the sampling rate of audio data, and converting the format of image and video data.

[0282] Step 4:

[0283] The emotion engine analyzes the pre-processed data, using natural language processing technology to perform sentiment analysis on the text data, speech recognition technology to convert the voice data into text data and analyze its tone, and computer vision technology to perform facial recognition and facial expression analysis on image and video data to identify the user's emotional state.

[0284] Step 5:

[0285] The server identifies the user's needs and emotional state based on the analysis results from the emotion engine. For example, if the text data is judged as "confused," a confused expression is also detected in the facial expression analysis.

[0286] Step 6:

[0287] The server generates customized content based on the identified needs and emotional state, for example, creating a text message detailing how to use a product or generating a reassuring video guide.

[0288] Step 7:

[0289] The server transmits the generated customized content to the terminal, compressing the data as necessary to improve the efficiency of data transmission.

[0290] Step 8:

[0291] The device displays the received customized content to the user, displaying text messages on the screen and playing video guidance.

[0292] Step 9:

[0293] The user reviews the customized content provided and takes action to resolve the issue.

[0294] Step 10:

[0295] The server collects feedback from the user, which may be received in text or audio format.

[0296] Step 11:

[0297] The server analyzes the feedback data and retrains it to improve the performance of the emotion engine and the overall system, thereby increasing the accuracy of future analysis and content generation.

[0298] In this way, incorporating an emotion engine makes it possible to perform advanced analysis of various user data and provide customized content tailored to needs in real time. This will significantly improve the quality and efficiency of customer service, helping to strengthen the competitiveness of small and medium-sized enterprises.

[0299] Example 2

[0300] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0301] Conventional customer support systems have struggled to efficiently collect and analyze diverse user input data and provide customized content tailored to the user's emotional state and needs. While systems that utilize multimodal input data to accurately grasp the user's emotions and generate and deliver appropriate responses are particularly needed, few such systems exist. This has led to lower user satisfaction and negatively impacted the efficiency of companies' customer support efforts.

[0302] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[0303] In this invention, the server includes: means for a user to provide multimodal input data including text, voice, images, and videos; means for temporarily storing the received multimodal input data and transmitting it to the server in real time; data preprocessing means for tokenizing, normalizing, formatting, and sampling rate converting the data received by the server; means for an emotion engine to analyze the user's emotions and needs using natural language processing, speech recognition, face recognition, and facial expression analysis techniques using the preprocessed data; means for generating customized content based on the analysis results; means for transmitting the generated content to a terminal and providing it to the user; and means for collecting user feedback and retraining the system's emotion engine to improve performance. This makes it possible to efficiently collect and analyze a variety of user input data and provide appropriate customized content in real time, thereby improving customer service efficiency and user satisfaction.

[0304] "Multimodal input data" refers to data in multiple formats, such as text, audio, images, and video.

[0305] "Preprocessing" refers to processing such as tokenization, normalization, format conversion, and sampling rate conversion to make data easier to analyze.

[0306] An "emotion engine" refers to a system that analyzes a user's emotional state by integrating natural language processing technology, voice recognition technology, facial recognition technology, and facial expression analysis technology.

[0307] "Natural language processing technology" refers to technology that enables computers to understand, analyze, and generate human language.

[0308] "Speech recognition technology" refers to the technology that converts speech into text data and analyzes its content.

[0309] "Facial recognition technology" refers to technology that detects and identifies human faces from images and videos.

[0310] "Facial expression analysis technology" refers to technology that analyzes facial expressions and identifies their emotional state.

[0311] "Customized content" refers to individually optimized information and responses generated based on a user's emotional state and needs.

[0312] "Feedback" refers to subsequent data such as ratings, opinions, and impressions collected from users.

[0313] "Retraining" refers to the process of retraining an existing learning model with new data to improve its performance.

[0314] "Server" refers to the central management system that collects, pre-processes, analyzes, generates content, gathers feedback, and improves performance of data.

[0315] "Terminal" refers to a device that acts as an interface with a user and is responsible for inputting and outputting data.

[0316] This invention relates to a system that uses multimodal input data such as text, voice, images, and video to provide customized content through advanced analysis of user needs and emotions. In particular, by combining an emotion engine, it becomes possible to accurately recognize the user's emotional state and generate appropriate responses. This system is composed of three entities: a "server," a "terminal," and a "user."

[0317] The server collects, preprocesses, and analyzes data, generates customized content, and collects feedback to improve performance. The terminal functions as an interface with the user, who provides multimodal input data and receives customized content based on the analysis results.

[0318] Hardware and software used

[0319] The servers are data centers or cloud computing services with high-performance computing capabilities, and the software uses the BERT model for natural language processing (NLP), the Google Speech-to-Text API for speech recognition, and OpenCV and Dlib libraries for facial recognition and facial expression analysis.

[0320] Specific examples

[0321] We will explain how this system can be applied to customer support for small and medium-sized enterprises.

[0322] User Input

[0323] A user can text the support chat saying, "I don't know how to use the product," and then upload a video showing a confused expression, which inputs the user's needs and emotions into the system.

[0324] Data collection

[0325] The device temporarily stores the multimodal data received from the user, including text, audio, images, and video, and transmits it to the server in real time. The data is temporarily cached to prevent data loss.

[0326] Data Preprocessing

[0327] The server performs preprocessing on the received data, including tokenization, normalization, format conversion, and sampling rate conversion. Specifically, text data is processed using the BERT model, and the sampling rate of audio data is converted to 16 kHz to optimize it for analysis. For image and video data, the resolution is changed to 1280x720 pixels using the OpenCV library and the data is converted to the required format.

[0328] Emotion and Needs Analysis

[0329] The emotion engine uses the pre-processed data to analyze the user's emotional state and needs. It uses natural language processing technology to perform sentiment analysis of the text, speech recognition technology to convert the voice data into text and analyze its tone, and computer vision technology to perform facial recognition and facial expression analysis from images and videos to understand the user's emotions.

[0330] Customized Content Generation

[0331] Based on the analysis results, the server generates customized content that corresponds to the user's emotional state and needs. For example, based on the emotion analysis result of "confused," a text message explaining detailed product usage and a reassuring video guide are generated.

[0332] Providing feedback

[0333] The server transmits the generated customized content to the terminal, which compresses the content as necessary for efficient transmission. The terminal then displays the received content to the user, including displaying a text message on the screen and playing a video guide.

[0334] Continuous learning and improvement

[0335] The server collects feedback from users, including text and voice ratings, and analyzes the collected feedback and re-trains the emotion engine using TensorFlow to improve the system's accuracy and response quality.

[0336] Prompt Sentence Examples

[0337] "If a customer is confused about how to use a product, use an emotion engine to analyze their emotions in real time and generate a text message explaining how to use the product and a reassuring video guide."

[0338] This makes it possible to efficiently collect and analyze a wide range of user input data and provide appropriate customized content in real time. This system will improve the efficiency of customer service and user satisfaction, strengthening the competitiveness of small and medium-sized enterprises. Similar benefits can be expected in the fields of education and research and development.

[0339] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0340] Step 1:

[0341] The user types a text message into the support chat and uploads a video showing a confused expression. The input data includes four types: text, audio, images, and video. This allows the user's needs and emotions to be input into the system. For example, the user can send a specific text message such as, "I don't know how to use the product."

[0342] Input: Text message, video (confused expression)

[0343] Output: Collected multimodal data

[0344] Step 2:

[0345] The device temporarily stores multimodal data received from the user, including text, audio, images, and video. The stored data is temporarily cached and then sent to the server in real time. This prevents data loss and enables analysis on the server side.

[0346] Input: Collected multimodal data

[0347] Output: Multimodal data sent to the server

[0348] Step 3:

[0349] The server preprocesses the received data. For text data, tokenization and normalization are performed using the BERT model. For audio data, the sampling rate is converted to 16 kHz to optimize the data for analysis. For image and video data, the resolution is changed to 1280x720 pixels using the OpenCV library and the data is converted to the required format.

[0350] Input: Multimodal data sent to the server

[0351] Output: Preprocessed data (tokenized, normalized, formatted, and sample rate converted)

[0352] Step 4:

[0353] The emotion engine uses the pre-processed data to analyze the user's emotional state and needs. It uses natural language processing technology to perform sentiment analysis of the text, speech recognition technology to convert the voice data into text and analyze its tone, and computer vision technology to perform facial recognition and facial expression analysis from images and videos to understand the user's emotions. The results of this analysis determine the user's specific emotions and needs.

[0354] Input: Preprocessed data

[0355] Output: User's emotional state and needs analysis results

[0356] Step 5:

[0357] Based on the analysis results, the server generates customized content that corresponds to the user's emotional state and needs. For example, based on the emotion analysis result of "confused," a text message explaining detailed product usage and a reassuring video guide are generated.

[0358] Input: User's emotional state and needs analysis results

[0359] Output: Customized content (text messages, video guidance)

[0360] Step 6:

[0361] The server transmits the generated customized content to the terminal, which compresses the content as necessary for efficient transmission. The terminal then displays the received content to the user, including displaying a text message on the screen and playing a video guide.

[0362] Input:Customized Content

[0363] Output: Customized content served to the user

[0364] Step 7:

[0365] The server collects feedback from users, including text and voice ratings, and analyzes the collected feedback and re-trains the emotion engine using TensorFlow to improve the system's accuracy and response quality.

[0366] Input: User feedback

[0367] Output: Performance improvement by retraining the emotion engine

[0368] The above are the specific processing steps of the program in this system.

[0369] (Application example 2)

[0370] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0371] Current content distribution services struggle to analyze a user's emotional state and needs in real time and provide appropriate entertainment content. Conventional systems often fail to provide personalized recommendations based on the user's emotions and state, resulting in reduced user satisfaction. The present invention aims to address these challenges and provide optimal content tailored to the user's emotional state and needs.

[0372] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[0373] In this invention, the server includes means for collecting multimodal input data, means for preprocessing the collected multimodal input data, means for analyzing emotions and needs using the preprocessed data, means for generating customized content based on the analysis results, means for providing the generated content to the user, means for collecting feedback from the user and improving the performance of the system, and means for recommending entertainment content based on the emotional state and needs of the user, thereby making it possible to provide personalized entertainment content according to the emotional state and needs of the user in real time.

[0374] "Multimodal input data" refers to data in multiple formats, such as text data, audio data, image data, and video data.

[0375] "Means for collection" refers to the devices and software that obtain and store multimodal input data from users.

[0376] "Preprocessing means" refers to devices or software that perform processes to convert collected multimodal input data into a format that is easy to analyze.

[0377] "Emotion and needs analysis means" refers to algorithms or software that identify a user's emotional state and needs based on pre-processed data.

[0378] "Means for generating customized content" refers to systems or software for creating optimal content for users based on the analysis results.

[0379] The "means for providing to the user" refers to a device or interface for showing the generated customized content to the user.

[0380] "Means for collecting user feedback" refers to devices or software that collect and record user ratings and opinions.

[0381] "Means for improving system performance" refers to methods for improving the behavior of algorithms or software based on collected feedback.

[0382] "Emotional state" refers to data or states that represent a user's mental and emotional state.

[0383] "Needs" refer to the requests and desires for information, functions, and services that users desire.

[0384] "Entertainment content" refers to various media content for entertainment purposes, such as movies, dramas, music, and games.

[0385] The present invention relates to a system for providing customized entertainment content based on a user's emotional state and needs, which is implemented in the following steps:

[0386] System configuration

[0387] This system consists of three components: a server, a terminal, and a user. The server collects data, preprocesses it, analyzes it, generates content, and collects feedback to improve performance. The terminal functions as an interface with the user. As a user of the system, the user provides multimodal input data and receives customized content based on the analysis results.

[0388] Hardware and software used

[0389] This system uses the following hardware and software:

[0390] Hardware:

[0391] Device: Smartphone or head-mounted display equipped with a camera and microphone

[0392] software:

[0393] Cloud services: Cloud services for data processing (e.g., AWS, Google Cloud)

[0394] Natural language processing engine: Google NLP

[0395] Speech recognition engine: Amazon Transcribe

[0396] Facial expression and face recognition engine: Microsoft Azure Face API

[0397] Data processing and calculation

[0398] The device captures multimodal input data (text, voice, image, and video) from the user. This data is temporarily cached and sent to the server. The server preprocesses the received data, tokenizing and normalizing the text data. The voice data undergoes sampling rate conversion and text conversion, and the image and video data undergo format conversion and facial expression analysis.

[0399] The following techniques are used for the analysis:

[0400] Sentiment analysis of text using natural language processing (NLP) technology

[0401] Voice recognition technology converts voice data into text and analyzes the tone

[0402] Uses computer vision technology to perform facial recognition and facial expression analysis

[0403] Content generation and delivery

[0404] Based on the analysis results, the server generates a recommendation list of entertainment content (movies, dramas, music) according to the user's emotional state and needs. This recommendation list is customized and accessible to the user via smartphone or head-mounted display. The user can view the recommended content and provide feedback.

[0405] Specific examples

[0406] For example, if a user uses a smartphone to type the text "I'm feeling a bit down today," input a low-pitched voice into the microphone, and capture an image of a gloomy facial expression with the camera, all of this data is immediately sent to the server and analyzed.

[0407] The server analyzes the user's emotional state (feeling depressed) and needs (content that will lift their spirits) and recommends the most suitable movies and music for the user.

[0408] Prompt Sentence Examples

[0409] If you send "I'm feeling a bit down today," enter the following prompt into the model:

[0410] text

[0411] User sent text: "I'm feeling a bit down today."

[0412] Voice data: User's voice is low and restless

[0413] Facial expression data: The user has a dull expression

[0414] Use this data to generate content recommendations for your users.

[0415] This makes it possible to provide entertainment content in real time that matches the user's emotional state and needs, thereby increasing user satisfaction and improving the value of content distribution services.

[0416] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0417] Step 1:

[0418] The device collects the user's input data. Specifically, the device's camera and microphone are used to capture the user's facial expressions and voice in real time, and the user inputs text. The input data at this time includes text data, voice data, and image data. This data is temporarily stored on the device and then sent to the server.

[0419] input:

[0420] Text data: Character information entered by the user

[0421] Voice data: recordings of your voice

[0422] Image data: photos or videos of the user's facial expressions

[0423] output:

[0424] Capture and temporary storage of multimodal input data (text, audio, images)

[0425] Step 2:

[0426] The server preprocesses the received multimodal input data, specifically tokenizing and normalizing text data, converting audio data to text through sampling rate conversion, and formatting and analyzing facial expressions on image and video data.

[0427] input:

[0428] Multimodal input data (text, audio, images)

[0429] output:

[0430] Preprocessed data (tokenized text, transcribed audio, analyzed images)

[0431] Step 3:

[0432] The server's emotion engine uses the pre-processed data to analyze emotions and needs, using natural language processing (NLP) technology to perform sentiment analysis of text, speech recognition technology to analyze the tone of audio data, and computer vision technology to perform facial recognition and facial expression analysis from images and videos.

[0433] input:

[0434] Preprocessed data (tokenized text, transcribed audio, analyzed images)

[0435] output:

[0436] Analysis of user's emotional state and needs

[0437] Step 4:

[0438] The server generates customized content based on the analysis results, specifically a recommended list of entertainment content (movies, dramas, music) that corresponds to the user's emotional state and needs.

[0439] input:

[0440] Analysis of user's emotional state and needs

[0441] output:

[0442] Customized entertainment content recommendations

[0443] Step 5:

[0444] The server sends the generated customized content to the device, which displays it to the user and provides recommended content. The user then views the recommended content and provides feedback on their impressions and ratings.

[0445] input:

[0446] Customized entertainment content recommendations

[0447] output:

[0448] Device screen showing recommended content and feedback data

[0449] Step 6:

[0450] The server collects user feedback and uses it to improve the system performance. The collected feedback data is analyzed and used to retrain the emotion engine algorithm.

[0451] input:

[0452] User feedback data

[0453] output:

[0454] Improved emotion engine algorithm and system performance improvement

[0455] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0456] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0457] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.

[0458] [Second embodiment]

[0459] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.

[0460] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0461] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0462] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.

[0463] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0464] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0465] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0466] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0467] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0468] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0469] In the smart glasses 214, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0470] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."

[0471] The present invention relates to a system that uses multimodal inputs such as text, voice, images, and videos to understand the needs and emotions of users and provide customized content and solutions. The program of this system is explained below in natural language.

[0472] System Overview

[0473] This system consists of three entities: a server, a terminal, and a user. The server is responsible for data collection, preprocessing, analysis, content generation, and feedback collection and performance improvement, while the terminal functions as an interface with the user. As a user of the system, the user provides multimodal input data and receives customized content as a result.

[0474] Specific examples

[0475] For example, a case where this system is applied to customer support for small and medium-sized enterprises will be described.

[0476] 1. User Input

[0477] A user sends a text message through the chat window saying, "I don't know how to use the product," and then uploads a video showing their confused expression.

[0478] 2. Data collection

[0479] The terminal temporarily stores the text, audio, image, and video data sent by the user and transmits it to the server in real time.

[0480] 3. Data Preprocessing

[0481] The server preprocesses the various data it receives, specifically by segmenting text data, normalizing words and phrases, adjusting the sampling rate of audio data, and converting image and video formats.

[0482] 4. Emotion and Needs Analysis

[0483] The server analyzes the pre-processed data, using natural language processing technology to perform sentiment analysis of the text, speech recognition technology to convert the voice data into text and analyze its tone, and computer vision technology to perform facial recognition and facial expression analysis from image and video data to identify the user's emotional state.

[0484] 5. Customized Content Generation

[0485] Based on the analysis, the server generates customized content that addresses the user's needs and emotions. In this example, it creates a text message detailing how to use the product and a reassuring video guide.

[0486] 6. Providing Feedback

[0487] The server sends the generated content to the terminal,

[0488] The device displays this information to the user, who then takes action to resolve the problem based on the information provided.

[0489] 7. Continuous learning and improvement

[0490] The server collects feedback data from users and retrains the AI ​​model, improving the accuracy of sentiment analysis and needs prediction.

[0491] In this way, this system can quickly provide highly customized content based on the user's complex input data. This will improve the efficiency of customer service, increase customer satisfaction, and strengthen the competitiveness of small and medium-sized enterprises. Similar benefits can also be expected in the fields of education and research and development.

[0492] The processing flow will be explained below.

[0493] Step 1:

[0494] A user inputs multimodal data such as text, voice, images, and video. For example, a user may enter text such as "I don't know how to use the product" into a chat box and upload a video of a confused expression.

[0495] Step 2:

[0496] The device temporarily stores the text, voice, image, and video data entered by the user and transmits it to the server in real time. At this time, the data is temporarily cached to prevent data loss.

[0497] Step 3:

[0498] The server performs preprocessing on the received multimodal data, specifically tokenizing and normalizing text data, sample rate conversion for audio data, and format conversion for image and video data.

[0499] Step 4:

[0500] The server prepares the preprocessed data for analysis, for example, using natural language processing (NLP) to perform sentiment analysis on the text, using speech recognition technology to convert audio data into text, and using computer vision technology to perform facial recognition and facial expression analysis on image and video data.

[0501] Step 5:

[0502] The server identifies the user's emotional state and needs based on the analysis results. For example, the result of sentiment analysis may be "confused," and facial expression analysis may also detect a confused expression.

[0503] Step 6:

[0504] The server generates customized content tailored to the user based on the identified emotional state and needs, such as a text message detailing how to use a product or a reassuring video guide.

[0505] Step 7:

[0506] The server transmits the generated customized content to the terminal, compressing the data as necessary during transmission to ensure efficient data transmission.

[0507] Step 8:

[0508] The device displays the received content to the user. For example, a text message is displayed on the chat screen, and a video guidance is played on a player.

[0509] Step 9:

[0510] The user reviews the customized content provided and takes the necessary action to resolve the issue.

[0511] Step 10:

[0512] The server collects feedback data from users, including text and audio ratings.

[0513] Step 11:

[0514] The server analyzes the collected feedback, evaluates the performance of the AI ​​model, and identifies areas for improvement, which are then used for retraining to improve the accuracy of sentiment analysis and needs prediction.

[0515] In this way, through a series of steps, the system can quickly and accurately provide customized content based on the user's multiple input data.

[0516] Example 1

[0517] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0518] Many modern systems rely on single-modal input data from users (text only or voice only), which makes it difficult to fully understand users' emotions and needs, making it difficult to respond appropriately or provide customized content. Furthermore, there is a lack of means to properly collect user feedback and improve system performance, making continuous improvement difficult.

[0519] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0520] In this invention, the server includes a means for a user to input input data into the terminal, a means for the terminal to temporarily store the input data and send it to the server, and a means for the server to pre-process the collected multimodal input data. This makes it possible to analyze the user's emotions and needs from multiple angles and quickly provide highly customized content. It is also possible to collect feedback from users and continuously improve the system's performance.

[0521] A "user" is a person or organization that utilizes the system to input data and receive services.

[0522] A "terminal" is a device that allows a user to input data and send it to a server. Examples include computers, smartphones, and tablets.

[0523] "Server" means a device or system that receives data sent from a terminal, pre-processes and analyzes it, and generates and provides customized content.

[0524] "Multimodal input data" refers to data in multiple formats, including text data, audio data, image data, and video data.

[0525] "Preprocessing" is the process of preparing received data so that it can be easily analyzed, since it is difficult to handle as it is. This includes morphological analysis and normalization of text, adjusting the sampling rate of audio, and converting the format of images and videos.

[0526] "Emotion and needs analysis" refers to identifying a user's emotional state and the information and solutions they need based on pre-processed data. This analysis uses natural language processing, voice recognition, facial recognition, and facial expression analysis technologies.

[0527] "Customized content" refers to information or guidance generated based on the analysis results to address the user's individual needs and emotions, such as text messages or video guidance.

[0528] "Feedback" refers to the reactions and evaluations given by users to the content provided, and is information that can be used to improve the performance of the system.

[0529] "Natural language processing technology" is a technology that processes and analyzes human language using a computer. It includes morphological analysis and sentiment analysis.

[0530] "Speech recognition technology" is a technology that converts voice data into text data and analyzes the content and tone of the text.

[0531] "Facial recognition technology" is a technology that detects faces from images and videos and analyzes their features.

[0532] "Facial expression analysis technology" is a technology that uses facial recognition technology to identify a user's emotional state from their facial expression.

[0533] A "generative AI model" is an artificial intelligence model that generates new information based on given data. Examples include GPT-3.

[0534] A "prompt sentence" is the text of an instruction or question that a user enters into a system.

[0535] MODE FOR CARRYING OUT THE INVENTION

[0536] The present invention relates to a system that uses multimodal input data such as text, audio, images, and videos to understand the needs and emotions of users and provide customized content. Specific embodiments of this system are described below.

[0537] System Overview

[0538] This system consists of three entities: a server, a terminal, and a user. The server is responsible for data collection, preprocessing, analysis, generation of customized content, feedback collection, and performance improvement. The terminal functions as an interface with the user, who, as a user of the system, provides multimodal input data and receives customized content as a result.

[0539] Hardware and software used

[0540] Hardware: Server machines, user devices (computers, smartphones, tablets, etc.)

[0541] software:

[0542] Natural language processing libraries: NLTK, spaCy

[0543] Speech Recognition Library: Google Speech-to-Text API

[0544] Image processing library: OpenCV

[0545] Facial Expression Recognition API:Facial Expression Recognition API

[0546] Generative AI model: GPT-3

[0547] User Input

[0548] Users can use the chat window on their device to type text messages, or upload audio messages and video files. For example, a user might type "I don't know how to use this product" and attach a video of themselves looking confused.

[0549] Prompt Sentence Examples

[0550] I don't know how to use the new product. How do I set it up?

[0551] Video file example

[0552] Approximately 30 seconds of video file (mp4 format) containing a confused expression

[0553] Data collection and transmission

[0554] The device temporarily stores the text, audio, image, and video data entered by the user and transmits it to the server in real time. Specifically, the device transmits this data to the server using an HTTP POST request.

[0555] Data Preprocessing

[0556] The server preprocesses the various types of data it receives. For text data, it uses a natural language processing library (e.g., NLTK or spaCy) to perform morphological analysis and normalization. For audio data, it adjusts the sampling rate and removes noise. For image and video data, it converts them into an appropriate format using OpenCV and performs preprocessing for face detection.

[0557] Emotion and Needs Analysis

[0558] The server analyzes the preprocessed data. For text data, it performs sentiment analysis using models such as the BERT model. For audio data, it converts the audio to text using the Google Speech-to-Text API and analyzes its tone. For image and video data, it uses the Facial Expression Recognition API for facial recognition and expression analysis.

[0559] Customized content generation

[0560] Based on the analysis results, the server generates customized content that addresses the user's needs and emotions. Using a generative AI model (e.g., GPT-3), it creates text messages detailing how to use the product and reassuring video guidance.

[0561] Providing content and gathering feedback

[0562] The server sends the generated content to the device, which displays it to the user. The user then takes action to solve the problem based on the information provided. The server also collects feedback data from users to continuously improve the system's performance.

[0563] In this way, this system can quickly provide highly customized content based on the user's complex input data, resulting in more efficient customer service and improved customer satisfaction, thereby enhancing a company's competitiveness.

[0564] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0565] Step 1:

[0566] The user provides multimodal input data such as text, voice, images, and video to the device. The input data includes a prompt such as "I don't know how to use this new product. How do I set it up?" and a video file (mp4 format) containing a confused facial expression. The device receives this data based on the user's input.

[0567] Step 2:

[0568] The device temporarily stores the received text, audio, image, and video data and sends it to the server in real time using an HTTP POST request. Input data is stored and sent as is.

[0569] Step 3:

[0570] The server preprocesses the received multimodal input data, specifically by:

[0571] Text data: Perform morphological analysis and normalization using a natural language processing library (e.g., NLTK, spaCy).

[0572] Audio data: Adjust the sampling rate and remove noise.

[0573] Image and video data: Use OpenCV to convert them into the appropriate format and pre-process them for face detection.

[0574] The output is preprocessed text, clear audio data, and format-converted images and videos.

[0575] Step 4:

[0576] The server analyzes the preprocessed data, specifically:

[0577] Text data: Sentiment analysis is performed using models such as BERT.

[0578] Audio data: Converted to text using the Google Speech-to-Text API, and tone analysis is performed based on that text.

[0579] Image and video data: Use the Facial Expression Recognition API for facial recognition and facial expression analysis.

[0580] The output is an analysis that indicates the user's emotional state and needs.

[0581] Step 5:

[0582] The server generates customized content based on the analysis results. Using a generative AI model (e.g., GPT-3), it creates text messages with specific actions to take based on the user's needs and reassuring video guidance. The output is a customized text message and video guidance.

[0583] Step 6:

[0584] The server then sends the generated customized content to the device, which then receives it and displays it to the user. Specifically, the device displays text messages in a chat window and plays video guidance. The user can then take action to resolve the problem based on this content.

[0585] Step 7:

[0586] The server collects feedback data from users. Specifically, it collects the ratings and opinions users have given about the content provided and sends that data to the server. The server then retrains the generative AI model based on the collected feedback data, continuously improving the system's performance. This improves the accuracy of user sentiment analysis and need prediction, providing a better user experience.

[0587] (Application example 1)

[0588] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0589] Conventional customer support systems are unable to fully utilize the complex input data (text, voice, images, video, etc.) from users, making it difficult to accurately grasp users' emotions and specific needs. Food delivery services, in particular, are required to provide quick and accurate solutions to issues such as cold food, but current systems have limitations. Therefore, there is an urgent need to develop a system that can analyze user input data at multiple stages and provide appropriate solutions.

[0590] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0591] In this invention, the server includes means for collecting multimodal input data, means for preprocessing the collected multimodal input data, means for analyzing emotions and needs using the preprocessed data, means for generating customized content based on the analysis results, means for providing the generated content to users, means for collecting feedback from users and improving system performance, and means for analyzing text and video data entered by users and providing corresponding solutions, including generating guide videos including instructions on how to reheat food when it has cooled and an apology message. This makes it possible to quickly and accurately solve specific user problems and improve customer satisfaction.

[0592] "Multimodal input data" refers to data in multiple different formats, such as text data, audio data, image data, and video data.

[0593] "Means of collection" refers to the mechanisms and technologies for acquiring multimodal input data provided by users.

[0594] "Preprocessing means" refers to the techniques and algorithms used to prepare collected multimodal input data in an analyzable form.

[0595] "Means for emotion and needs analysis" refers to analytical tools and techniques for discovering user emotions and specific needs based on processed data.

[0596] "Means for generating customized content" refers to technology that creates specific information or guides tailored to the user's needs based on the analysis results.

[0597] "Means for providing" refers to the mechanisms and technologies for displaying or notifying the user of the generated customized content.

[0598] "Means of collecting feedback" refers to methods and techniques for obtaining responses and reactions from users and using them to improve the system.

[0599] "Means of providing solutions" refers to mechanisms and technologies for presenting users with specific solutions and proposals for solving problems or issues.

[0600] "Means for generating guide videos" refers to technology for creating video content that includes guidance and instructions that meet the needs of users.

[0601] System configuration

[0602] This system is composed of three elements: a server, a terminal, and a user. Each element is explained in detail below.

[0603] server

[0604] The server is responsible for data collection, preprocessing, analysis, content generation, and feedback collection and performance improvement. Specific software used includes TensorFlow, PyTorch, OpenCV, NLTK, and Transformer-based NLP models (e.g., BERT).

[0605] Terminal

[0606] The terminal is a device such as a smartphone that functions as an interface with the user. Specifically, it is built using React Native and plays a role in sending data entered by the user to the server and displaying the content received from the server.

[0607] User

[0608] Users are the users of the system and provide multimodal input data (text, audio, images, video), which is analyzed and customized content is returned to the user.

[0609] Processing flow

[0610] 1. User input:

[0611] The user inputs a text message and uploads video data via the device. For example, the user can provide a message such as "My food has arrived but it's cold and I'm worried" along with a video of the food being cold.

[0612] 2. Data Collection:

[0613] The device temporarily stores the text and video data sent by the user and then transmits it to the server. Data transmission is performed using real-time communication.

[0614] 3. Data preprocessing:

[0615] The server preprocesses the received data: text data is segmented using NLTK, and video data is converted to the appropriate format using OpenCV.

[0616] 4. Emotion and Needs Analysis:

[0617] The server analyzes the preprocessed data: text data is subjected to sentiment analysis using natural language processing techniques such as the BERT model, and video data is subjected to facial recognition and facial expression analysis using OpenCV.

[0618] 5. Customized Content Generation:

[0619] Based on the analysis results, it generates customized content that addresses the user's needs and emotions. In this example, it generates a guide video explaining how to reheat cold food and an apology message.

[0620] 6. Providing Feedback:

[0621] The generated content is sent to the device, which displays it to the user, who then takes action to solve the problem based on the information provided.

[0622] 7. Continuous learning and improvement:

[0623] The server collects feedback data from users and retrains the AI ​​model, improving the accuracy of sentiment analysis and needs prediction.

[0624] Examples and prompts

[0625] Specific examples

[0626] If a user texts, "My food arrived but it's cold and I'm worried," and uploads a video showing the food cold, the system will provide a video guide on how to reheat it and an apology message.

[0627] Prompt Sentence Examples

[0628] User input: The text "My food arrived but it's cold and I'm worried" and a video of the food being cold.

[0629] Desired output: An apology message and a guided video with specific steps for reheating the food.

[0630] The system of the present invention makes it possible to quickly and accurately solve specific user problems and improve customer satisfaction.

[0631] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0632] Step 1:

[0633] User input:

[0634] The user inputs a text message and uploads video data through the device. For example, the user can provide a text message saying, "My food arrived, but it's cold and I'm worried," along with a video of the cold food. The device then receives this input and prepares it for the next step.

[0635] input:

[0636] Text messages, video data

[0637] output:

[0638] Temporarily saved text messages and video data

[0639] Step 2:

[0640] Data collection:

[0641] The device temporarily stores the text and video data sent by the user and transmits it to the server in real time. This data transmission requires a stable internet connection.

[0642] input:

[0643] Temporarily saved text messages and video data

[0644] output:

[0645] Text messages and video data sent to the server

[0646] Step 3:

[0647] Data preprocessing:

[0648] The server preprocesses the received data: text data is segmented using natural language processing technology (NLTK), and video data is converted into a format that can be analyzed frame by frame using OpenCV.

[0649] input:

[0650] Text messages and video data sent to the server

[0651] output:

[0652] Segmented text data, converted video data

[0653] Step 4:

[0654] Emotion and Needs Analysis:

[0655] The server analyzes emotions and needs based on the preprocessed data. It uses the BERT model for sentiment analysis on text data, and OpenCV for facial recognition and facial expression analysis on video data to identify the user's emotional state.

[0656] input:

[0657] Segmented text data, converted video data

[0658] output:

[0659] Analyzed emotion and needs data

[0660] Step 5:

[0661] Customized Content Generation:

[0662] Based on the analysis results, the system generates customized content that responds to the user's needs and emotions, such as a guide video explaining how to reheat cold food and an apology message.

[0663] input:

[0664] Analyzed emotion and needs data

[0665] output:

[0666] Guide video, apology message

[0667] Step 6:

[0668] Providing feedback:

[0669] The server sends the generated content to the device, which displays it to the user, who then acts to solve the problem based on the information provided.

[0670] input:

[0671] Guide video, apology message

[0672] output:

[0673] Guide video and apology message displayed to users

[0674] Step 7:

[0675] Continuous learning and improvement:

[0676] The server collects feedback data from users and uses it to retrain the AI ​​model, improving the accuracy of sentiment analysis and needs prediction.

[0677] input:

[0678] User feedback data

[0679] output:

[0680] Improved AI models

[0681] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[0682] This invention relates to a system that provides customized content by utilizing multimodal inputs such as text, voice, images, and video, and by performing advanced analysis of user needs and emotions. In particular, by combining an emotion engine, it is possible to accurately recognize the user's emotional state and generate appropriate responses.

[0683] System Overview

[0684] This system is composed of three entities: a server, a terminal, and a user. The server collects data, preprocesses it, analyzes it, generates customized content, and collects feedback and improves performance, while the terminal functions as an interface with the user. As a user of the system, the user provides multimodal input data and receives customized content based on the analysis results.

[0685] Incorporating an emotion engine

[0686] One of the features of this system is the incorporation of an emotion engine, which integrates natural language processing (NLP), speech recognition, facial recognition, and facial expression analysis technologies to analyze the user's emotions in real time.

[0687] Specific examples

[0688] The application of this system to customer support for small and medium-sized enterprises will be explained.

[0689] 1. User Input

[0690] A user texts the support chat saying, "I don't know how to use the product," and then uploads a video showing a confused expression.

[0691] 2. Data collection

[0692] The device temporarily stores the multimodal data sent by the user, including text, audio, images, and video, and transmits it to the server in real time. The data is temporarily cached to prevent data loss.

[0693] 3. Data Preprocessing

[0694] The server preprocesses the various data it receives, specifically tokenizing and normalizing text data, converting the sampling rate of audio data, and converting the format of image and video data.

[0695] 4. Emotion and Needs Analysis

[0696] The emotion engine uses the pre-processed data to analyze the user's emotional state and needs, using natural language processing technology to perform sentiment analysis of the text, speech recognition technology to convert the voice data into text and analyze its tone, and computer vision technology to perform facial recognition and facial expression analysis from images and videos.

[0697] 5. Customized Content Generation

[0698] Based on the analysis, the server generates customized content that responds to the user's emotional state and needs, such as a text message detailing how to use the product and a reassuring video guide.

[0699] 6. Providing Feedback

[0700] The server transmits the generated customized content to the terminal, compressing the data as necessary to transmit the data efficiently.

[0701] The device displays the received content to the user, displaying text messages on the screen and playing video guidance.

[0702] 7. Continuous learning and improvement

[0703] The server collects feedback from the user, which may include a text or audio rating.

[0704] The server analyzes the collected feedback and retrains the emotion engine algorithm, thereby improving the system's accuracy and response quality.

[0705] This system can accurately grasp the user's emotional state and needs in real time based on a variety of input data, and quickly provide highly customized content. This can improve the efficiency of customer service and customer satisfaction, strengthening the competitiveness of small and medium-sized enterprises. Similar benefits can be expected in the fields of education and research and development.

[0706] The processing flow will be explained below.

[0707] Step 1:

[0708] A user inputs multimodal data such as text, audio, images, and video. For example, a user texts a support chat message saying, "I don't know how to use the product," and uploads a video showing a confused expression.

[0709] Step 2:

[0710] The device temporarily stores the text, voice, image, and video data entered by the user and transmits it to the server in real time, using a cache process to prevent data loss.

[0711] Step 3:

[0712] The server preprocesses the multimodal data it receives, specifically tokenizing and normalizing text data, converting the sampling rate of audio data, and converting the format of image and video data.

[0713] Step 4:

[0714] The emotion engine analyzes the pre-processed data, using natural language processing technology to perform sentiment analysis on the text data, speech recognition technology to convert the voice data into text data and analyze its tone, and computer vision technology to perform facial recognition and facial expression analysis on image and video data to identify the user's emotional state.

[0715] Step 5:

[0716] The server identifies the user's needs and emotional state based on the analysis results from the emotion engine. For example, if the text data is judged as "confused," a confused expression is also detected in the facial expression analysis.

[0717] Step 6:

[0718] The server generates customized content based on the identified needs and emotional state, for example, creating a text message detailing how to use a product or generating a reassuring video guide.

[0719] Step 7:

[0720] The server transmits the generated customized content to the terminal, compressing the data as necessary to improve the efficiency of data transmission.

[0721] Step 8:

[0722] The device displays the received customized content to the user, displaying text messages on the screen and playing video guidance.

[0723] Step 9:

[0724] The user reviews the customized content provided and takes action to resolve the issue.

[0725] Step 10:

[0726] The server collects feedback from the user, which may be received in text or audio format.

[0727] Step 11:

[0728] The server analyzes the feedback data and retrains it to improve the performance of the emotion engine and the overall system, thereby increasing the accuracy of future analysis and content generation.

[0729] In this way, incorporating an emotion engine makes it possible to perform advanced analysis of various user data and provide customized content tailored to needs in real time. This will significantly improve the quality and efficiency of customer service, helping to strengthen the competitiveness of small and medium-sized enterprises.

[0730] Example 2

[0731] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0732] Conventional customer support systems have struggled to efficiently collect and analyze diverse user input data and provide customized content tailored to the user's emotional state and needs. While systems that utilize multimodal input data to accurately grasp the user's emotions and generate and deliver appropriate responses are particularly needed, few such systems exist. This has led to lower user satisfaction and negatively impacted the efficiency of companies' customer support efforts.

[0733] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[0734] In this invention, the server includes: means for a user to provide multimodal input data including text, voice, images, and videos; means for temporarily storing the received multimodal input data and transmitting it to the server in real time; data preprocessing means for tokenizing, normalizing, formatting, and sampling rate converting the data received by the server; means for an emotion engine to analyze the user's emotions and needs using natural language processing, speech recognition, face recognition, and facial expression analysis techniques using the preprocessed data; means for generating customized content based on the analysis results; means for transmitting the generated content to a terminal and providing it to the user; and means for collecting user feedback and retraining the system's emotion engine to improve performance. This makes it possible to efficiently collect and analyze a variety of user input data and provide appropriate customized content in real time, thereby improving customer service efficiency and user satisfaction.

[0735] "Multimodal input data" refers to data in multiple formats, such as text, audio, images, and video.

[0736] "Preprocessing" refers to processing such as tokenization, normalization, format conversion, and sampling rate conversion to make data easier to analyze.

[0737] An "emotion engine" refers to a system that analyzes a user's emotional state by integrating natural language processing technology, voice recognition technology, facial recognition technology, and facial expression analysis technology.

[0738] "Natural language processing technology" refers to technology that enables computers to understand, analyze, and generate human language.

[0739] "Speech recognition technology" refers to the technology that converts speech into text data and analyzes its content.

[0740] "Facial recognition technology" refers to technology that detects and identifies human faces from images and videos.

[0741] "Facial expression analysis technology" refers to technology that analyzes facial expressions and identifies their emotional state.

[0742] "Customized content" refers to individually optimized information and responses generated based on a user's emotional state and needs.

[0743] "Feedback" refers to subsequent data such as ratings, opinions, and impressions collected from users.

[0744] "Retraining" refers to the process of retraining an existing learning model with new data to improve its performance.

[0745] "Server" refers to the central management system that collects, pre-processes, analyzes, generates content, gathers feedback, and improves performance of data.

[0746] "Terminal" refers to a device that acts as an interface with a user and is responsible for inputting and outputting data.

[0747] This invention relates to a system that uses multimodal input data such as text, voice, images, and video to provide customized content through advanced analysis of user needs and emotions. In particular, by combining an emotion engine, it becomes possible to accurately recognize the user's emotional state and generate appropriate responses. This system is composed of three entities: a "server," a "terminal," and a "user."

[0748] The server collects, preprocesses, and analyzes data, generates customized content, and collects feedback to improve performance. The terminal functions as an interface with the user, who provides multimodal input data and receives customized content based on the analysis results.

[0749] Hardware and software used

[0750] The servers are data centers or cloud computing services with high-performance computing capabilities, and the software uses the BERT model for natural language processing (NLP), the Google Speech-to-Text API for speech recognition, and OpenCV and Dlib libraries for facial recognition and facial expression analysis.

[0751] Specific examples

[0752] We will explain how this system can be applied to customer support for small and medium-sized enterprises.

[0753] User Input

[0754] A user can text the support chat saying, "I don't know how to use the product," and then upload a video showing a confused expression, which inputs the user's needs and emotions into the system.

[0755] Data collection

[0756] The device temporarily stores the multimodal data received from the user, including text, audio, images, and video, and transmits it to the server in real time. The data is temporarily cached to prevent data loss.

[0757] Data Preprocessing

[0758] The server performs preprocessing on the received data, including tokenization, normalization, format conversion, and sampling rate conversion. Specifically, text data is processed using the BERT model, and the sampling rate of audio data is converted to 16 kHz to optimize it for analysis. For image and video data, the resolution is changed to 1280x720 pixels using the OpenCV library and the data is converted to the required format.

[0759] Emotion and Needs Analysis

[0760] The emotion engine uses the pre-processed data to analyze the user's emotional state and needs. It uses natural language processing technology to perform sentiment analysis of the text, speech recognition technology to convert the voice data into text and analyze its tone, and computer vision technology to perform facial recognition and facial expression analysis from images and videos to understand the user's emotions.

[0761] Customized Content Generation

[0762] Based on the analysis results, the server generates customized content that corresponds to the user's emotional state and needs. For example, based on the emotion analysis result of "confused," a text message explaining detailed product usage and a reassuring video guide are generated.

[0763] Providing feedback

[0764] The server transmits the generated customized content to the terminal, which compresses the content as necessary for efficient transmission. The terminal then displays the received content to the user, including displaying a text message on the screen and playing a video guide.

[0765] Continuous learning and improvement

[0766] The server collects feedback from users, including text and voice ratings, and analyzes the collected feedback and re-trains the emotion engine using TensorFlow to improve the system's accuracy and response quality.

[0767] Prompt Sentence Examples

[0768] "If a customer is confused about how to use a product, use an emotion engine to analyze their emotions in real time and generate a text message explaining how to use the product and a reassuring video guide."

[0769] This makes it possible to efficiently collect and analyze a wide range of user input data and provide appropriate customized content in real time. This system will improve the efficiency of customer service and user satisfaction, strengthening the competitiveness of small and medium-sized enterprises. Similar benefits can be expected in the fields of education and research and development.

[0770] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0771] Step 1:

[0772] The user types a text message into the support chat and uploads a video showing a confused expression. The input data includes four types: text, audio, images, and video. This allows the user's needs and emotions to be input into the system. For example, the user can send a specific text message such as, "I don't know how to use the product."

[0773] Input: Text message, video (confused expression)

[0774] Output: Collected multimodal data

[0775] Step 2:

[0776] The device temporarily stores multimodal data received from the user, including text, audio, images, and video. The stored data is temporarily cached and then sent to the server in real time. This prevents data loss and enables analysis on the server side.

[0777] Input: Collected multimodal data

[0778] Output: Multimodal data sent to the server

[0779] Step 3:

[0780] The server preprocesses the received data. For text data, tokenization and normalization are performed using the BERT model. For audio data, the sampling rate is converted to 16 kHz to optimize the data for analysis. For image and video data, the resolution is changed to 1280x720 pixels using the OpenCV library and the data is converted to the required format.

[0781] Input: Multimodal data sent to the server

[0782] Output: Preprocessed data (tokenized, normalized, formatted, and sample rate converted)

[0783] Step 4:

[0784] The emotion engine uses the pre-processed data to analyze the user's emotional state and needs. It uses natural language processing technology to perform sentiment analysis of the text, speech recognition technology to convert the voice data into text and analyze its tone, and computer vision technology to perform facial recognition and facial expression analysis from images and videos to understand the user's emotions. The results of this analysis determine the user's specific emotions and needs.

[0785] Input: Preprocessed data

[0786] Output: User's emotional state and needs analysis results

[0787] Step 5:

[0788] Based on the analysis results, the server generates customized content that corresponds to the user's emotional state and needs. For example, based on the emotion analysis result of "confused," a text message explaining detailed product usage and a reassuring video guide are generated.

[0789] Input: User's emotional state and needs analysis results

[0790] Output: Customized content (text messages, video guidance)

[0791] Step 6:

[0792] The server transmits the generated customized content to the terminal, which compresses the content as necessary for efficient transmission. The terminal then displays the received content to the user, including displaying a text message on the screen and playing a video guide.

[0793] Input:Customized Content

[0794] Output: Customized content served to the user

[0795] Step 7:

[0796] The server collects feedback from users, including text and voice ratings, and analyzes the collected feedback and re-trains the emotion engine using TensorFlow to improve the system's accuracy and response quality.

[0797] Input: User feedback

[0798] Output: Performance improvement by retraining the emotion engine

[0799] The above are the specific processing steps of the program in this system.

[0800] (Application example 2)

[0801] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0802] Current content distribution services struggle to analyze a user's emotional state and needs in real time and provide appropriate entertainment content. Conventional systems often fail to provide personalized recommendations based on the user's emotions and state, resulting in reduced user satisfaction. The present invention aims to address these challenges and provide optimal content tailored to the user's emotional state and needs.

[0803] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[0804] In this invention, the server includes means for collecting multimodal input data, means for preprocessing the collected multimodal input data, means for analyzing emotions and needs using the preprocessed data, means for generating customized content based on the analysis results, means for providing the generated content to the user, means for collecting feedback from the user and improving the performance of the system, and means for recommending entertainment content based on the emotional state and needs of the user, thereby making it possible to provide personalized entertainment content according to the emotional state and needs of the user in real time.

[0805] "Multimodal input data" refers to data in multiple formats, such as text data, audio data, image data, and video data.

[0806] "Means for collection" refers to the devices and software that obtain and store multimodal input data from users.

[0807] "Preprocessing means" refers to devices or software that perform processes to convert collected multimodal input data into a format that is easy to analyze.

[0808] "Emotion and needs analysis means" refers to algorithms or software that identify a user's emotional state and needs based on pre-processed data.

[0809] "Means for generating customized content" refers to systems or software for creating optimal content for users based on the analysis results.

[0810] The "means for providing to the user" refers to a device or interface for showing the generated customized content to the user.

[0811] "Means for collecting user feedback" refers to devices or software that collect and record user ratings and opinions.

[0812] "Means for improving system performance" refers to methods for improving the behavior of algorithms or software based on collected feedback.

[0813] "Emotional state" refers to data or states that represent a user's mental and emotional state.

[0814] "Needs" refer to the requests and desires for information, functions, and services that users desire.

[0815] "Entertainment content" refers to various media content for entertainment purposes, such as movies, dramas, music, and games.

[0816] The present invention relates to a system for providing customized entertainment content based on a user's emotional state and needs, which is implemented in the following steps:

[0817] System configuration

[0818] This system consists of three components: a server, a terminal, and a user. The server collects data, preprocesses it, analyzes it, generates content, and collects feedback to improve performance. The terminal functions as an interface with the user. As a user of the system, the user provides multimodal input data and receives customized content based on the analysis results.

[0819] Hardware and software used

[0820] This system uses the following hardware and software:

[0821] Hardware:

[0822] Device: Smartphone or head-mounted display equipped with a camera and microphone

[0823] software:

[0824] Cloud services: Cloud services for data processing (e.g., AWS, Google Cloud)

[0825] Natural language processing engine: Google NLP

[0826] Speech recognition engine: Amazon Transcribe

[0827] Facial expression and face recognition engine: Microsoft Azure Face API

[0828] Data processing and calculation

[0829] The device captures multimodal input data (text, voice, image, and video) from the user. This data is temporarily cached and sent to the server. The server preprocesses the received data, tokenizing and normalizing the text data. The voice data undergoes sampling rate conversion and text conversion, and the image and video data undergo format conversion and facial expression analysis.

[0830] The following techniques are used for the analysis:

[0831] Sentiment analysis of text using natural language processing (NLP) technology

[0832] Voice recognition technology converts voice data into text and analyzes the tone

[0833] Uses computer vision technology to perform facial recognition and facial expression analysis

[0834] Content generation and delivery

[0835] Based on the analysis results, the server generates a recommendation list of entertainment content (movies, dramas, music) according to the user's emotional state and needs. This recommendation list is customized and accessible to the user via smartphone or head-mounted display. The user can view the recommended content and provide feedback.

[0836] Specific examples

[0837] For example, if a user uses a smartphone to type the text "I'm feeling a bit down today," input a low-pitched voice into the microphone, and capture an image of a gloomy facial expression with the camera, all of this data is immediately sent to the server and analyzed.

[0838] The server analyzes the user's emotional state (feeling depressed) and needs (content that will lift their spirits) and recommends the most suitable movies and music for the user.

[0839] Prompt Sentence Examples

[0840] If you send "I'm feeling a bit down today," enter the following prompt into the model:

[0841] text

[0842] User sent text: "I'm feeling a bit down today."

[0843] Voice data: User's voice is low and restless

[0844] Facial expression data: The user has a dull expression

[0845] Use this data to generate content recommendations for your users.

[0846] This makes it possible to provide entertainment content in real time that matches the user's emotional state and needs, thereby increasing user satisfaction and improving the value of content distribution services.

[0847] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0848] Step 1:

[0849] The device collects the user's input data. Specifically, the device's camera and microphone are used to capture the user's facial expressions and voice in real time, and the user inputs text. The input data at this time includes text data, voice data, and image data. This data is temporarily stored on the device and then sent to the server.

[0850] input:

[0851] Text data: Character information entered by the user

[0852] Voice data: recordings of your voice

[0853] Image data: photos or videos of the user's facial expressions

[0854] output:

[0855] Capture and temporary storage of multimodal input data (text, audio, images)

[0856] Step 2:

[0857] The server preprocesses the received multimodal input data, specifically tokenizing and normalizing text data, converting audio data to text through sampling rate conversion, and formatting and analyzing facial expressions on image and video data.

[0858] input:

[0859] Multimodal input data (text, audio, images)

[0860] output:

[0861] Preprocessed data (tokenized text, transcribed audio, analyzed images)

[0862] Step 3:

[0863] The server's emotion engine uses the pre-processed data to analyze emotions and needs, using natural language processing (NLP) technology to perform sentiment analysis of text, speech recognition technology to analyze the tone of audio data, and computer vision technology to perform facial recognition and facial expression analysis from images and videos.

[0864] input:

[0865] Preprocessed data (tokenized text, transcribed audio, analyzed images)

[0866] output:

[0867] Analysis of user's emotional state and needs

[0868] Step 4:

[0869] The server generates customized content based on the analysis results, specifically a recommended list of entertainment content (movies, dramas, music) that corresponds to the user's emotional state and needs.

[0870] input:

[0871] Analysis of user's emotional state and needs

[0872] output:

[0873] Customized entertainment content recommendations

[0874] Step 5:

[0875] The server sends the generated customized content to the device, which displays it to the user and provides recommended content. The user then views the recommended content and provides feedback on their impressions and ratings.

[0876] input:

[0877] Customized entertainment content recommendations

[0878] output:

[0879] Device screen showing recommended content and feedback data

[0880] Step 6:

[0881] The server collects user feedback and uses it to improve the system performance. The collected feedback data is analyzed and used to retrain the emotion engine algorithm.

[0882] input:

[0883] User feedback data

[0884] output:

[0885] Improved emotion engine algorithm and system performance improvement

[0886] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0887] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0888] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.

[0889] [Third embodiment]

[0890] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.

[0891] 5, the data processing system 310 includes the data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.

[0892] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0893] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.

[0894] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0895] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0896] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0897] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0898] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0899] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0900] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0901] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."

[0902] The present invention relates to a system that uses multimodal inputs such as text, voice, images, and videos to understand the needs and emotions of users and provide customized content and solutions. The program of this system is explained below in natural language.

[0903] System Overview

[0904] This system consists of three entities: a server, a terminal, and a user. The server is responsible for data collection, preprocessing, analysis, content generation, and feedback collection and performance improvement, while the terminal functions as an interface with the user. As a user of the system, the user provides multimodal input data and receives customized content as a result.

[0905] Specific examples

[0906] For example, a case where this system is applied to customer support for small and medium-sized enterprises will be described.

[0907] 1. User Input

[0908] A user sends a text message through the chat window saying, "I don't know how to use the product," and then uploads a video showing their confused expression.

[0909] 2. Data collection

[0910] The terminal temporarily stores the text, audio, image, and video data sent by the user and transmits it to the server in real time.

[0911] 3. Data Preprocessing

[0912] The server preprocesses the various data it receives, specifically by segmenting text data, normalizing words and phrases, adjusting the sampling rate of audio data, and converting image and video formats.

[0913] 4. Emotion and Needs Analysis

[0914] The server analyzes the pre-processed data, using natural language processing technology to perform sentiment analysis of the text, speech recognition technology to convert the voice data into text and analyze its tone, and computer vision technology to perform facial recognition and facial expression analysis from image and video data to identify the user's emotional state.

[0915] 5. Customized Content Generation

[0916] Based on the analysis, the server generates customized content that addresses the user's needs and emotions. In this example, it creates a text message detailing how to use the product and a reassuring video guide.

[0917] 6. Providing Feedback

[0918] The server sends the generated content to the terminal,

[0919] The device displays this information to the user, who then takes action to resolve the problem based on the information provided.

[0920] 7. Continuous learning and improvement

[0921] The server collects feedback data from users and retrains the AI ​​model, improving the accuracy of sentiment analysis and needs prediction.

[0922] In this way, this system can quickly provide highly customized content based on the user's complex input data. This will improve the efficiency of customer service, increase customer satisfaction, and strengthen the competitiveness of small and medium-sized enterprises. Similar benefits can also be expected in the fields of education and research and development.

[0923] The processing flow will be explained below.

[0924] Step 1:

[0925] A user inputs multimodal data such as text, voice, images, and video. For example, a user may enter text such as "I don't know how to use the product" into a chat box and upload a video of a confused expression.

[0926] Step 2:

[0927] The device temporarily stores the text, voice, image, and video data entered by the user and transmits it to the server in real time. At this time, the data is temporarily cached to prevent data loss.

[0928] Step 3:

[0929] The server performs preprocessing on the received multimodal data, specifically tokenizing and normalizing text data, sample rate conversion for audio data, and format conversion for image and video data.

[0930] Step 4:

[0931] The server prepares the preprocessed data for analysis, for example, using natural language processing (NLP) to perform sentiment analysis on the text, using speech recognition technology to convert audio data into text, and using computer vision technology to perform facial recognition and facial expression analysis on image and video data.

[0932] Step 5:

[0933] The server identifies the user's emotional state and needs based on the analysis results. For example, the result of sentiment analysis may be "confused," and facial expression analysis may also detect a confused expression.

[0934] Step 6:

[0935] The server generates customized content tailored to the user based on the identified emotional state and needs, such as a text message detailing how to use a product or a reassuring video guide.

[0936] Step 7:

[0937] The server transmits the generated customized content to the terminal, compressing the data as necessary during transmission to ensure efficient data transmission.

[0938] Step 8:

[0939] The device displays the received content to the user. For example, a text message is displayed on the chat screen, and a video guidance is played on a player.

[0940] Step 9:

[0941] The user reviews the customized content provided and takes the necessary action to resolve the issue.

[0942] Step 10:

[0943] The server collects feedback data from users, including text and audio ratings.

[0944] Step 11:

[0945] The server analyzes the collected feedback, evaluates the performance of the AI ​​model, and identifies areas for improvement, which are then used for retraining to improve the accuracy of sentiment analysis and needs prediction.

[0946] In this way, through a series of steps, the system can quickly and accurately provide customized content based on the user's multiple input data.

[0947] Example 1

[0948] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[0949] Many modern systems rely on single-modal input data from users (text only or voice only), which makes it difficult to fully understand users' emotions and needs, making it difficult to respond appropriately or provide customized content. Furthermore, there is a lack of means to properly collect user feedback and improve system performance, making continuous improvement difficult.

[0950] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0951] In this invention, the server includes a means for a user to input input data into the terminal, a means for the terminal to temporarily store the input data and send it to the server, and a means for the server to pre-process the collected multimodal input data. This makes it possible to analyze the user's emotions and needs from multiple angles and quickly provide highly customized content. It is also possible to collect feedback from users and continuously improve the system's performance.

[0952] A "user" is a person or organization that utilizes the system to input data and receive services.

[0953] A "terminal" is a device that allows a user to input data and send it to a server. Examples include computers, smartphones, and tablets.

[0954] "Server" means a device or system that receives data sent from a terminal, pre-processes and analyzes it, and generates and provides customized content.

[0955] "Multimodal input data" refers to data in multiple formats, including text data, audio data, image data, and video data.

[0956] "Preprocessing" is the process of preparing received data so that it can be easily analyzed, since it is difficult to handle as it is. This includes morphological analysis and normalization of text, adjusting the sampling rate of audio, and converting the format of images and videos.

[0957] "Emotion and needs analysis" refers to identifying a user's emotional state and the information and solutions they need based on pre-processed data. This analysis uses natural language processing, voice recognition, facial recognition, and facial expression analysis technologies.

[0958] "Customized content" refers to information or guidance generated based on the analysis results to address the user's individual needs and emotions, such as text messages or video guidance.

[0959] "Feedback" refers to the reactions and evaluations given by users to the content provided, and is information that can be used to improve the performance of the system.

[0960] "Natural language processing technology" is a technology that processes and analyzes human language using a computer. It includes morphological analysis and sentiment analysis.

[0961] "Speech recognition technology" is a technology that converts voice data into text data and analyzes the content and tone of the text.

[0962] "Facial recognition technology" is a technology that detects faces from images and videos and analyzes their features.

[0963] "Facial expression analysis technology" is a technology that uses facial recognition technology to identify a user's emotional state from their facial expression.

[0964] A "generative AI model" is an artificial intelligence model that generates new information based on given data. Examples include GPT-3.

[0965] A "prompt sentence" is the text of an instruction or question that a user enters into a system.

[0966] MODE FOR CARRYING OUT THE INVENTION

[0967] The present invention relates to a system that uses multimodal input data such as text, audio, images, and videos to understand the needs and emotions of users and provide customized content. Specific embodiments of this system are described below.

[0968] System Overview

[0969] This system consists of three entities: a server, a terminal, and a user. The server is responsible for data collection, preprocessing, analysis, generation of customized content, feedback collection, and performance improvement. The terminal functions as an interface with the user, who, as a user of the system, provides multimodal input data and receives customized content as a result.

[0970] Hardware and software used

[0971] Hardware: Server machines, user devices (computers, smartphones, tablets, etc.)

[0972] software:

[0973] Natural language processing libraries: NLTK, spaCy

[0974] Speech Recognition Library: Google Speech-to-Text API

[0975] Image processing library: OpenCV

[0976] Facial Expression Recognition API:Facial Expression Recognition API

[0977] Generative AI model: GPT-3

[0978] User Input

[0979] Users can use the chat window on their device to type text messages, or upload audio messages and video files. For example, a user might type "I don't know how to use this product" and attach a video of themselves looking confused.

[0980] Prompt Sentence Examples

[0981] I don't know how to use the new product. How do I set it up?

[0982] Video file example

[0983] Approximately 30 seconds of video file (mp4 format) containing a confused expression

[0984] Data collection and transmission

[0985] The device temporarily stores the text, audio, image, and video data entered by the user and transmits it to the server in real time. Specifically, the device transmits this data to the server using an HTTP POST request.

[0986] Data Preprocessing

[0987] The server preprocesses the various types of data it receives. For text data, it uses a natural language processing library (e.g., NLTK or spaCy) to perform morphological analysis and normalization. For audio data, it adjusts the sampling rate and removes noise. For image and video data, it converts them into an appropriate format using OpenCV and performs preprocessing for face detection.

[0988] Emotion and Needs Analysis

[0989] The server analyzes the preprocessed data. For text data, it performs sentiment analysis using models such as the BERT model. For audio data, it converts the audio to text using the Google Speech-to-Text API and analyzes its tone. For image and video data, it uses the Facial Expression Recognition API for facial recognition and expression analysis.

[0990] Customized content generation

[0991] Based on the analysis results, the server generates customized content that addresses the user's needs and emotions. Using a generative AI model (e.g., GPT-3), it creates text messages detailing how to use the product and reassuring video guidance.

[0992] Providing content and gathering feedback

[0993] The server sends the generated content to the device, which displays it to the user. The user then takes action to solve the problem based on the information provided. The server also collects feedback data from users to continuously improve the system's performance.

[0994] In this way, this system can quickly provide highly customized content based on the user's complex input data, resulting in more efficient customer service and improved customer satisfaction, thereby enhancing a company's competitiveness.

[0995] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0996] Step 1:

[0997] The user provides multimodal input data such as text, voice, images, and video to the device. The input data includes a prompt such as "I don't know how to use this new product. How do I set it up?" and a video file (mp4 format) containing a confused facial expression. The device receives this data based on the user's input.

[0998] Step 2:

[0999] The device temporarily stores the received text, audio, image, and video data and sends it to the server in real time using an HTTP POST request. Input data is stored and sent as is.

[1000] Step 3:

[1001] The server preprocesses the received multimodal input data, specifically by:

[1002] Text data: Perform morphological analysis and normalization using a natural language processing library (e.g., NLTK, spaCy).

[1003] Audio data: Adjust the sampling rate and remove noise.

[1004] Image and video data: Use OpenCV to convert them into the appropriate format and pre-process them for face detection.

[1005] The output is preprocessed text, clear audio data, and format-converted images and videos.

[1006] Step 4:

[1007] The server analyzes the preprocessed data, specifically:

[1008] Text data: Sentiment analysis is performed using models such as BERT.

[1009] Audio data: Converted to text using the Google Speech-to-Text API, and tone analysis is performed based on that text.

[1010] Image and video data: Use the Facial Expression Recognition API for facial recognition and facial expression analysis.

[1011] The output is an analysis that indicates the user's emotional state and needs.

[1012] Step 5:

[1013] The server generates customized content based on the analysis results. Using a generative AI model (e.g., GPT-3), it creates text messages with specific actions to take based on the user's needs and reassuring video guidance. The output is a customized text message and video guidance.

[1014] Step 6:

[1015] The server then sends the generated customized content to the device, which then receives it and displays it to the user. Specifically, the device displays text messages in a chat window and plays video guidance. The user can then take action to resolve the problem based on this content.

[1016] Step 7:

[1017] The server collects feedback data from users. Specifically, it collects the ratings and opinions users have given about the content provided and sends that data to the server. The server then retrains the generative AI model based on the collected feedback data, continuously improving the system's performance. This improves the accuracy of user sentiment analysis and need prediction, providing a better user experience.

[1018] (Application example 1)

[1019] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1020] Conventional customer support systems are unable to fully utilize the complex input data (text, voice, images, video, etc.) from users, making it difficult to accurately grasp users' emotions and specific needs. Food delivery services, in particular, are required to provide quick and accurate solutions to issues such as cold food, but current systems have limitations. Therefore, there is an urgent need to develop a system that can analyze user input data at multiple stages and provide appropriate solutions.

[1021] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[1022] In this invention, the server includes means for collecting multimodal input data, means for preprocessing the collected multimodal input data, means for analyzing emotions and needs using the preprocessed data, means for generating customized content based on the analysis results, means for providing the generated content to users, means for collecting feedback from users and improving system performance, and means for analyzing text and video data entered by users and providing corresponding solutions, including generating guide videos including instructions on how to reheat food when it has cooled and an apology message. This makes it possible to quickly and accurately solve specific user problems and improve customer satisfaction.

[1023] "Multimodal input data" refers to data in multiple different formats, such as text data, audio data, image data, and video data.

[1024] "Means of collection" refers to the mechanisms and technologies for acquiring multimodal input data provided by users.

[1025] "Preprocessing means" refers to the techniques and algorithms used to prepare collected multimodal input data in an analyzable form.

[1026] "Means for emotion and needs analysis" refers to analytical tools and techniques for discovering user emotions and specific needs based on processed data.

[1027] "Means for generating customized content" refers to technology that creates specific information or guides tailored to the user's needs based on the analysis results.

[1028] "Means for providing" refers to the mechanisms and technologies for displaying or notifying the user of the generated customized content.

[1029] "Means of collecting feedback" refers to methods and techniques for obtaining responses and reactions from users and using them to improve the system.

[1030] "Means of providing solutions" refers to mechanisms and technologies for presenting users with specific solutions and proposals for solving problems or issues.

[1031] "Means for generating guide videos" refers to technology for creating video content that includes guidance and instructions that meet the needs of users.

[1032] System configuration

[1033] This system is composed of three elements: a server, a terminal, and a user. Each element is explained in detail below.

[1034] server

[1035] The server is responsible for data collection, preprocessing, analysis, content generation, and feedback collection and performance improvement. Specific software used includes TensorFlow, PyTorch, OpenCV, NLTK, and Transformer-based NLP models (e.g., BERT).

[1036] Terminal

[1037] The terminal is a device such as a smartphone that functions as an interface with the user. Specifically, it is built using React Native and plays a role in sending data entered by the user to the server and displaying the content received from the server.

[1038] User

[1039] Users are the users of the system and provide multimodal input data (text, audio, images, video), which is analyzed and customized content is returned to the user.

[1040] Processing flow

[1041] 1. User input:

[1042] The user inputs a text message and uploads video data via the device. For example, the user can provide a message such as "My food has arrived but it's cold and I'm worried" along with a video of the food being cold.

[1043] 2. Data Collection:

[1044] The device temporarily stores the text and video data sent by the user and then transmits it to the server. Data transmission is performed using real-time communication.

[1045] 3. Data preprocessing:

[1046] The server preprocesses the received data: text data is segmented using NLTK, and video data is converted to the appropriate format using OpenCV.

[1047] 4. Emotion and Needs Analysis:

[1048] The server analyzes the preprocessed data: text data is subjected to sentiment analysis using natural language processing techniques such as the BERT model, and video data is subjected to facial recognition and facial expression analysis using OpenCV.

[1049] 5. Customized Content Generation:

[1050] Based on the analysis results, it generates customized content that addresses the user's needs and emotions. In this example, it generates a guide video explaining how to reheat cold food and an apology message.

[1051] 6. Providing Feedback:

[1052] The generated content is sent to the device, which displays it to the user, who then takes action to solve the problem based on the information provided.

[1053] 7. Continuous learning and improvement:

[1054] The server collects feedback data from users and retrains the AI ​​model, improving the accuracy of sentiment analysis and needs prediction.

[1055] Examples and prompts

[1056] Specific examples

[1057] If a user texts, "My food arrived but it's cold and I'm worried," and uploads a video showing the food cold, the system will provide a video guide on how to reheat it and an apology message.

[1058] Prompt Sentence Examples

[1059] User input: The text "My food arrived but it's cold and I'm worried" and a video of the food being cold.

[1060] Desired output: An apology message and a guided video with specific steps for reheating the food.

[1061] The system of the present invention makes it possible to quickly and accurately solve specific user problems and improve customer satisfaction.

[1062] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[1063] Step 1:

[1064] User input:

[1065] The user inputs a text message and uploads video data through the device. For example, the user can provide a text message saying, "My food arrived, but it's cold and I'm worried," along with a video of the cold food. The device then receives this input and prepares it for the next step.

[1066] input:

[1067] Text messages, video data

[1068] output:

[1069] Temporarily saved text messages and video data

[1070] Step 2:

[1071] Data collection:

[1072] The device temporarily stores the text and video data sent by the user and transmits it to the server in real time. This data transmission requires a stable internet connection.

[1073] input:

[1074] Temporarily saved text messages and video data

[1075] output:

[1076] Text messages and video data sent to the server

[1077] Step 3:

[1078] Data preprocessing:

[1079] The server preprocesses the received data: text data is segmented using natural language processing technology (NLTK), and video data is converted into a format that can be analyzed frame by frame using OpenCV.

[1080] input:

[1081] Text messages and video data sent to the server

[1082] output:

[1083] Segmented text data, converted video data

[1084] Step 4:

[1085] Emotion and Needs Analysis:

[1086] The server analyzes emotions and needs based on the preprocessed data. It uses the BERT model for sentiment analysis on text data, and OpenCV for facial recognition and facial expression analysis on video data to identify the user's emotional state.

[1087] input:

[1088] Segmented text data, converted video data

[1089] output:

[1090] Analyzed emotion and needs data

[1091] Step 5:

[1092] Customized Content Generation:

[1093] Based on the analysis results, the system generates customized content that responds to the user's needs and emotions, such as a guide video explaining how to reheat cold food and an apology message.

[1094] input:

[1095] Analyzed emotion and needs data

[1096] output:

[1097] Guide video, apology message

[1098] Step 6:

[1099] Providing feedback:

[1100] The server sends the generated content to the device, which displays it to the user, who then acts to solve the problem based on the information provided.

[1101] input:

[1102] Guide video, apology message

[1103] output:

[1104] Guide video and apology message displayed to users

[1105] Step 7:

[1106] Continuous learning and improvement:

[1107] The server collects feedback data from users and uses it to retrain the AI ​​model, improving the accuracy of sentiment analysis and needs prediction.

[1108] input:

[1109] User feedback data

[1110] output:

[1111] Improved AI models

[1112] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[1113] This invention relates to a system that provides customized content by utilizing multimodal inputs such as text, voice, images, and video, and by performing advanced analysis of user needs and emotions. In particular, by combining an emotion engine, it is possible to accurately recognize the user's emotional state and generate appropriate responses.

[1114] System Overview

[1115] This system is composed of three entities: a server, a terminal, and a user. The server collects data, preprocesses it, analyzes it, generates customized content, and collects feedback and improves performance, while the terminal functions as an interface with the user. As a user of the system, the user provides multimodal input data and receives customized content based on the analysis results.

[1116] Incorporating an emotion engine

[1117] One of the features of this system is the incorporation of an emotion engine, which integrates natural language processing (NLP), speech recognition, facial recognition, and facial expression analysis technologies to analyze the user's emotions in real time.

[1118] Specific examples

[1119] The application of this system to customer support for small and medium-sized enterprises will be explained.

[1120] 1. User Input

[1121] A user texts the support chat saying, "I don't know how to use the product," and then uploads a video showing a confused expression.

[1122] 2. Data collection

[1123] The device temporarily stores the multimodal data sent by the user, including text, audio, images, and video, and transmits it to the server in real time. The data is temporarily cached to prevent data loss.

[1124] 3. Data Preprocessing

[1125] The server preprocesses the various data it receives, specifically tokenizing and normalizing text data, converting the sampling rate of audio data, and converting the format of image and video data.

[1126] 4. Emotion and Needs Analysis

[1127] The emotion engine uses the pre-processed data to analyze the user's emotional state and needs, using natural language processing technology to perform sentiment analysis of the text, speech recognition technology to convert the voice data into text and analyze its tone, and computer vision technology to perform facial recognition and facial expression analysis from images and videos.

[1128] 5. Customized Content Generation

[1129] Based on the analysis, the server generates customized content that responds to the user's emotional state and needs, such as a text message detailing how to use the product and a reassuring video guide.

[1130] 6. Providing Feedback

[1131] The server transmits the generated customized content to the terminal, compressing the data as necessary to transmit the data efficiently.

[1132] The device displays the received content to the user, displaying text messages on the screen and playing video guidance.

[1133] 7. Continuous learning and improvement

[1134] The server collects feedback from the user, which may include a text or audio rating.

[1135] The server analyzes the collected feedback and retrains the emotion engine algorithm, thereby improving the system's accuracy and response quality.

[1136] This system can accurately grasp the user's emotional state and needs in real time based on a variety of input data, and quickly provide highly customized content. This can improve the efficiency of customer service and customer satisfaction, strengthening the competitiveness of small and medium-sized enterprises. Similar benefits can be expected in the fields of education and research and development.

[1137] The processing flow will be explained below.

[1138] Step 1:

[1139] A user inputs multimodal data such as text, audio, images, and video. For example, a user texts a support chat message saying, "I don't know how to use the product," and uploads a video showing a confused expression.

[1140] Step 2:

[1141] The device temporarily stores the text, voice, image, and video data entered by the user and transmits it to the server in real time, using a cache process to prevent data loss.

[1142] Step 3:

[1143] The server preprocesses the multimodal data it receives, specifically tokenizing and normalizing text data, converting the sampling rate of audio data, and converting the format of image and video data.

[1144] Step 4:

[1145] The emotion engine analyzes the pre-processed data, using natural language processing technology to perform sentiment analysis on the text data, speech recognition technology to convert the voice data into text data and analyze its tone, and computer vision technology to perform facial recognition and facial expression analysis on image and video data to identify the user's emotional state.

[1146] Step 5:

[1147] The server identifies the user's needs and emotional state based on the analysis results from the emotion engine. For example, if the text data is judged as "confused," a confused expression is also detected in the facial expression analysis.

[1148] Step 6:

[1149] The server generates customized content based on the identified needs and emotional state, for example, creating a text message detailing how to use a product or generating a reassuring video guide.

[1150] Step 7:

[1151] The server transmits the generated customized content to the terminal, compressing the data as necessary to improve the efficiency of data transmission.

[1152] Step 8:

[1153] The device displays the received customized content to the user, displaying text messages on the screen and playing video guidance.

[1154] Step 9:

[1155] The user reviews the customized content provided and takes action to resolve the issue.

[1156] Step 10:

[1157] The server collects feedback from the user, which may be received in text or audio format.

[1158] Step 11:

[1159] The server analyzes the feedback data and retrains it to improve the performance of the emotion engine and the overall system, thereby increasing the accuracy of future analysis and content generation.

[1160] In this way, incorporating an emotion engine makes it possible to perform advanced analysis of various user data and provide customized content tailored to needs in real time. This will significantly improve the quality and efficiency of customer service, helping to strengthen the competitiveness of small and medium-sized enterprises.

[1161] Example 2

[1162] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1163] Conventional customer support systems have struggled to efficiently collect and analyze diverse user input data and provide customized content tailored to the user's emotional state and needs. While systems that utilize multimodal input data to accurately grasp the user's emotions and generate and deliver appropriate responses are particularly needed, few such systems exist. This has led to lower user satisfaction and negatively impacted the efficiency of companies' customer support efforts.

[1164] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[1165] In this invention, the server includes: means for a user to provide multimodal input data including text, voice, images, and videos; means for temporarily storing the received multimodal input data and transmitting it to the server in real time; data preprocessing means for tokenizing, normalizing, formatting, and sampling rate converting the data received by the server; means for an emotion engine to analyze the user's emotions and needs using natural language processing, speech recognition, face recognition, and facial expression analysis techniques using the preprocessed data; means for generating customized content based on the analysis results; means for transmitting the generated content to a terminal and providing it to the user; and means for collecting user feedback and retraining the system's emotion engine to improve performance. This makes it possible to efficiently collect and analyze a variety of user input data and provide appropriate customized content in real time, thereby improving customer service efficiency and user satisfaction.

[1166] "Multimodal input data" refers to data in multiple formats, such as text, audio, images, and video.

[1167] "Preprocessing" refers to processing such as tokenization, normalization, format conversion, and sampling rate conversion to make data easier to analyze.

[1168] An "emotion engine" refers to a system that analyzes a user's emotional state by integrating natural language processing technology, voice recognition technology, facial recognition technology, and facial expression analysis technology.

[1169] "Natural language processing technology" refers to technology that enables computers to understand, analyze, and generate human language.

[1170] "Speech recognition technology" refers to the technology that converts speech into text data and analyzes its content.

[1171] "Facial recognition technology" refers to technology that detects and identifies human faces from images and videos.

[1172] "Facial expression analysis technology" refers to technology that analyzes facial expressions and identifies their emotional state.

[1173] "Customized content" refers to individually optimized information and responses generated based on a user's emotional state and needs.

[1174] "Feedback" refers to subsequent data such as ratings, opinions, and impressions collected from users.

[1175] "Retraining" refers to the process of retraining an existing learning model with new data to improve its performance.

[1176] "Server" refers to the central management system that collects, pre-processes, analyzes, generates content, gathers feedback, and improves performance of data.

[1177] "Terminal" refers to a device that acts as an interface with a user and is responsible for inputting and outputting data.

[1178] This invention relates to a system that uses multimodal input data such as text, voice, images, and video to provide customized content through advanced analysis of user needs and emotions. In particular, by combining an emotion engine, it becomes possible to accurately recognize the user's emotional state and generate appropriate responses. This system is composed of three entities: a "server," a "terminal," and a "user."

[1179] The server collects, preprocesses, and analyzes data, generates customized content, and collects feedback to improve performance. The terminal functions as an interface with the user, who provides multimodal input data and receives customized content based on the analysis results.

[1180] Hardware and software used

[1181] The servers are data centers or cloud computing services with high-performance computing capabilities, and the software uses the BERT model for natural language processing (NLP), the Google Speech-to-Text API for speech recognition, and OpenCV and Dlib libraries for facial recognition and facial expression analysis.

[1182] Specific examples

[1183] We will explain how this system can be applied to customer support for small and medium-sized enterprises.

[1184] User Input

[1185] A user can text the support chat saying, "I don't know how to use the product," and then upload a video showing a confused expression, which inputs the user's needs and emotions into the system.

[1186] Data collection

[1187] The device temporarily stores the multimodal data received from the user, including text, audio, images, and video, and transmits it to the server in real time. The data is temporarily cached to prevent data loss.

[1188] Data Preprocessing

[1189] The server performs preprocessing on the received data, including tokenization, normalization, format conversion, and sampling rate conversion. Specifically, text data is processed using the BERT model, and the sampling rate of audio data is converted to 16 kHz to optimize it for analysis. For image and video data, the resolution is changed to 1280x720 pixels using the OpenCV library and the data is converted to the required format.

[1190] Emotion and Needs Analysis

[1191] The emotion engine uses the pre-processed data to analyze the user's emotional state and needs. It uses natural language processing technology to perform sentiment analysis of the text, speech recognition technology to convert the voice data into text and analyze its tone, and computer vision technology to perform facial recognition and facial expression analysis from images and videos to understand the user's emotions.

[1192] Customized Content Generation

[1193] Based on the analysis results, the server generates customized content that corresponds to the user's emotional state and needs. For example, based on the emotion analysis result of "confused," a text message explaining detailed product usage and a reassuring video guide are generated.

[1194] Providing feedback

[1195] The server transmits the generated customized content to the terminal, which compresses the content as necessary for efficient transmission. The terminal then displays the received content to the user, including displaying a text message on the screen and playing a video guide.

[1196] Continuous learning and improvement

[1197] The server collects feedback from users, including text and voice ratings, and analyzes the collected feedback and re-trains the emotion engine using TensorFlow to improve the system's accuracy and response quality.

[1198] Prompt Sentence Examples

[1199] "If a customer is confused about how to use a product, use an emotion engine to analyze their emotions in real time and generate a text message explaining how to use the product and a reassuring video guide."

[1200] This makes it possible to efficiently collect and analyze a wide range of user input data and provide appropriate customized content in real time. This system will improve the efficiency of customer service and user satisfaction, strengthening the competitiveness of small and medium-sized enterprises. Similar benefits can be expected in the fields of education and research and development.

[1201] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1202] Step 1:

[1203] The user types a text message into the support chat and uploads a video showing a confused expression. The input data includes four types: text, audio, images, and video. This allows the user's needs and emotions to be input into the system. For example, the user can send a specific text message such as, "I don't know how to use the product."

[1204] Input: Text message, video (confused expression)

[1205] Output: Collected multimodal data

[1206] Step 2:

[1207] The device temporarily stores multimodal data received from the user, including text, audio, images, and video. The stored data is temporarily cached and then sent to the server in real time. This prevents data loss and enables analysis on the server side.

[1208] Input: Collected multimodal data

[1209] Output: Multimodal data sent to the server

[1210] Step 3:

[1211] The server preprocesses the received data. For text data, tokenization and normalization are performed using the BERT model. For audio data, the sampling rate is converted to 16 kHz to optimize the data for analysis. For image and video data, the resolution is changed to 1280x720 pixels using the OpenCV library and the data is converted to the required format.

[1212] Input: Multimodal data sent to the server

[1213] Output: Preprocessed data (tokenized, normalized, formatted, and sample rate converted)

[1214] Step 4:

[1215] The emotion engine uses the pre-processed data to analyze the user's emotional state and needs. It uses natural language processing technology to perform sentiment analysis of the text, speech recognition technology to convert the voice data into text and analyze its tone, and computer vision technology to perform facial recognition and facial expression analysis from images and videos to understand the user's emotions. The results of this analysis determine the user's specific emotions and needs.

[1216] Input: Preprocessed data

[1217] Output: User's emotional state and needs analysis results

[1218] Step 5:

[1219] Based on the analysis results, the server generates customized content that corresponds to the user's emotional state and needs. For example, based on the emotion analysis result of "confused," a text message explaining detailed product usage and a reassuring video guide are generated.

[1220] Input: User's emotional state and needs analysis results

[1221] Output: Customized content (text messages, video guidance)

[1222] Step 6:

[1223] The server transmits the generated customized content to the terminal, which compresses the content as necessary for efficient transmission. The terminal then displays the received content to the user, including displaying a text message on the screen and playing a video guide.

[1224] Input:Customized Content

[1225] Output: Customized content served to the user

[1226] Step 7:

[1227] The server collects feedback from users, including text and voice ratings, and analyzes the collected feedback and re-trains the emotion engine using TensorFlow to improve the system's accuracy and response quality.

[1228] Input: User feedback

[1229] Output: Performance improvement by retraining the emotion engine

[1230] The above are the specific processing steps of the program in this system.

[1231] (Application example 2)

[1232] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1233] Current content distribution services struggle to analyze a user's emotional state and needs in real time and provide appropriate entertainment content. Conventional systems often fail to provide personalized recommendations based on the user's emotions and state, resulting in reduced user satisfaction. The present invention aims to address these challenges and provide optimal content tailored to the user's emotional state and needs.

[1234] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[1235] In this invention, the server includes means for collecting multimodal input data, means for preprocessing the collected multimodal input data, means for analyzing emotions and needs using the preprocessed data, means for generating customized content based on the analysis results, means for providing the generated content to the user, means for collecting feedback from the user and improving the performance of the system, and means for recommending entertainment content based on the emotional state and needs of the user, thereby making it possible to provide personalized entertainment content according to the emotional state and needs of the user in real time.

[1236] "Multimodal input data" refers to data in multiple formats, such as text data, audio data, image data, and video data.

[1237] "Means for collection" refers to the devices and software that obtain and store multimodal input data from users.

[1238] "Preprocessing means" refers to devices or software that perform processes to convert collected multimodal input data into a format that is easy to analyze.

[1239] "Emotion and needs analysis means" refers to algorithms or software that identify a user's emotional state and needs based on pre-processed data.

[1240] "Means for generating customized content" refers to systems or software for creating optimal content for users based on the analysis results.

[1241] The "means for providing to the user" refers to a device or interface for showing the generated customized content to the user.

[1242] "Means for collecting user feedback" refers to devices or software that collect and record user ratings and opinions.

[1243] "Means for improving system performance" refers to methods for improving the behavior of algorithms or software based on collected feedback.

[1244] "Emotional state" refers to data or states that represent a user's mental and emotional state.

[1245] "Needs" refer to the requests and desires for information, functions, and services that users desire.

[1246] "Entertainment content" refers to various media content for entertainment purposes, such as movies, dramas, music, and games.

[1247] The present invention relates to a system for providing customized entertainment content based on a user's emotional state and needs, which is implemented in the following steps:

[1248] System configuration

[1249] This system consists of three components: a server, a terminal, and a user. The server collects data, preprocesses it, analyzes it, generates content, and collects feedback to improve performance. The terminal functions as an interface with the user. As a user of the system, the user provides multimodal input data and receives customized content based on the analysis results.

[1250] Hardware and software used

[1251] This system uses the following hardware and software:

[1252] Hardware:

[1253] Device: Smartphone or head-mounted display equipped with a camera and microphone

[1254] software:

[1255] Cloud services: Cloud services for data processing (e.g., AWS, Google Cloud)

[1256] Natural language processing engine: Google NLP

[1257] Speech recognition engine: Amazon Transcribe

[1258] Facial expression and face recognition engine: Microsoft Azure Face API

[1259] Data processing and calculation

[1260] The device captures multimodal input data (text, voice, image, and video) from the user. This data is temporarily cached and sent to the server. The server preprocesses the received data, tokenizing and normalizing the text data. The voice data undergoes sampling rate conversion and text conversion, and the image and video data undergo format conversion and facial expression analysis.

[1261] The following techniques are used for the analysis:

[1262] Sentiment analysis of text using natural language processing (NLP) technology

[1263] Voice recognition technology converts voice data into text and analyzes the tone

[1264] Uses computer vision technology to perform facial recognition and facial expression analysis

[1265] Content generation and delivery

[1266] Based on the analysis results, the server generates a recommendation list of entertainment content (movies, dramas, music) according to the user's emotional state and needs. This recommendation list is customized and accessible to the user via smartphone or head-mounted display. The user can view the recommended content and provide feedback.

[1267] Specific examples

[1268] For example, if a user uses a smartphone to type the text "I'm feeling a bit down today," input a low-pitched voice into the microphone, and capture an image of a gloomy facial expression with the camera, all of this data is immediately sent to the server and analyzed.

[1269] The server analyzes the user's emotional state (feeling depressed) and needs (content that will lift their spirits) and recommends the most suitable movies and music for the user.

[1270] Prompt Sentence Examples

[1271] If you send "I'm feeling a bit down today," enter the following prompt into the model:

[1272] text

[1273] User sent text: "I'm feeling a bit down today."

[1274] Voice data: User's voice is low and restless

[1275] Facial expression data: The user has a dull expression

[1276] Use this data to generate content recommendations for your users.

[1277] This makes it possible to provide entertainment content in real time that matches the user's emotional state and needs, thereby increasing user satisfaction and improving the value of content distribution services.

[1278] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1279] Step 1:

[1280] The device collects the user's input data. Specifically, the device's camera and microphone are used to capture the user's facial expressions and voice in real time, and the user inputs text. The input data at this time includes text data, voice data, and image data. This data is temporarily stored on the device and then sent to the server.

[1281] input:

[1282] Text data: Character information entered by the user

[1283] Voice data: recordings of your voice

[1284] Image data: photos or videos of the user's facial expressions

[1285] output:

[1286] Capture and temporary storage of multimodal input data (text, audio, images)

[1287] Step 2:

[1288] The server preprocesses the received multimodal input data, specifically tokenizing and normalizing text data, converting audio data to text through sampling rate conversion, and formatting and analyzing facial expressions on image and video data.

[1289] input:

[1290] Multimodal input data (text, audio, images)

[1291] output:

[1292] Preprocessed data (tokenized text, transcribed audio, analyzed images)

[1293] Step 3:

[1294] The server's emotion engine uses the pre-processed data to analyze emotions and needs, using natural language processing (NLP) technology to perform sentiment analysis of text, speech recognition technology to analyze the tone of audio data, and computer vision technology to perform facial recognition and facial expression analysis from images and videos.

[1295] input:

[1296] Preprocessed data (tokenized text, transcribed audio, analyzed images)

[1297] output:

[1298] Analysis of user's emotional state and needs

[1299] Step 4:

[1300] The server generates customized content based on the analysis results, specifically a recommended list of entertainment content (movies, dramas, music) that corresponds to the user's emotional state and needs.

[1301] input:

[1302] Analysis of user's emotional state and needs

[1303] output:

[1304] Customized entertainment content recommendations

[1305] Step 5:

[1306] The server sends the generated customized content to the device, which displays it to the user and provides recommended content. The user then views the recommended content and provides feedback on their impressions and ratings.

[1307] input:

[1308] Customized entertainment content recommendations

[1309] output:

[1310] Device screen showing recommended content and feedback data

[1311] Step 6:

[1312] The server collects user feedback and uses it to improve the system performance. The collected feedback data is analyzed and used to retrain the emotion engine algorithm.

[1313] input:

[1314] User feedback data

[1315] output:

[1316] Improved emotion engine algorithm and system performance improvement

[1317] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[1318] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1319] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.

[1320] [Fourth embodiment]

[1321] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.

[1322] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[1323] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[1324] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.

[1325] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[1326] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[1327] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[1328] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.

[1329] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[1330] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[1331] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[1332] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[1333] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1334] The present invention relates to a system that uses multimodal inputs such as text, voice, images, and videos to understand the needs and emotions of users and provide customized content and solutions. The program of this system is explained below in natural language.

[1335] System Overview

[1336] This system consists of three entities: a server, a terminal, and a user. The server is responsible for data collection, preprocessing, analysis, content generation, and feedback collection and performance improvement, while the terminal functions as an interface with the user. As a user of the system, the user provides multimodal input data and receives customized content as a result.

[1337] Specific examples

[1338] For example, a case where this system is applied to customer support for small and medium-sized enterprises will be described.

[1339] 1. User Input

[1340] A user sends a text message through the chat window saying, "I don't know how to use the product," and then uploads a video showing their confused expression.

[1341] 2. Data collection

[1342] The terminal temporarily stores the text, audio, image, and video data sent by the user and transmits it to the server in real time.

[1343] 3. Data Preprocessing

[1344] The server preprocesses the various data it receives, specifically by segmenting text data, normalizing words and phrases, adjusting the sampling rate of audio data, and converting image and video formats.

[1345] 4. Emotion and Needs Analysis

[1346] The server analyzes the pre-processed data, using natural language processing technology to perform sentiment analysis of the text, speech recognition technology to convert the voice data into text and analyze its tone, and computer vision technology to perform facial recognition and facial expression analysis from image and video data to identify the user's emotional state.

[1347] 5. Customized Content Generation

[1348] Based on the analysis, the server generates customized content that addresses the user's needs and emotions. In this example, it creates a text message detailing how to use the product and a reassuring video guide.

[1349] 6. Providing Feedback

[1350] The server sends the generated content to the terminal,

[1351] The device displays this information to the user, who then takes action to resolve the problem based on the information provided.

[1352] 7. Continuous learning and improvement

[1353] The server collects feedback data from users and retrains the AI ​​model, improving the accuracy of sentiment analysis and needs prediction.

[1354] In this way, this system can quickly provide highly customized content based on the user's complex input data. This will improve the efficiency of customer service, increase customer satisfaction, and strengthen the competitiveness of small and medium-sized enterprises. Similar benefits can also be expected in the fields of education and research and development.

[1355] The processing flow will be explained below.

[1356] Step 1:

[1357] A user inputs multimodal data such as text, voice, images, and video. For example, a user may enter text such as "I don't know how to use the product" into a chat box and upload a video of a confused expression.

[1358] Step 2:

[1359] The device temporarily stores the text, voice, image, and video data entered by the user and transmits it to the server in real time. At this time, the data is temporarily cached to prevent data loss.

[1360] Step 3:

[1361] The server performs preprocessing on the received multimodal data, specifically tokenizing and normalizing text data, sample rate conversion for audio data, and format conversion for image and video data.

[1362] Step 4:

[1363] The server prepares the preprocessed data for analysis, for example, using natural language processing (NLP) to perform sentiment analysis on the text, using speech recognition technology to convert audio data into text, and using computer vision technology to perform facial recognition and facial expression analysis on image and video data.

[1364] Step 5:

[1365] The server identifies the user's emotional state and needs based on the analysis results. For example, the result of sentiment analysis may be "confused," and facial expression analysis may also detect a confused expression.

[1366] Step 6:

[1367] The server generates customized content tailored to the user based on the identified emotional state and needs, such as a text message detailing how to use a product or a reassuring video guide.

[1368] Step 7:

[1369] The server transmits the generated customized content to the terminal, compressing the data as necessary during transmission to ensure efficient data transmission.

[1370] Step 8:

[1371] The device displays the received content to the user. For example, a text message is displayed on the chat screen, and a video guidance is played on a player.

[1372] Step 9:

[1373] The user reviews the customized content provided and takes the necessary action to resolve the issue.

[1374] Step 10:

[1375] The server collects feedback data from users, including text and audio ratings.

[1376] Step 11:

[1377] The server analyzes the collected feedback, evaluates the performance of the AI ​​model, and identifies areas for improvement, which are then used for retraining to improve the accuracy of sentiment analysis and needs prediction.

[1378] In this way, through a series of steps, the system can quickly and accurately provide customized content based on the user's multiple input data.

[1379] Example 1

[1380] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1381] Many modern systems rely on single-modal input data from users (text only or voice only), which makes it difficult to fully understand users' emotions and needs, making it difficult to respond appropriately or provide customized content. Furthermore, there is a lack of means to properly collect user feedback and improve system performance, making continuous improvement difficult.

[1382] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[1383] In this invention, the server includes a means for a user to input input data into the terminal, a means for the terminal to temporarily store the input data and send it to the server, and a means for the server to pre-process the collected multimodal input data. This makes it possible to analyze the user's emotions and needs from multiple angles and quickly provide highly customized content. It is also possible to collect feedback from users and continuously improve the system's performance.

[1384] A "user" is a person or organization that utilizes the system to input data and receive services.

[1385] A "terminal" is a device that allows a user to input data and send it to a server. Examples include computers, smartphones, and tablets.

[1386] "Server" means a device or system that receives data sent from a terminal, pre-processes and analyzes it, and generates and provides customized content.

[1387] "Multimodal input data" refers to data in multiple formats, including text data, audio data, image data, and video data.

[1388] "Preprocessing" is the process of preparing received data so that it can be easily analyzed, since it is difficult to handle as it is. This includes morphological analysis and normalization of text, adjusting the sampling rate of audio, and converting the format of images and videos.

[1389] "Emotion and needs analysis" refers to identifying a user's emotional state and the information and solutions they need based on pre-processed data. This analysis uses natural language processing, voice recognition, facial recognition, and facial expression analysis technologies.

[1390] "Customized content" refers to information or guidance generated based on the analysis results to address the user's individual needs and emotions, such as text messages or video guidance.

[1391] "Feedback" refers to the reactions and evaluations given by users to the content provided, and is information that can be used to improve the performance of the system.

[1392] "Natural language processing technology" is a technology that processes and analyzes human language using a computer. It includes morphological analysis and sentiment analysis.

[1393] "Speech recognition technology" is a technology that converts voice data into text data and analyzes the content and tone of the text.

[1394] "Facial recognition technology" is a technology that detects faces from images and videos and analyzes their features.

[1395] "Facial expression analysis technology" is a technology that uses facial recognition technology to identify a user's emotional state from their facial expression.

[1396] A "generative AI model" is an artificial intelligence model that generates new information based on given data. Examples include GPT-3.

[1397] A "prompt sentence" is the text of an instruction or question that a user enters into a system.

[1398] MODE FOR CARRYING OUT THE INVENTION

[1399] The present invention relates to a system that uses multimodal input data such as text, audio, images, and videos to understand the needs and emotions of users and provide customized content. Specific embodiments of this system are described below.

[1400] System Overview

[1401] This system consists of three entities: a server, a terminal, and a user. The server is responsible for data collection, preprocessing, analysis, generation of customized content, feedback collection, and performance improvement. The terminal functions as an interface with the user, who, as a user of the system, provides multimodal input data and receives customized content as a result.

[1402] Hardware and software used

[1403] Hardware: Server machines, user devices (computers, smartphones, tablets, etc.)

[1404] software:

[1405] Natural language processing libraries: NLTK, spaCy

[1406] Speech Recognition Library: Google Speech-to-Text API

[1407] Image processing library: OpenCV

[1408] Facial Expression Recognition API:Facial Expression Recognition API

[1409] Generative AI model: GPT-3

[1410] User Input

[1411] Users can use the chat window on their device to type text messages, or upload audio messages and video files. For example, a user might type "I don't know how to use this product" and attach a video of themselves looking confused.

[1412] Prompt Sentence Examples

[1413] I don't know how to use the new product. How do I set it up?

[1414] Video file example

[1415] Approximately 30 seconds of video file (mp4 format) containing a confused expression

[1416] Data collection and transmission

[1417] The device temporarily stores the text, audio, image, and video data entered by the user and transmits it to the server in real time. Specifically, the device transmits this data to the server using an HTTP POST request.

[1418] Data Preprocessing

[1419] The server preprocesses the various types of data it receives. For text data, it uses a natural language processing library (e.g., NLTK or spaCy) to perform morphological analysis and normalization. For audio data, it adjusts the sampling rate and removes noise. For image and video data, it converts them into an appropriate format using OpenCV and performs preprocessing for face detection.

[1420] Emotion and Needs Analysis

[1421] The server analyzes the preprocessed data. For text data, it performs sentiment analysis using models such as the BERT model. For audio data, it converts the audio to text using the Google Speech-to-Text API and analyzes its tone. For image and video data, it uses the Facial Expression Recognition API for facial recognition and expression analysis.

[1422] Customized content generation

[1423] Based on the analysis results, the server generates customized content that addresses the user's needs and emotions. Using a generative AI model (e.g., GPT-3), it creates text messages detailing how to use the product and reassuring video guidance.

[1424] Providing content and gathering feedback

[1425] The server sends the generated content to the device, which displays it to the user. The user then takes action to solve the problem based on the information provided. The server also collects feedback data from users to continuously improve the system's performance.

[1426] In this way, this system can quickly provide highly customized content based on the user's complex input data, resulting in more efficient customer service and improved customer satisfaction, thereby enhancing a company's competitiveness.

[1427] The flow of the identification process in the first embodiment will be described with reference to FIG.

[1428] Step 1:

[1429] The user provides multimodal input data such as text, voice, images, and video to the device. The input data includes a prompt such as "I don't know how to use this new product. How do I set it up?" and a video file (mp4 format) containing a confused facial expression. The device receives this data based on the user's input.

[1430] Step 2:

[1431] The device temporarily stores the received text, audio, image, and video data and sends it to the server in real time using an HTTP POST request. Input data is stored and sent as is.

[1432] Step 3:

[1433] The server preprocesses the received multimodal input data, specifically by:

[1434] Text data: Perform morphological analysis and normalization using a natural language processing library (e.g., NLTK, spaCy).

[1435] Audio data: Adjust the sampling rate and remove noise.

[1436] Image and video data: Use OpenCV to convert them into the appropriate format and pre-process them for face detection.

[1437] The output is preprocessed text, clear audio data, and format-converted images and videos.

[1438] Step 4:

[1439] The server analyzes the preprocessed data, specifically:

[1440] Text data: Sentiment analysis is performed using models such as BERT.

[1441] Audio data: Converted to text using the Google Speech-to-Text API, and tone analysis is performed based on that text.

[1442] Image and video data: Use the Facial Expression Recognition API for facial recognition and facial expression analysis.

[1443] The output is an analysis that indicates the user's emotional state and needs.

[1444] Step 5:

[1445] The server generates customized content based on the analysis results. Using a generative AI model (e.g., GPT-3), it creates text messages with specific actions to take based on the user's needs and reassuring video guidance. The output is a customized text message and video guidance.

[1446] Step 6:

[1447] The server then sends the generated customized content to the device, which then receives it and displays it to the user. Specifically, the device displays text messages in a chat window and plays video guidance. The user can then take action to resolve the problem based on this content.

[1448] Step 7:

[1449] The server collects feedback data from users. Specifically, it collects the ratings and opinions users have given about the content provided and sends that data to the server. The server then retrains the generative AI model based on the collected feedback data, continuously improving the system's performance. This improves the accuracy of user sentiment analysis and need prediction, providing a better user experience.

[1450] (Application example 1)

[1451] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1452] Conventional customer support systems are unable to fully utilize the complex input data (text, voice, images, video, etc.) from users, making it difficult to accurately grasp users' emotions and specific needs. Food delivery services, in particular, are required to provide quick and accurate solutions to issues such as cold food, but current systems have limitations. Therefore, there is an urgent need to develop a system that can analyze user input data at multiple stages and provide appropriate solutions.

[1453] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[1454] In this invention, the server includes means for collecting multimodal input data, means for preprocessing the collected multimodal input data, means for analyzing emotions and needs using the preprocessed data, means for generating customized content based on the analysis results, means for providing the generated content to users, means for collecting feedback from users and improving system performance, and means for analyzing text and video data entered by users and providing corresponding solutions, including generating guide videos including instructions on how to reheat food when it has cooled and an apology message. This makes it possible to quickly and accurately solve specific user problems and improve customer satisfaction.

[1455] "Multimodal input data" refers to data in multiple different formats, such as text data, audio data, image data, and video data.

[1456] "Means of collection" refers to the mechanisms and technologies for acquiring multimodal input data provided by users.

[1457] "Preprocessing means" refers to the techniques and algorithms used to prepare collected multimodal input data in an analyzable form.

[1458] "Means for emotion and needs analysis" refers to analytical tools and techniques for discovering user emotions and specific needs based on processed data.

[1459] "Means for generating customized content" refers to technology that creates specific information or guides tailored to the user's needs based on the analysis results.

[1460] "Means for providing" refers to the mechanisms and technologies for displaying or notifying the user of the generated customized content.

[1461] "Means of collecting feedback" refers to methods and techniques for obtaining responses and reactions from users and using them to improve the system.

[1462] "Means of providing solutions" refers to mechanisms and technologies for presenting users with specific solutions and proposals for solving problems or issues.

[1463] "Means for generating guide videos" refers to technology for creating video content that includes guidance and instructions that meet the needs of users.

[1464] System configuration

[1465] This system is composed of three elements: a server, a terminal, and a user. Each element is explained in detail below.

[1466] server

[1467] The server is responsible for data collection, preprocessing, analysis, content generation, and feedback collection and performance improvement. Specific software used includes TensorFlow, PyTorch, OpenCV, NLTK, and Transformer-based NLP models (e.g., BERT).

[1468] Terminal

[1469] The terminal is a device such as a smartphone that functions as an interface with the user. Specifically, it is built using React Native and plays a role in sending data entered by the user to the server and displaying the content received from the server.

[1470] User

[1471] Users are the users of the system and provide multimodal input data (text, audio, images, video), which is analyzed and customized content is returned to the user.

[1472] Processing flow

[1473] 1. User input:

[1474] The user inputs a text message and uploads video data via the device. For example, the user can provide a message such as "My food has arrived but it's cold and I'm worried" along with a video of the food being cold.

[1475] 2. Data Collection:

[1476] The device temporarily stores the text and video data sent by the user and then transmits it to the server. Data transmission is performed using real-time communication.

[1477] 3. Data preprocessing:

[1478] The server preprocesses the received data: text data is segmented using NLTK, and video data is converted to the appropriate format using OpenCV.

[1479] 4. Emotion and Needs Analysis:

[1480] The server analyzes the preprocessed data: text data is subjected to sentiment analysis using natural language processing techniques such as the BERT model, and video data is subjected to facial recognition and facial expression analysis using OpenCV.

[1481] 5. Customized Content Generation:

[1482] Based on the analysis results, it generates customized content that addresses the user's needs and emotions. In this example, it generates a guide video explaining how to reheat cold food and an apology message.

[1483] 6. Providing Feedback:

[1484] The generated content is sent to the device, which displays it to the user, who then takes action to solve the problem based on the information provided.

[1485] 7. Continuous learning and improvement:

[1486] The server collects feedback data from users and retrains the AI ​​model, improving the accuracy of sentiment analysis and needs prediction.

[1487] Examples and prompts

[1488] Specific examples

[1489] If a user texts, "My food arrived but it's cold and I'm worried," and uploads a video showing the food cold, the system will provide a video guide on how to reheat it and an apology message.

[1490] Prompt Sentence Examples

[1491] User input: The text "My food arrived but it's cold and I'm worried" and a video of the food being cold.

[1492] Desired output: An apology message and a guided video with specific steps for reheating the food.

[1493] The system of the present invention makes it possible to quickly and accurately solve specific user problems and improve customer satisfaction.

[1494] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[1495] Step 1:

[1496] User input:

[1497] The user inputs a text message and uploads video data through the device. For example, the user can provide a text message saying, "My food arrived, but it's cold and I'm worried," along with a video of the cold food. The device then receives this input and prepares it for the next step.

[1498] input:

[1499] Text messages, video data

[1500] output:

[1501] Temporarily saved text messages and video data

[1502] Step 2:

[1503] Data collection:

[1504] The device temporarily stores the text and video data sent by the user and transmits it to the server in real time. This data transmission requires a stable internet connection.

[1505] input:

[1506] Temporarily saved text messages and video data

[1507] output:

[1508] Text messages and video data sent to the server

[1509] Step 3:

[1510] Data preprocessing:

[1511] The server preprocesses the received data: text data is segmented using natural language processing technology (NLTK), and video data is converted into a format that can be analyzed frame by frame using OpenCV.

[1512] input:

[1513] Text messages and video data sent to the server

[1514] output:

[1515] Segmented text data, converted video data

[1516] Step 4:

[1517] Emotion and Needs Analysis:

[1518] The server analyzes emotions and needs based on the preprocessed data. It uses the BERT model for sentiment analysis on text data, and OpenCV for facial recognition and facial expression analysis on video data to identify the user's emotional state.

[1519] input:

[1520] Segmented text data, converted video data

[1521] output:

[1522] Analyzed emotion and needs data

[1523] Step 5:

[1524] Customized Content Generation:

[1525] Based on the analysis results, the system generates customized content that responds to the user's needs and emotions, such as a guide video explaining how to reheat cold food and an apology message.

[1526] input:

[1527] Analyzed emotion and needs data

[1528] output:

[1529] Guide video, apology message

[1530] Step 6:

[1531] Providing feedback:

[1532] The server sends the generated content to the device, which displays it to the user, who then acts to solve the problem based on the information provided.

[1533] input:

[1534] Guide video, apology message

[1535] output:

[1536] Guide video and apology message displayed to users

[1537] Step 7:

[1538] Continuous learning and improvement:

[1539] The server collects feedback data from users and uses it to retrain the AI ​​model, improving the accuracy of sentiment analysis and needs prediction.

[1540] input:

[1541] User feedback data

[1542] output:

[1543] Improved AI models

[1544] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[1545] This invention relates to a system that provides customized content by utilizing multimodal inputs such as text, voice, images, and video, and by performing advanced analysis of user needs and emotions. In particular, by combining an emotion engine, it is possible to accurately recognize the user's emotional state and generate appropriate responses.

[1546] System Overview

[1547] This system is composed of three entities: a server, a terminal, and a user. The server collects data, preprocesses it, analyzes it, generates customized content, and collects feedback and improves performance, while the terminal functions as an interface with the user. As a user of the system, the user provides multimodal input data and receives customized content based on the analysis results.

[1548] Incorporating an emotion engine

[1549] One of the features of this system is the incorporation of an emotion engine, which integrates natural language processing (NLP), speech recognition, facial recognition, and facial expression analysis technologies to analyze the user's emotions in real time.

[1550] Specific examples

[1551] The application of this system to customer support for small and medium-sized enterprises will be explained.

[1552] 1. User Input

[1553] A user texts the support chat saying, "I don't know how to use the product," and then uploads a video showing a confused expression.

[1554] 2. Data collection

[1555] The device temporarily stores the multimodal data sent by the user, including text, audio, images, and video, and transmits it to the server in real time. The data is temporarily cached to prevent data loss.

[1556] 3. Data Preprocessing

[1557] The server preprocesses the various data it receives, specifically tokenizing and normalizing text data, converting the sampling rate of audio data, and converting the format of image and video data.

[1558] 4. Emotion and Needs Analysis

[1559] The emotion engine uses the pre-processed data to analyze the user's emotional state and needs, using natural language processing technology to perform sentiment analysis of the text, speech recognition technology to convert the voice data into text and analyze its tone, and computer vision technology to perform facial recognition and facial expression analysis from images and videos.

[1560] 5. Customized Content Generation

[1561] Based on the analysis, the server generates customized content that responds to the user's emotional state and needs, such as a text message detailing how to use the product and a reassuring video guide.

[1562] 6. Providing Feedback

[1563] The server transmits the generated customized content to the terminal, compressing the data as necessary to transmit the data efficiently.

[1564] The device displays the received content to the user, displaying text messages on the screen and playing video guidance.

[1565] 7. Continuous learning and improvement

[1566] The server collects feedback from the user, which may include a text or audio rating.

[1567] The server analyzes the collected feedback and retrains the emotion engine algorithm, thereby improving the system's accuracy and response quality.

[1568] This system can accurately grasp the user's emotional state and needs in real time based on a variety of input data, and quickly provide highly customized content. This can improve the efficiency of customer service and customer satisfaction, strengthening the competitiveness of small and medium-sized enterprises. Similar benefits can be expected in the fields of education and research and development.

[1569] The processing flow will be explained below.

[1570] Step 1:

[1571] A user inputs multimodal data such as text, audio, images, and video. For example, a user texts a support chat message saying, "I don't know how to use the product," and uploads a video showing a confused expression.

[1572] Step 2:

[1573] The device temporarily stores the text, voice, image, and video data entered by the user and transmits it to the server in real time, using a cache process to prevent data loss.

[1574] Step 3:

[1575] The server preprocesses the multimodal data it receives, specifically tokenizing and normalizing text data, converting the sampling rate of audio data, and converting the format of image and video data.

[1576] Step 4:

[1577] The emotion engine analyzes the pre-processed data, using natural language processing technology to perform sentiment analysis on the text data, speech recognition technology to convert the voice data into text data and analyze its tone, and computer vision technology to perform facial recognition and facial expression analysis on image and video data to identify the user's emotional state.

[1578] Step 5:

[1579] The server identifies the user's needs and emotional state based on the analysis results from the emotion engine. For example, if the text data is judged as "confused," a confused expression is also detected in the facial expression analysis.

[1580] Step 6:

[1581] The server generates customized content based on the identified needs and emotional state, for example, creating a text message detailing how to use a product or generating a reassuring video guide.

[1582] Step 7:

[1583] The server transmits the generated customized content to the terminal, compressing the data as necessary to improve the efficiency of data transmission.

[1584] Step 8:

[1585] The device displays the received customized content to the user, displaying text messages on the screen and playing video guidance.

[1586] Step 9:

[1587] The user reviews the customized content provided and takes action to resolve the issue.

[1588] Step 10:

[1589] The server collects feedback from the user, which may be received in text or audio format.

[1590] Step 11:

[1591] The server analyzes the feedback data and retrains it to improve the performance of the emotion engine and the overall system, thereby increasing the accuracy of future analysis and content generation.

[1592] In this way, incorporating an emotion engine makes it possible to perform advanced analysis of various user data and provide customized content tailored to needs in real time. This will significantly improve the quality and efficiency of customer service, helping to strengthen the competitiveness of small and medium-sized enterprises.

[1593] Example 2

[1594] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1595] Conventional customer support systems have struggled to efficiently collect and analyze diverse user input data and provide customized content tailored to the user's emotional state and needs. While systems that utilize multimodal input data to accurately grasp the user's emotions and generate and deliver appropriate responses are particularly needed, few such systems exist. This has led to lower user satisfaction and negatively impacted the efficiency of companies' customer support efforts.

[1596] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[1597] In this invention, the server includes: means for a user to provide multimodal input data including text, voice, images, and videos; means for temporarily storing the received multimodal input data and transmitting it to the server in real time; data preprocessing means for tokenizing, normalizing, formatting, and sampling rate converting the data received by the server; means for an emotion engine to analyze the user's emotions and needs using natural language processing, speech recognition, face recognition, and facial expression analysis techniques using the preprocessed data; means for generating customized content based on the analysis results; means for transmitting the generated content to a terminal and providing it to the user; and means for collecting user feedback and retraining the system's emotion engine to improve performance. This makes it possible to efficiently collect and analyze a variety of user input data and provide appropriate customized content in real time, thereby improving customer service efficiency and user satisfaction.

[1598] "Multimodal input data" refers to data in multiple formats, such as text, audio, images, and video.

[1599] "Preprocessing" refers to processing such as tokenization, normalization, format conversion, and sampling rate conversion to make data easier to analyze.

[1600] An "emotion engine" refers to a system that analyzes a user's emotional state by integrating natural language processing technology, voice recognition technology, facial recognition technology, and facial expression analysis technology.

[1601] "Natural language processing technology" refers to technology that enables computers to understand, analyze, and generate human language.

[1602] "Speech recognition technology" refers to the technology that converts speech into text data and analyzes its content.

[1603] "Facial recognition technology" refers to technology that detects and identifies human faces from images and videos.

[1604] "Facial expression analysis technology" refers to technology that analyzes facial expressions and identifies their emotional state.

[1605] "Customized content" refers to individually optimized information and responses generated based on a user's emotional state and needs.

[1606] "Feedback" refers to subsequent data such as ratings, opinions, and impressions collected from users.

[1607] "Retraining" refers to the process of retraining an existing learning model with new data to improve its performance.

[1608] "Server" refers to the central management system that collects, pre-processes, analyzes, generates content, gathers feedback, and improves performance of data.

[1609] "Terminal" refers to a device that acts as an interface with a user and is responsible for inputting and outputting data.

[1610] This invention relates to a system that uses multimodal input data such as text, voice, images, and video to provide customized content through advanced analysis of user needs and emotions. In particular, by combining an emotion engine, it becomes possible to accurately recognize the user's emotional state and generate appropriate responses. This system is composed of three entities: a "server," a "terminal," and a "user."

[1611] The server collects, preprocesses, and analyzes data, generates customized content, and collects feedback to improve performance. The terminal functions as an interface with the user, who provides multimodal input data and receives customized content based on the analysis results.

[1612] Hardware and software used

[1613] The servers are data centers or cloud computing services with high-performance computing capabilities, and the software uses the BERT model for natural language processing (NLP), the Google Speech-to-Text API for speech recognition, and OpenCV and Dlib libraries for facial recognition and facial expression analysis.

[1614] Specific examples

[1615] We will explain how this system can be applied to customer support for small and medium-sized enterprises.

[1616] User Input

[1617] A user can text the support chat saying, "I don't know how to use the product," and then upload a video showing a confused expression, which inputs the user's needs and emotions into the system.

[1618] Data collection

[1619] The device temporarily stores the multimodal data received from the user, including text, audio, images, and video, and transmits it to the server in real time. The data is temporarily cached to prevent data loss.

[1620] Data Preprocessing

[1621] The server performs preprocessing on the received data, including tokenization, normalization, format conversion, and sampling rate conversion. Specifically, text data is processed using the BERT model, and the sampling rate of audio data is converted to 16 kHz to optimize it for analysis. For image and video data, the resolution is changed to 1280x720 pixels using the OpenCV library and the data is converted to the required format.

[1622] Emotion and Needs Analysis

[1623] The emotion engine uses the pre-processed data to analyze the user's emotional state and needs. It uses natural language processing technology to perform sentiment analysis of the text, speech recognition technology to convert the voice data into text and analyze its tone, and computer vision technology to perform facial recognition and facial expression analysis from images and videos to understand the user's emotions.

[1624] Customized Content Generation

[1625] Based on the analysis results, the server generates customized content that corresponds to the user's emotional state and needs. For example, based on the emotion analysis result of "confused," a text message explaining detailed product usage and a reassuring video guide are generated.

[1626] Providing feedback

[1627] The server transmits the generated customized content to the terminal, which compresses the content as necessary for efficient transmission. The terminal then displays the received content to the user, including displaying a text message on the screen and playing a video guide.

[1628] Continuous learning and improvement

[1629] The server collects feedback from users, including text and voice ratings, and analyzes the collected feedback and re-trains the emotion engine using TensorFlow to improve the system's accuracy and response quality.

[1630] Prompt Sentence Examples

[1631] "If a customer is confused about how to use a product, use an emotion engine to analyze their emotions in real time and generate a text message explaining how to use the product and a reassuring video guide."

[1632] This makes it possible to efficiently collect and analyze a wide range of user input data and provide appropriate customized content in real time. This system will improve the efficiency of customer service and user satisfaction, strengthening the competitiveness of small and medium-sized enterprises. Similar benefits can be expected in the fields of education and research and development.

[1633] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1634] Step 1:

[1635] The user types a text message into the support chat and uploads a video showing a confused expression. The input data includes four types: text, audio, images, and video. This allows the user's needs and emotions to be input into the system. For example, the user can send a specific text message such as, "I don't know how to use the product."

[1636] Input: Text message, video (confused expression)

[1637] Output: Collected multimodal data

[1638] Step 2:

[1639] The device temporarily stores multimodal data received from the user, including text, audio, images, and video. The stored data is temporarily cached and then sent to the server in real time. This prevents data loss and enables analysis on the server side.

[1640] Input: Collected multimodal data

[1641] Output: Multimodal data sent to the server

[1642] Step 3:

[1643] The server preprocesses the received data. For text data, tokenization and normalization are performed using the BERT model. For audio data, the sampling rate is converted to 16 kHz to optimize the data for analysis. For image and video data, the resolution is changed to 1280x720 pixels using the OpenCV library and the data is converted to the required format.

[1644] Input: Multimodal data sent to the server

[1645] Output: Preprocessed data (tokenized, normalized, formatted, and sample rate converted)

[1646] Step 4:

[1647] The emotion engine uses the pre-processed data to analyze the user's emotional state and needs. It uses natural language processing technology to perform sentiment analysis of the text, speech recognition technology to convert the voice data into text and analyze its tone, and computer vision technology to perform facial recognition and facial expression analysis from images and videos to understand the user's emotions. The results of this analysis determine the user's specific emotions and needs.

[1648] Input: Preprocessed data

[1649] Output: User's emotional state and needs analysis results

[1650] Step 5:

[1651] Based on the analysis results, the server generates customized content that corresponds to the user's emotional state and needs. For example, based on the emotion analysis result of "confused," a text message explaining detailed product usage and a reassuring video guide are generated.

[1652] Input: User's emotional state and needs analysis results

[1653] Output: Customized content (text messages, video guidance)

[1654] Step 6:

[1655] The server transmits the generated customized content to the terminal, which compresses the content as necessary for efficient transmission. The terminal then displays the received content to the user, including displaying a text message on the screen and playing a video guide.

[1656] Input:Customized Content

[1657] Output: Customized content served to the user

[1658] Step 7:

[1659] The server collects feedback from users, including text and voice ratings, and analyzes the collected feedback and re-trains the emotion engine using TensorFlow to improve the system's accuracy and response quality.

[1660] Input: User feedback

[1661] Output: Performance improvement by retraining the emotion engine

[1662] The above are the specific processing steps of the program in this system.

[1663] (Application example 2)

[1664] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1665] Current content distribution services struggle to analyze a user's emotional state and needs in real time and provide appropriate entertainment content. Conventional systems often fail to provide personalized recommendations based on the user's emotions and state, resulting in reduced user satisfaction. The present invention aims to address these challenges and provide optimal content tailored to the user's emotional state and needs.

[1666] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[1667] In this invention, the server includes means for collecting multimodal input data, means for preprocessing the collected multimodal input data, means for analyzing emotions and needs using the preprocessed data, means for generating customized content based on the analysis results, means for providing the generated content to the user, means for collecting feedback from the user and improving the performance of the system, and means for recommending entertainment content based on the emotional state and needs of the user, thereby making it possible to provide personalized entertainment content according to the emotional state and needs of the user in real time.

[1668] "Multimodal input data" refers to data in multiple formats, such as text data, audio data, image data, and video data.

[1669] "Means for collection" refers to the devices and software that obtain and store multimodal input data from users.

[1670] "Preprocessing means" refers to devices or software that perform processes to convert collected multimodal input data into a format that is easy to analyze.

[1671] "Emotion and needs analysis means" refers to algorithms or software that identify a user's emotional state and needs based on pre-processed data.

[1672] "Means for generating customized content" refers to systems or software for creating optimal content for users based on the analysis results.

[1673] The "means for providing to the user" refers to a device or interface for showing the generated customized content to the user.

[1674] "Means for collecting user feedback" refers to devices or software that collect and record user ratings and opinions.

[1675] "Means for improving system performance" refers to methods for improving the behavior of algorithms or software based on collected feedback.

[1676] "Emotional state" refers to data or states that represent a user's mental and emotional state.

[1677] "Needs" refer to the requests and desires for information, functions, and services that users desire.

[1678] "Entertainment content" refers to various media content for entertainment purposes, such as movies, dramas, music, and games.

[1679] The present invention relates to a system for providing customized entertainment content based on a user's emotional state and needs, which is implemented in the following steps:

[1680] System configuration

[1681] This system consists of three components: a server, a terminal, and a user. The server collects data, preprocesses it, analyzes it, generates content, and collects feedback to improve performance. The terminal functions as an interface with the user. As a user of the system, the user provides multimodal input data and receives customized content based on the analysis results.

[1682] Hardware and software used

[1683] This system uses the following hardware and software:

[1684] Hardware:

[1685] Device: Smartphone or head-mounted display equipped with a camera and microphone

[1686] software:

[1687] Cloud services: Cloud services for data processing (e.g., AWS, Google Cloud)

[1688] Natural language processing engine: Google NLP

[1689] Speech recognition engine: Amazon Transcribe

[1690] Facial expression and face recognition engine: Microsoft Azure Face API

[1691] Data processing and calculation

[1692] The device captures multimodal input data (text, voice, image, and video) from the user. This data is temporarily cached and sent to the server. The server preprocesses the received data, tokenizing and normalizing the text data. The voice data undergoes sampling rate conversion and text conversion, and the image and video data undergo format conversion and facial expression analysis.

[1693] The following techniques are used for the analysis:

[1694] Sentiment analysis of text using natural language processing (NLP) technology

[1695] Voice recognition technology converts voice data into text and analyzes the tone

[1696] Uses computer vision technology to perform facial recognition and facial expression analysis

[1697] Content generation and delivery

[1698] Based on the analysis results, the server generates a recommendation list of entertainment content (movies, dramas, music) according to the user's emotional state and needs. This recommendation list is customized and accessible to the user via smartphone or head-mounted display. The user can view the recommended content and provide feedback.

[1699] Specific examples

[1700] For example, if a user uses a smartphone to type the text "I'm feeling a bit down today," input a low-pitched voice into the microphone, and capture an image of a gloomy facial expression with the camera, all of this data is immediately sent to the server and analyzed.

[1701] The server analyzes the user's emotional state (feeling depressed) and needs (content that will lift their spirits) and recommends the most suitable movies and music for the user.

[1702] Prompt Sentence Examples

[1703] If you send "I'm feeling a bit down today," enter the following prompt into the model:

[1704] text

[1705] User sent text: "I'm feeling a bit down today."

[1706] Voice data: User's voice is low and restless

[1707] Facial expression data: The user has a dull expression

[1708] Use this data to generate content recommendations for your users.

[1709] This makes it possible to provide entertainment content in real time that matches the user's emotional state and needs, thereby increasing user satisfaction and improving the value of content distribution services.

[1710] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1711] Step 1:

[1712] The device collects the user's input data. Specifically, the device's camera and microphone are used to capture the user's facial expressions and voice in real time, and the user inputs text. The input data at this time includes text data, voice data, and image data. This data is temporarily stored on the device and then sent to the server.

[1713] input:

[1714] Text data: Character information entered by the user

[1715] Voice data: recordings of your voice

[1716] Image data: photos or videos of the user's facial expressions

[1717] output:

[1718] Capture and temporary storage of multimodal input data (text, audio, images)

[1719] Step 2:

[1720] The server preprocesses the received multimodal input data, specifically tokenizing and normalizing text data, converting audio data to text through sampling rate conversion, and formatting and analyzing facial expressions on image and video data.

[1721] input:

[1722] Multimodal input data (text, audio, images)

[1723] output:

[1724] Preprocessed data (tokenized text, transcribed audio, analyzed images)

[1725] Step 3:

[1726] The server's emotion engine uses the pre-processed data to analyze emotions and needs, using natural language processing (NLP) technology to perform sentiment analysis of text, speech recognition technology to analyze the tone of audio data, and computer vision technology to perform facial recognition and facial expression analysis from images and videos.

[1727] input:

[1728] Preprocessed data (tokenized text, transcribed audio, analyzed images)

[1729] output:

[1730] Analysis of user's emotional state and needs

[1731] Step 4:

[1732] The server generates customized content based on the analysis results, specifically a recommended list of entertainment content (movies, dramas, music) that corresponds to the user's emotional state and needs.

[1733] input:

[1734] Analysis of user's emotional state and needs

[1735] output:

[1736] Customized entertainment content recommendations

[1737] Step 5:

[1738] The server sends the generated customized content to the device, which displays it to the user and provides recommended content. The user then views the recommended content and provides feedback on their impressions and ratings.

[1739] input:

[1740] Customized entertainment content recommendations

[1741] output:

[1742] Device screen showing recommended content and feedback data

[1743] Step 6:

[1744] The server collects user feedback and uses it to improve the system performance. The collected feedback data is analyzed and used to retrain the emotion engine algorithm.

[1745] input:

[1746] User feedback data

[1747] output:

[1748] Improved emotion engine algorithm and system performance improvement

[1749] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.

[1750] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1751] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.

[1752] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[1753] FIG. 9 is a diagram illustrating an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and actions arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.

[1754] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.

[1755] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).

[1756] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.

[1757] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."

[1758] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values ​​indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.

[1759] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).

[1760] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.

[1761] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.

[1762] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[1763] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.

[1764] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.

[1765] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.

[1766] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.

[1767] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.

[1768] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.

[1769] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.

[1770] The following is further disclosed regarding the above embodiment.

[1771] (Claim 1)

[1772] a means for collecting multimodal input data;

[1773] means for preprocessing the collected multimodal input data;

[1774] means for analyzing emotions and needs using the preprocessed data;

[1775] means for generating customized content based on the analysis results;

[1776] means for providing the generated content to a user;

[1777] a means of collecting user feedback and improving the performance of the system;

[1778] A system including:

[1779] (Claim 2)

[1780] 10. The system of claim 1, wherein the multimodal input data includes text data, audio data, image data, or video data.

[1781] (Claim 3)

[1782] 2. The system according to claim 1, wherein natural language processing technology, voice recognition technology, face recognition technology, and facial expression analysis technology are used to analyze emotions and needs.

[1783] "Example 1"

[1784] (Claim 1)

[1785] means for a user to input input data into a terminal;

[1786] A means for the terminal to temporarily store input data and transmit it to a server;

[1787] means for the server to preprocess the collected multimodal input data;

[1788] means for analyzing emotions and needs using the preprocessed data;

[1789] means for generating customized content based on the analysis results;

[1790] means for providing the generated content to a user;

[1791] a means of collecting user feedback and improving the performance of the system;

[1792] A system including:

[1793] (Claim 2)

[1794] 10. The system of claim 1, wherein the multimodal input data includes text data, audio data, image data, or video data.

[1795] (Claim 3)

[1796] 2. The system according to claim 1, wherein natural language processing technology, voice recognition technology, face recognition technology, and facial expression analysis technology are used to analyze emotions and needs.

[1797] "Application Example 1"

[1798] Yes, here are the new patent claims incorporating the application example:

[1799] (Claim 1)

[1800] a means for collecting multimodal input data;

[1801] means for preprocessing the collected multimodal input data;

[1802] means for analyzing emotions and needs using the preprocessed data;

[1803] means for generating customized content based on the analysis results;

[1804] means for providing the generated content to a user;

[1805] a means of collecting user feedback and improving the performance of the system;

[1806] A means for analyzing the text and video data entered by the user and generating a guide video including a reheating method when the food has cooled down and an apology message as a means for providing a corresponding solution;

[1807] A system including:

[1808] (Claim 2)

[1809] 10. The system of claim 1, wherein the multimodal input data includes text data, audio data, image data, or video data.

[1810] (Claim 3)

[1811] 2. The system according to claim 1, wherein natural language processing technology, voice recognition technology, face recognition technology, and facial expression analysis technology are used to analyze emotions and needs.

[1812] "Example 2: Combining Emotion Engines"

[1813] (Claim 1)

[1814] a means for a user to provide multimodal input data including text, audio, images, and video;

[1815] a means for temporarily storing the received multimodal input data and transmitting it to a server in real time;

[1816] a data preprocessing means for tokenizing, normalizing, formatting, and sampling rate converting data received by the server;

[1817] A means for an emotion engine to analyze the emotions and needs of a user using natural language processing technology, voice recognition technology, face recognition technology, and facial expression analysis technology using the preprocessed data;

[1818] means for generating customized content based on the analysis results;

[1819] means for transmitting the generated content to a terminal and providing it to a user;

[1820] a means of collecting user feedback and retraining the system's emotion engine to improve performance; and

[1821] A system including:

[1822] (Claim 2)

[1823] 10. The system of claim 1, wherein the multimodal input data includes text, audio, images, and video.

[1824] (Claim 3)

[1825] 10. The system of claim 1, wherein the generated customized content includes text messages providing detailed instructions on how to use the product and reassuring video guidance.

[1826] "Application example 2 when combining emotion engines"

[1827] New Claims:

[1828] (Claim 1)

[1829] a means for collecting multimodal input data;

[1830] means for preprocessing the collected multimodal input data;

[1831] means for analyzing emotions and needs using the preprocessed data;

[1832] means for generating customized content based on the analysis results;

[1833] means for providing the generated content to a user;

[1834] a means of collecting user feedback and improving the performance of the system;

[1835] means for recommending entertainment content based on a user's emotional state and needs;

[1836] A system including:

[1837] (Claim 2)

[1838] 10. The system of claim 1, wherein the multimodal input data includes text data, audio data, image data, or video data.

[1839] (Claim 3)

[1840] 2. The system according to claim 1, wherein natural language processing technology, voice recognition technology, face recognition technology, and facial expression analysis technology are used to analyze emotions and needs. [Explanation of symbols]

[1841] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Device 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robot< / url:> < / url:> < / url:> < / url:>

Claims

1. a means for collecting multimodal input data; means for preprocessing the collected multimodal input data; means for analyzing emotions and needs using the preprocessed data; means for generating customized content based on the analysis results; means for providing the generated content to a user; a means of collecting user feedback and improving the performance of the system; A system including:

2. The system of claim 1 , wherein the multimodal input data includes text data, audio data, image data, or video data.

3. 2. The system according to claim 1, wherein natural language processing technology, voice recognition technology, face recognition technology, and facial expression analysis technology are used to analyze emotions and needs.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A